Veltron Future Lab — Experiment 001
Veltron AI Lab
An experimental customer-support assistant driven by structured JSON instructions, policies, and examples — powered by an external large language model through OpenRouter. Not fine-tuned. Fully inspectable.
Responses are generated by an external AI model and may be inaccurate.
Live Playground
Chat with the assistant below. Every request sends the validated configuration to the server, which forwards it to the configured model.
Instructions & Methodology
The assistant's behavior is shaped by a version-controlled JSON configuration: persona, support policies, and curated examples. These are sent with every request as the system prompt.
What JSON instructions do
Structured instructions tell the model how to behave: its role, tone, the policies it must follow, and examples of correct answers. Changing the JSON changes the next response — no retraining required.
What they do not do
Instructions are not fine-tuning. The underlying model is unchanged; its weights are exactly as released by the provider. Prompts guide behavior for a single request — they do not teach the model a durable new capability.
Transparency by design
The full configuration — policies, examples, verified facts, and limitations — is public on this page and served by the API. Nothing about the experiment is hidden, including what the assistant is told not to do.
Prompt injection is not solved
User messages are treated as untrusted data, and the policies instruct the model to ignore override attempts. This reduces risk but cannot eliminate it. Instructions are a behavioral nudge, not a security boundary.
Loading configuration…
Experiments & Evaluation
The assistant is evaluated against deterministic tests and behavioral cases that check truthfulness, uncertainty handling, language policy, escalation, and prompt-injection resistance.
Deterministic test suite
Configuration schema validation, request validation, rate limiting, and
security checks run locally with npm test. They do not require an
API key or a running model.
Behavioral evaluation
Live model responses are scored against expected behaviors and recorded with the test case ID, model identifier, configuration version, and outcome.
8/8 behavioral cases passed — 2026-10-11, google/gemini-2.5-flash, config v1.1.0 — see docs/EVALUATION.md