Contents
  1. 1. Why You Can’t Point an LLM at a Broker API
  2. 2. The 5-Layer Architecture
    1. End-to-End: The Voice-to-Trade Pipeline
  3. 3. Mastra as the Orchestrator (L2)
    1. Why Mastra
    2. The Orchestrator and Its Agents
    3. User Memory: Why the System Gets Smarter
  4. 4. The Neuro-Symbolic Boundary: DSL + Sandbox (L3)
    1. The Strategy Intermediate Representation (SIR)
    2. The SIR Compiler
    3. The Sandbox: Isolated Execution
    4. The Explainability Engine (Reverse Parser)
  5. 5. The Compliance Firewall: OPA + Risk Engine (L5)
    1. OPA/Rego: Compliance as Code
    2. The Risk Engine: Dynamic Market Risk
    3. SEBI Compliance Mapping
    4. The Hard Question: Black-Box Algo Classification
  6. 6. Backtesting with VectorBT and the ClickHouse Data Backbone
    1. VectorBT: Deterministic Strategy Validation
    2. ClickHouse: The Analytics Backbone
  7. 7. The Connector Layer (L4): Abstracting Broker Fragmentation
  8. 8. What I’d Do Differently and Open Questions
    1. Constrained decoding vs. post-hoc validation
    2. Semantic cache and OPA re-evaluation
    3. The “zero hallucinated tool calls” claim
    4. Where the guarantee stops
    5. A note on the “neuro-symbolic” label
    6. What’s next
  9. Summary

Most LLM demos end at “the model called a function.” In regulated finance, that’s where the problems start.

When I set out to build āagman — an AI trading platform for Indian markets — the core constraint wasn’t “can an LLM generate a trading strategy.” It was: can we prove, to a regulator, that a language model never caused a non-compliant trade?

That’s a different problem. The first is an ML problem. The second is a systems-design problem. This post is about the second.

āagman is a SEBI-registered Investment Adviser (INA000021951) and an NSE-empanelled algo provider under NSE Circular Ref. 40/2026. Real money flows through it. Real brokers depend on it. The architecture I describe here isn’t a thought experiment — it’s what runs in production, and it exists because the alternative was unacceptable.

The one-line thesis: LLMs propose. Deterministic engines decide. The architecture enforces this — not the prompt.

1. Why You Can’t Point an LLM at a Broker API#

Let me start with the failure modes that shaped every architectural decision.

Hallucinated parameters. A user says “buy 10 Reliance” and the LLM emits qty: 100. The order is structurally valid. It passes any naive schema check. It’s just wrong — and it’s 10x the intended exposure.

Prompt injection reaching execution. If the LLM has direct access to the broker API, a sufficiently adversarial input can cause it to place trades the user never intended. This isn’t theoretical. Any system where a language model holds credentials to a side-effecting API is one injection away from an incident.

No audit trail. SEBI’s algo trading framework requires that every algo order carry a unique identifier, and exchanges retain kill-switch authority over any algo ID. If your execution path runs through an LLM’s unstructured output, you cannot guarantee the order metadata is present, correct, or immutable.

Regulatory liability lands on the broker. Under SEBI’s February 2025 circular on safer participation of retail investors in algo trading, algo providers are agents of the broker. The broker is the principal. The broker carries the regulatory exposure. When a broker connects to your system, they need a guarantee — not a “the model usually gets it right.”

These failure modes don’t have prompt-engineering solutions. They have architecture solutions.

2. The 5-Layer Architecture#

The platform follows a layered microservices architecture. This design decouples the high-latency activities of the AI agents (reasoning, planning) from the low-latency requirements of market data processing and trade execution. Each layer enforces a different class of safety property. Remove any one layer and a specific failure mode becomes possible.

graph TD
    A["L1 · User Surfaces\nIntent Capture · Visualization"] --> B["L2 · Agentic Co-pilot\nMastra AI · Reasoning · Planning"]
    B --> C["L3 · Trading Intelligence\nDSL Engine · Sandbox · Explainability"]
    C --> D["L5 · Compliance Firewall\nOPA/Rego · Risk Engine"]
    D -->|allow| E["L4 · Connector Layer\nBroker Adaptors · Execution"]
    D -->|deny| F["Rejection\n+ Reason Logging"]

    class A,B,C,D amber
    class E green
    class F red
Fig 1 — The 5-layer pipeline. L5 (Compliance) sits in the blocking path between L3 (Trading Intelligence) and L4 (Connectors). Every broker-bound action must pass through it. Read-only queries exit earlier.

Here’s what each layer enforces:

Layer Component Core Tech What breaks if you remove it
L1 · User Surfaces Web (Next.js), Mobile (React Native), Browser Extension Intent capture, visualization, voice input No structured intent reaches the agents — raw text hits the planner directly.
L2 · Agentic Co-pilot Mastra AI, LLMs, pgvector Reasoning, memory, multi-step planning No decomposition of complex requests. “Hedge my Nifty position with BankNifty calls” becomes a single, unmanageable LLM call.
L3 · Trading Intelligence DSL Engine, E2B Sandbox, Explainability Engine Translation, validation, sandboxed execution The LLM’s probabilistic output reaches the broker without deterministic translation. No human-readable explanation of why a trade happened.
L4 · Connector Layer Broker Adaptors (Strategy Pattern), Market Data Aggregator Protocol normalization, execution Broker-specific fragmentation leaks into the upper layers. A Zerodha error code means something different than an Angel One error code.
L5 · Compliance Firewall OPA/Rego (static rules), Risk Engine (dynamic risk) Deterministic rule enforcement, real-time margin checks A structurally valid but non-compliant order reaches the broker. Position limits, segment restrictions, and kill switches don’t fire.

The key architectural property: the LLM process holds no broker credentials and has no network path to the broker. Its only output channel is a string that the orchestrator parses. “LLM suggests” is a property of the network topology, not a prompt instruction.

End-to-End: The Voice-to-Trade Pipeline#

To make this concrete, here’s what happens when a user driving their car speaks into the mobile app: “Buy 50 shares of Tata Motors if it breaks the 15-minute high.”

L1 (User Surfaces): The app’s local Voice Activity Detection module — running edge AI on the device — detects speech in 20ms frames, calculates zero-crossing rate and energy levels. Only when VAD probability exceeds 0.8 does it open a WebSocket and begin streaming Opus-encoded audio. This reduces bandwidth by over 90% compared to continuous streaming. A circular buffer retains the last 1000ms so the initial consonant isn’t lost.

L2 (Agentic Co-pilot): Mastra Voice transcribes the audio. The Intent Router — a lightweight distilled model (GPT-4o-mini), not the reasoning model — classifies this as STRATEGY_GEN with a CONDITIONAL_ENTRY subtype in under 50ms. Before the reasoning model is called, the router checks the semantic cache (Redis, cosine > 0.92). Cache miss. The Orchestrator retrieves the user’s context from pgvector: preferences, position sizing, risk-reward ratio. It constructs a multi-step plan: Get_15min_High → Create_Trigger → Execute_Buy.

L3 (Trading Intelligence): The DSL Engine converts this abstract plan into a concrete JSON definition: {"symbol": "TATAMOTORS", "trigger": "PRICE > HIGH_15M", "action": "BUY", "qty": 50}. This is validated against a strict schema — Zod on the TypeScript side, Pydantic on the Python side. The Explainability Engine generates a push notification: “I will place a trigger order for Tata Motors at ₹980 (15-min high). I verified this fits your ₹1L limit. Is this okay?”

L5 (Compliance Firewall): The JSON object is passed to OPA, deployed as a sidecar to the Execution Service. The policy engine verifies: Is the user allowed to trade equity? Is the order value within free cash? Is Tata Motors on the ban list? Is the F&O segment activated if needed? If valid, it returns ALLOW.

L4 (Connector Layer): The user confirms via push notification. The Connector Layer selects the correct Broker Adaptor (e.g., ZerodhaAdaptor), authenticates using the stored daily token, and places a GTT (Good Till Triggered) order. The Explainability Engine generates a confirmation: “I’ve placed a trigger order for Tata Motors at ₹980 (15-min high).”

The entire pipeline — from voice to confirmed order — completes with the user never leaving their steering wheel.

3. Mastra as the Orchestrator (L2)#

The choice of orchestration framework was one of the more consequential decisions in the stack. We went with Mastra — a TypeScript-native agent framework — over LangGraph, CrewAI, and a hand-rolled solution.

Why Mastra#

Three reasons, in order of importance:

Durable workflows with typed schemas. Mastra lets you define a workflow as a directed acyclic graph where each step has a typed input and output schema. If an agent emits output that doesn’t match the schema, the step fails before the next step sees it. This is critical — we needed the orchestrator itself to enforce type boundaries, not rely on downstream validation alone.

Agent separation as a first-class concept. In Mastra, each agent is a discrete unit with its own system prompt, tool access, and output schema. You don’t get a monolithic “agent that does everything.” You get composable specialists with explicit interfaces. This maps directly to our security model: the Strategy Agent cannot call execution tools. The Execution Agent cannot reason about markets.

TypeScript-native with unified codebase. Our backend and Next.js frontend are TypeScript. Running the orchestrator in the same language means no serialization boundary between the agent layer and the execution layer. When OPA returns a deny, the error propagates through the same type system as the rest of the application. Mastra’s strong typing across the entire stack — from workflow definitions to tool schemas — catches integration bugs at compile time.

The Orchestrator and Its Agents#

The Orchestrator decomposes complex trading requests into a DAG of tasks and routes them to domain-specific agents. When a user says “Build a strategy to buy Nifty if it drops 1% but hedge it with BankNifty calls,” the Orchestrator breaks this into: Market_Scanner → Option_Chain_Analyzer → Strategy_Compiler, with Mastra’s workflow engine propagating typed state between steps.

sequenceDiagram
    participant U as User
    participant IR as Intent Router
    participant O as Orchestrator
    participant SA as Strategy Agent
    participant RA as Risk Analyst
    participant OPA as OPA (Sidecar)
    participant EA as Execution Agent
    participant BA as Broker Adaptor

    U->>IR: "Buy 50 Tata Motors if it breaks 15m high"
    IR->>IR: Classify: STRATEGY_GEN
    IR->>O: intent + user context (pgvector)
    O->>SA: Plan: Get_15m_High → Create_Trigger → Execute_Buy
    SA->>SA: Formulate DSL via JSON Schema
    SA->>RA: Validate against portfolio
    RA->>RA: VaR check · sector correlation · margin
    RA->>U: "Confirm BUY 50 TATAMOTORS trigger at ₹980?"
    U->>RA: ✓ Confirmed
    RA->>OPA: DSL + user context + market state
    OPA-->>OPA: Evaluate Rego policies
    OPA->>EA: ALLOW + reasons
    EA->>BA: ZerodhaAdaptor.placeGTT()
    BA->>EA: order_id confirmed
    EA->>U: "Trigger order placed for Tata Motors at ₹980"
Fig 2 — The full workflow from intent to execution. The Intent Router uses a lightweight model for classification. The Strategy Agent formulates plans. The Risk Analyst validates them. OPA gates execution. The Broker Adaptor handles protocol normalization.

Five specialized agents divide the work:

Strategy Agent — Formulates deterministic plans from natural language. It drafts strategies, runs backtests via VectorBT, and outputs structured plans against a strict schema. It does not place orders or talk to brokers.

Market Scanner — Interfaces with the Market Data Aggregator to filter the instrument universe. “Find all stocks with RSI < 30” becomes a vectorized scan across NIFTY 500 symbols. Returns structured match results with indicator snapshots.

Risk Analyst — Validates every trade against the user’s portfolio before it reaches OPA. Calculates Value at Risk (VaR), checks sector correlation (“You’re already long on 3 auto stocks — adding Tata Motors increases concentration risk”), and verifies real-time margin requirements against the broker’s API.

Execution Agent — Interfaces with the Connector Layer. Converts confirmed plans into broker-ready payloads. Handles retries, executor reassignment, and failure recovery. Does not reason about markets.

Performance Analyst — Post-trade analysis. Generates “lessons learned” entries that feed back into the user’s memory via pgvector, ensuring the Strategy Agent adapts to the user’s evolving preferences over time.

User Memory: Why the System Gets Smarter#

The agents don’t operate in a vacuum. A multilayered memory architecture ensures each interaction is informed by the user’s trading personality:

Resource-Scoped Memory — Permanent facts retrieved for every session: “User is risk-averse,” “User prefers limit orders over market orders,” “User holds 1000 shares of ITC.”

Thread-Scoped Memory — Ephemeral conversation context: “We just discussed Tata Steel.”

RAG Pipeline — When the Strategy Agent drafts a plan, it queries pgvector: “Retrieve past strategies rejected by this user.” If the user previously rejected high-beta strategies, this constraint is injected into the prompt context. The system learns what you don’t want, not just what you ask for.

4. The Neuro-Symbolic Boundary: DSL + Sandbox (L3)#

This is the core of the system — the layer that translates the LLM’s probabilistic output into deterministic, auditable logic. Layer 3 exists because of a critical security constraint: AI never executes code directly. Allowing an LLM to generate Python for execution is a vulnerability. Prompt injection could lead to malicious code execution.

The Strategy Intermediate Representation (SIR)#

The LLM does not emit free-form text that reaches any interpreter. It emits a JSON object conforming to a strict schema — what we call the Strategy Intermediate Representation (SIR). Think of it as an AST for trading strategies, not a config file.

The SIR uses a closed operator set. Operations are enums, not free strings. Indicators are resolved through a registry. Every field is typed and validated. Here’s what a real SIR looks like:

JSON
{
  "strategy_id": "gen_uuid_123",
  "trigger_condition": {
    "type": "indicator_comparison",
    "left": {"type": "indicator", "name": "SMA", "period": 50},
    "operator": "CROSS_ABOVE",
    "right": {"type": "indicator", "name": "SMA", "period": 200}
  },
  "execution_logic": {
    "action": "BUY",
    "order_type": "LIMIT",
    "limit_price": "LAST_TRADED_PRICE * 1.001"
  }
}

This is far more expressive than a simple {op: "PLACE_ORDER", qty: 50} — it encodes conditional triggers, indicator-based logic, and execution parameters. But every element maps to a pre-compiled, safe function. CROSS_ABOVE maps to a specific math function. SMA maps to a pinned, deterministic implementation (pandas rolling mean, adjust=False for EMA). The LLM selects the blocks; the engine assembles them.

The SIR Compiler#

The SIR goes through a compilation step before execution — exactly like source code:

  1. Schema validation (Zod on TypeScript, Pydantic on Python). additionalProperties: false. Numeric fields have type and bound constraints. Operators are an enum. symbol is checked against the instrument master — not accepted as free text.
  2. Default resolution. Fees and slippage default to 0. Session timing is normalized to UTC. A deterministic sir_hash is computed so identical strategies always produce identical plans.
  3. Enum resolution. Indicator types and operators are resolved against the indicator registry. Unsupported features fail fast with clear errors — no silent fallbacks.
  4. Compilation to ExecutionPlan. The SIR compiles into a strongly typed ExecutionPlan (Pydantic models: UniverseSpec, IndicatorPlan, RulePlan, RiskPlan, SizingPlan). This is the internal representation that the execution adapters consume. The same SIR compiles identically for both backtesting and screening — the intent field (BACKTEST | SCREENER) determines the execution path, not the logic.

On validation failure at any step, the request is rejected, not repaired. A retry is allowed, but the retried output goes through the full compiler again. The LLM never patches its own output into the execution path.

The Sandbox: Isolated Execution#

Even with a typed DSL, the evaluation of triggers — calculating indicators on live data, evaluating cross-above conditions — requires computation. That computation runs inside a sandbox with no escape path.

The DSL interpreter runs inside an E2B sandbox (with Firecracker as the isolation backend). The constraints:

  • No network access. The sandbox cannot reach the internet, the broker, or any internal service.
  • No filesystem access beyond the mounted data volume (read-only) and output directory (write-only).
  • No environment variables. No leaked credentials, no configuration injection.
  • Deterministic I/O. The sandbox accepts input (market data candles from mounted Parquet files) and returns output (boolean trigger signals, performance metrics). Nothing else.

This means that even if a strategy is malformed — even under active prompt injection — the worst case is a wrong computation that the sandbox reports as output. It cannot compromise the host system, exfiltrate data, or reach a broker API.

For backtesting specifically, the sandbox mounts historical OHLCV data as read-only Parquet files and writes results (summary JSON, trades Parquet, equity curve Parquet) to a write-only output directory. The orchestration layer retrieves these artifacts after the sandbox exits, stores time-series data in ClickHouse and summaries in PostgreSQL.

graph LR
    A["LLM Output\n(SIR JSON)"] --> B{"SIR Compiler\nSchema + Enums + Defaults"}
    B -->|invalid| C["Reject\n+ clear error"]
    B -->|valid| D["ExecutionPlan\n(Pydantic models)"]
    D --> E{"E2B Sandbox\nNo network · No FS · No env"}
    E --> F["Indicator Registry\nRSI · SMA · EMA"]
    F --> G["Rule Engine\ncrosses_above · comparisons"]
    G --> H{"OPA Sidecar\nRego Policies"}
    H -->|deny| I["Reject\n+ deny reasons"]
    H -->|allow| J["Broker Adaptor\nExecute"]

    class B,H amber
    class E violet
    class C,I red
    class J green
Fig 3 — The full path from LLM output to execution. Three independent barriers (compiler, sandbox, OPA) each have the authority to reject. Default is deny.

The Explainability Engine (Reverse Parser)#

Trust is established when the AI explains why it acted — in plain language, not JSON.

The Explainability Engine takes trigger signals (e.g., RSI=28, Volume=1.5x_avg) and reverse-parses them into human-readable narratives using specialized prompt templates: “I entered this trade because the RSI dropped to 28, indicating an oversold condition, while volume spiked 20% above the 10-day average, suggesting institutional buying interest.”

This reverse-parsed explanation is what the user sees in push notifications and in the trade log. It’s also what gets stored in the audit trail — a human-readable record alongside the raw DSL, so a compliance review can trace both the machine logic and the plain-English reasoning.

5. The Compliance Firewall: OPA + Risk Engine (L5)#

Layer 5 is the platform’s key differentiator. While Layer 2 (Mastra) is designed to be creative, Layer 5 is designed to be obstructive. It is a deterministic firewall that blocks any action violating regulatory or safety rules.

The compliance firewall has two components that enforce different things:

Component Enforces Type of Safety
OPA/Rego Static rule compliance: SEBI regulations, segment restrictions, ban lists, exposure limits Regulatory
Risk Engine Dynamic market risk: real-time margin, VaR, sector correlation, kill switch Market risk

OPA/Rego: Compliance as Code#

Open Policy Agent is deployed as a sidecar to the Execution Service. Every API call to place an order must pass through an OPA authorization check — there is no code path that bypasses it. The sidecar pattern means the execution service literally cannot reach the broker without OPA’s approval.

The workflow:

  1. Execution Service receives DSL payload: Order(User=Ajit, Symbol=TATASTEEL, Qty=100).
  2. Execution Service queries OPA: POST /v1/data/trading/allow {input: {user: Ajit, order: Order}}.
  3. OPA evaluates Rego policies against the input.
  4. OPA returns: {"result": false, "reason": "Exposure limit exceeded"}.
  5. Execution Service throws an exception. The order never reaches the broker.

The critical design choice: OPA receives its input from trusted code, not from the LLM. The policy input is assembled by deterministic services:

  • op and args come from the validated SIR (the LLM’s contribution, post-compilation)
  • user.id, user.segment, user.limits come from the identity service
  • context.ltp (last traded price), context.market_open, context.vix come from the market data service

The LLM never populates trust-bearing fields. If it could, the boundary would be fake.

Here are concrete Rego policies mapping to specific SEBI requirements (illustrative, not our production policies):

Exposure check — SEBI requires brokers to ensure clients don’t exceed exposure limits based on reported net worth:

Rego
package trading.compliance
default allow = false

allow {
    input.user.is_active
    exposure_check
}

exposure_check {
    projected := input.user.current_exposure + (input.order.price * input.order.quantity)
    projected <= input.user.max_exposure_limit
}

F&O segment activation — Users cannot trade derivatives unless they’ve uploaded income proof and activated the segment:

Rego
deny[msg] {
    input.order.segment == "FNO"
    not input.user.segments.fno_active
    msg := "Trade Rejected: F&O segment is not activated for this user."
}

Ban list enforcement — No fresh positions in securities under the F&O ban period:

Rego
deny[msg] {
    input.order.action == "OPEN"
    data.banned_securities[input.order.symbol]
    msg := sprintf("Trade Rejected: %v is currently in F&O ban period.", [input.order.symbol])
}

The defaults matter: OPA timeouts, errors, and undefined results all resolve to deny. The system is fail-closed.

The Risk Engine: Dynamic Market Risk#

While OPA handles static regulatory rules, the Risk Engine handles what changes by the millisecond:

Real-time margin calculation. Before execution, the engine queries the broker’s API to get exact margin requirements for the specific order (including spanning). It compares required margin against free cash. If Required > Free, it blocks the trade — preventing broker-side rejection penalties.

Kill switch. Per SEBI guidelines, the system includes a kill switch. If the Risk Engine detects a logic loop (e.g., an algo placing 10+ orders per second), or if the user hits a global “Max Loss” for the day, the kill switch triggers OPA to update a global allow = false policy for that user, instantly freezing all activity. This works even if the LLM, the orchestrator, and the interpreter all “want” to execute — because OPA is in the blocking path and defaults to deny.

Portfolio correlation. The Risk Analyst agent feeds into this: “You are already long on 3 auto stocks; adding Tata Motors increases sector concentration risk.” This is a soft warning, not a hard block — but it’s surfaced to the user in the confirmation step.

SEBI Compliance Mapping#

āagman operates under SEBI’s algo trading framework, anchored by the “Safer Participation of Retail Investors in Algorithmic Trading” circular (4 February 2025). The framework became mandatory for all brokers on 1 August 2025, after a timeline extension. Under this framework, algo providers are agents of the broker, and the broker carries the regulatory exposure.

Here’s how each requirement maps to an enforcement layer:

graph TD
    subgraph "OPA/Rego (Static Compliance)"
        A["Permitted instruments\nper user segment"]
        B["Exposure limits\n(net-worth based)"]
        C["F&O segment activation"]
        D["Ban list enforcement"]
        E["Market hours check"]
    end

    subgraph "Risk Engine (Dynamic)"
        F["Real-time margin\ncalculation"]
        G["Kill switch\n(loss cap · order rate)"]
        H["Portfolio VaR\n+ correlation"]
    end

    subgraph "Rate Limiter → OPA Input"
        I["10 OPS threshold\nmonitoring"]
    end

    subgraph "Interpreter / Order Gateway"
        J["Algo ID stamped\non every order"]
    end

    subgraph "Infrastructure"
        K["Static IP whitelisting"]
        L["Per-client API key"]
        M["OAuth + 2FA"]
    end

    subgraph "Organizational"
        N["SEBI IA: INA000021951"]
        O["NSE empanelment: 40/2026"]
    end

    class A,B,C,D,E amber
    class F,G,H blue
    class I violet
    class J green
    class K,L,M gray
    class N,O red
Fig 4 — SEBI compliance enforcement map. OPA handles static regulatory rules. The Risk Engine handles dynamic market risk. The rate limiter is stateful and feeds its count into OPA as input.

The non-obvious ones worth explaining:

10 OPS threshold. SEBI’s framework requires that algo orders exceeding 10 orders per rolling second per exchange must be registered with a strategy ID. This is a stateful computation — you need a counter, not a policy rule. Our rate limiter tracks the count and feeds it into OPA as an input field. OPA itself doesn’t count; it just checks input.rate.current_ops > 10 and enforces the appropriate tagging.

Algo ID on every order. The interpreter stamps the algo ID on every order — place, modify, cancel. OPA has a secondary check: if the algo ID field is absent or malformed, the order is denied. Belt and suspenders.

Immutable audit trail. OPA decision logs — every evaluation’s input, the policy version that was applied, and the result — feed into ClickHouse. This gives you a complete, time-ordered, immutable record of every compliance decision. When a broker needs to answer “why did this order execute?” in a regulatory review, the answer is a ClickHouse query, not a post-hoc reconstruction.

The Hard Question: Black-Box Algo Classification#

SEBI’s framework distinguishes between white-box algos (transparent, replicable logic) and black-box algos (logic not known to the user or not replicable). Black-box providers must register as research analysts and maintain a research report per algo.

The open question: when a user says “buy 10 Reliance” and the LLM parses that into an order, is that a black-box algo? If the LLM chooses the strategy or parameters — rather than the user — the logic is arguably not replicable, and the RA registration question opens.

I’m not going to pretend this is settled. The circular’s definitions are the only text that decides it, and there’s an ongoing clarification process. Our position: when the user explicitly specifies the instrument, quantity, and order type, āagman is an execution tool. When the system generates strategy parameters autonomously (e.g., from a screener), the classification may be different, and we’ve built the system to support per-algo registration for that scenario.

6. Backtesting with VectorBT and the ClickHouse Data Backbone#

Two parallel pipelines complete the system: one validates strategies before any money is at risk, and the other makes every production decision auditable after the fact. Both converge on ClickHouse.

VectorBT: Deterministic Strategy Validation#

VectorBT handles all backtesting. The engine runs inside the E2B sandbox, consuming historical OHLCV data from Parquet files mounted read-only. It’s pinned to a specific version (0.26.2) because unpinned versions can cause result drift — backtesting demands bit-for-bit reproducibility.

The execution flow for a backtest:

  1. SIR compilation produces an ExecutionPlan with intent: BACKTEST.
  2. Parquet path resolver deterministically generates file paths from universe, timeframe, and date range: /parquet/venue=NSE/asset=EQUITY/tf=1m/date=2026-01-03/symbol=RELIANCE.parquet.
  3. OHLCV loader reads Parquet files, concatenates daily files, sorts by timestamp, enforces schema (open, high, low, close, volume), sets UTC index, and applies warmup bars.
  4. Indicator registry computes indicators (RSI, SMA, EMA) — each pinned to a specific implementation (e.g., pandas.ewm(span=period, adjust=False) for EMA) to prevent variant drift.
  5. Rule engine evaluates conditions (crosses_above, crosses_below, comparisons) into boolean Series, evaluated at bar close. No eval(), no dynamic code — operator whitelist only.
  6. BacktestAdapter shifts entry signals by +1 bar (entries fill at next bar open, not the signal bar — a critical realism constraint), applies stop-loss and take-profit logic intrabar with SL priority, and runs VectorBT Portfolio.from_signals().
  7. Serialization writes summary JSON, trades Parquet, and equity curve Parquet to the output directory.

Golden tests ensure determinism: fixed Parquet fixtures with hand-verified expected values. The same SIR always produces the same metrics. Any change to the engine requires explicit test updates and approval. This is how you protect user trust.

For screening (not backtesting), the same SIR compiles but routes to the ScreenerAdapter instead. It processes the NIFTY 500 universe using batch-vectorized evaluation: symbols are partitioned into chunks of 100, OHLCV data is pivoted from long to wide format, and indicators are computed across all columns simultaneously — no per-symbol loops for math. Memory is explicitly cleared between batches. This processes 500+ symbols in under 2GB of memory with throughput 10x faster than sequential evaluation.

ClickHouse: The Analytics Backbone#

ClickHouse isn’t just the audit trail — it’s the time-series backbone for the entire platform. It stores:

Market data. OHLCV time-series for equities, commodities (MCX, NCDEX), and F&O contracts. Mutual fund NAV history. All using ReplacingMergeTree for idempotent ingestion with ingested_at as the version column.

Backtest artifacts. Time-series results (equity curves, trade logs) from every backtest run, stored alongside the SIR hash for reproducibility.

OPA decision logs. Every compliance evaluation — input, policy version, result, deny reasons — forms the regulatory audit trail.

Order lifecycle events. Every state transition of every order, with timestamps and upstream context.

The Parquet export pipeline bridges ClickHouse and VectorBT: finalized bars are exported daily from ClickHouse to partitioned Parquet on object storage, which the backtesting sandbox mounts read-only. Live writes go to ClickHouse continuously; Parquet exports write “yesterday” to avoid rewriting files.

graph LR
    subgraph "Data Ingestion"
        A1["NSE / MCX / NCDEX\nBhavcopy"] --> CH["ClickHouse\n(ReplacingMergeTree)"]
        A2["AMFI\nMF NAV History"] --> CH
        A3["Broker WebSocket\nLive Tick Data"] --> CH
    end

    subgraph "Parquet Export"
        CH --> PQ["Partitioned Parquet\n/venue/asset/tf/date/symbol"]
    end

    subgraph "Backtesting (Sandbox)"
        PQ -->|read-only mount| VBT["VectorBT\nE2B Sandbox"]
        VBT --> RES["Results: JSON + Parquet"]
        RES --> CH
    end

    subgraph "Compliance Audit"
        OPA["OPA Decision Logs"] --> CH
        ORD["Order Lifecycle"] --> CH
        OTEL["Mastra OTEL Traces"] --> CH
        CH --> API["Explainability API\nfor broker audits"]
    end

    class CH blue
    class VBT amber
    class PQ,API green
Fig 5 — ClickHouse is the convergence point: market data ingestion, Parquet exports for backtesting, backtest results, and compliance audit logs all flow through it.

The explainability API is what makes the system auditable for broker customers. When a broker asks “why did order X execute at time T?”, the API returns the full decision chain: the user’s original message, the Strategy Agent’s output, the parsed SIR, the confirmation step, the OPA evaluation (with policy version and all inputs), and the execution result. This isn’t reconstructed — it’s a query over structured, immutable logs.

7. The Connector Layer (L4): Abstracting Broker Fragmentation#

A quick note on Layer 4, because the problem it solves is invisible when it works — and catastrophic when it doesn’t.

There is no standard API for Indian brokers. Zerodha (Kite Connect), Groww, Angel One (SmartAPI), and Upstox all use different authentication flows, order payloads, error codes, and WebSocket formats.

The Connector Layer implements the Strategy Pattern: a generic Broker interface (placeOrder, cancelOrder, getHoldings) with concrete adaptors — ZerodhaAdaptor, AngelAdaptor, etc. Each adaptor is responsible for:

  • Payload normalization: mapping the platform’s canonical order format to the broker’s expected structure.
  • Error normalization: mapping broker-specific error codes (“Network Exception” vs “Insufficient Funds”) into a standard PlatformError enum. This lets the Agentic Layer understand why a trade failed — if “Insufficient Funds,” the agent suggests lowering quantity; if “Network Exception,” it retries.
  • Authentication management: storing and refreshing daily tokens, handling session expiry.

The Market Data Aggregator maintains persistent WebSocket connections to multiple brokers and market data providers. It de-duplicates feeds (if multiple brokers provide RELIANCE data, it arbitrates to the most timely feed), synthesizes tick data into 1-minute, 5-minute, and 15-minute OHLCV candles in real-time memory buffers, and publishes completed candles to Redis Pub/Sub channels that the Strategy Agents subscribe to.

For brokers without public APIs — or users on legacy brokerage plans — the platform provides a Chrome Extension that acts as a robotic DOM interface. The extension uses Manifest V3 with signed messaging: every execution command from the backend is digitally signed, and the Content Script verifies the signature using a baked-in public key before executing a DOM click. This prevents man-in-the-browser attacks where a compromised web page fakes a command.

8. What I’d Do Differently and Open Questions#

No architecture post is complete without the honest section. Here’s what I’d reconsider, and what remains genuinely unsettled.

Constrained decoding vs. post-hoc validation#

Our current approach validates the LLM’s output after generation. An alternative — constrained decoding — would force the LLM to only generate tokens that conform to the DSL grammar during inference. This is theoretically cleaner (malformed output is impossible, not just rejected) but practically harder with hosted model APIs where you don’t control the decoding loop. If we ever move to self-hosted models, this is the first change I’d make.

Semantic cache and OPA re-evaluation#

We use a semantic cache (cosine similarity > 0.92, keyed on market-state hash) to short-circuit the LLM for repeated queries. The design question: should a cache hit on an executable action return the cached DSL and skip OPA? The answer is no — policy depends on live state (market hours, current position size, daily loss), so a cached decision would be stale. The cache returns the DSL; OPA re-evaluates against current state. This is the correct design, but it means the cache saves LLM latency, not total latency.

The “zero hallucinated tool calls” claim#

This is the claim I’m most careful with. What the architecture actually guarantees: no malformed or non-compliant action can execute. What it does not guarantee: the LLM never proposes a wrong trade. The confirmation step, position limits, and OPA bounds close most of the residual risk, but they don’t eliminate it. The scoped version of the claim is the only one I’d defend in a technical review.

Where the guarantee stops#

The hostile question: “Zero hallucinated tool calls. What about a valid call with hallucinated arguments?”

The DSL makes invalid operators unrepresentable. It cannot make valid operators correct. If the user says “buy 10 Reliance” and the LLM emits qty: 100, symbol: RELIANCE, that passes the schema. It passes OPA if 100 shares of Reliance is within the user’s limits.

The mitigations: the confirmation step (user always sees the parsed order), OPA bounds (position limits and notional caps), and explainability logging (the full chain is traceable). The scoped claim that survives scrutiny: “The LLM can’t cause an action outside the allowed operator set, and can’t cause an allowed action outside policy limits. Residual risk is a wrong-but-permitted parameter, bounded by limits and caught by the confirmation step.”

A note on the “neuro-symbolic” label#

Classic neuro-symbolic AI integrates symbolic reasoning into the learning or inference process. This system uses a symbolic constraint layer around a neural component. The honest framing: neural proposal with symbolic verification. The value is the guarantee — not the label.

What’s next#

The architecture I’ve described is what runs today. What’s coming: execution algorithms (TWAP and VWAP are live; iceberg and POV are next), multi-leg options support for up to four legs placed in sync, and deeper integration with the VectorBT pipeline for continuous strategy evaluation in production — running the same backtesting engine against live data to detect strategy drift in real time.

Summary#

The system works because of a single architectural principle: the LLM has no path to the broker. Everything else — the typed SIR, the compiler, the sandbox, OPA as a sidecar, the Risk Engine, the confirmation step, the audit trail — is a consequence of taking that principle seriously and closing every gap between “the LLM emitted some text” and “money moved.”

If you’re building agentic systems in regulated domains, the framework generalizes: separate proposers from executors at the network level. Make the execution boundary typed, policy-gated, and auditable. Accept that the LLM will sometimes be wrong, and design the system so that “wrong” is bounded and visible, not silent and unbounded.

The labels — neuro-symbolic, agentic AI, whatever — matter less than the properties. Can you prove, to a regulator, that no non-compliant action was executed? Can you show the full decision chain for any order? Can you guarantee that the system fails closed?

If yes, ship it. If no, fix the architecture, not the prompt.


āagman is built by KotiLabs. I’m Aman Jain. If you’re building agentic systems under regulatory constraints, I’d love to talk — reach me at amanjain.codes.