vaish.ai
Connect
Intelligence OS · Runtime v0.3 · Independent research studio

The prefrontal cortex
for generative AI.

Agentic systems that reason causally, not just respond. Built at the intersection of frontier research and production-scale engineering.

Reasoning engine
ARCH · in progress
Reliability suite
CORTEX · in design
Adversarial layer
CART · in build
Status
All systems nominal
00What is Vaish.ai

An intelligence OS
for high-stakes reasoning.

Structured, causal reasoning on top of statistical engines. Built for decisions that cannot afford to be wrong.

HIGH-STAKES DECISIONS reasoned · verified · on the record VAISH.AI · INTELLIGENCE OS ARCH reasons · causally CORTEX measures · reliability CART attacks · continuously necessary friction STATISTICAL ENGINE frozen model · pattern matching · never retrained
OS/FIG-01 · The operating layer sits between a frozen engine and the decisions it informs. ARCH reasons, CORTEX measures, CART attacks.
/01

Prototype lab

Research hypotheses, engineered into production-grade systems.

/02

Research studio

Peer-reviewed work in causal reasoning, RL, and reliability.

/03

Intelligence OS

One runtime: ARCH reasons, CORTEX measures, CART attacks.

01ARCH · The reasoning engine
Framework stage · Whitepaper in progress

ARCH

Autonomous Reasoning & Causal Hierarchy

A metacognitive executive layer for frozen models. A learned controller plans over an explicit structural causal model, and no conclusion commits without passing a counterfactual consistency gate. The base model is never retrained.

MechanismModel-based planning over an explicit SCM
ValidationCounterfactual gate · commit / abstain / refine
ControlRL executive (PPO / GRPO) · sole writer to memory
StatusGate + buffer specified · reward design open
Spec sheet
StackPyTorch · JAX · PPO / GRPO · frozen-LLM API
Interfacesproposer ↔ executive · typed memory buffer · SCM store
Gate checkabduction → action → prediction · consistency or abstain
Benchmarkscounterfactual consistency suite · in design
PROPOSER frozen model · never retrained hypotheses queries EXECUTIVE model-based planner · PPO / GRPO plan · simulate · check · commit MEMORY typed state executive-owned candidate plan CAUSAL VALIDATION LAYER structural causal model · do-calculus Z X Y do(X) cuts Z → X · effect isolated GATE consistent? COMMIT abstain refine
ARCH/FIG-01 · Model-based planning above, causal validation below. Nothing commits without surviving the gate; failure refines or abstains, never fails silently.
02Inside ARCH · Pearl's causal hierarchy

Climbing the ladder of causation.

Standard LLMs are trapped on the first rung: advanced statistics engines, pattern-matching over what they have seen. ARCH is built to operate on the third, where a conclusion must survive the question: what if we had acted differently?

each rung asks a harder question than the one below it ASSOCIATION · OBSERVING P(y | x) simple pattern matching over history, where standard LLMs are trapped "What do past loan defaults look like when interest rates are 5%?" INTERVENTION · DOING P(y | do(x)) mapping direct action-to-reaction pathways · acting, not recalling "What happens to our default rate if we set rates to 4.5% tomorrow?" COUNTERFACTUALS · IMAGINING P(yx' | x, y) retrospective simulation · isolating one variable against historical constants "Given that defaults were 3% at a 5% rate, what would have happened if we had lowered rates last quarter?" ↑ the domain of ARCH's counterfactual gate LEVEL 1 LEVEL 2 LEVEL 3
ARCH/FIG-02 · Pearl's ladder of causation. Correlation lives on the bottom step; ARCH commits nothing that has not survived the top one.

The counterfactual gate · eliminating statistical guesswork

01 · DETECTION 02 · INTERCEPTION 03 · CAUSAL STRUCTURING 04 · RESOLUTION raw query · tangled correlations flagged: level-3 counterfactual halts the raw prompt: no semantic guessing RATES history held constant AFFORDABILITY mediator DEFAULT RISK modified variable · Δ 0.5% a strict causal graph: only the isolated variable moves mechanically logical output, or abstain
ARCH/FIG-03 · The counterfactual gate, stage by stage: parse and flag, halt before the model can guess, rewrite against the causal graph, answer only through the constrained pathway.
03CORTEX · The reliability suite
Design stage · Benchmark report pending

CORTEX

Model Reliability Evaluation Suite

An open evaluation harness that measures reliability, not capability: the same tasks re-run under drift, tool degradation, and adversarial content, scored as a degradation profile instead of a single number. ARCH is the first test subject.

MeasuresDegradation slope · calibration under shift
MethodBaseline → perturb → adversarial search
OutputA reliability profile, not a scalar score
StatusTaxonomy specified · harness unbuilt
Spec sheet
HarnessPython · containerized task suites · open source
Perturbationsdrift/ · tools/ · inject/ · state/
Profile outputdegradation slope · confident-error rate · abstention quality
First targetARCH · then open agentic models
cortex/ open source · evaluation harness baseline/ reference run · capability substrate perturb/ four classes of production pressure drift/APIs · schemas · distributions tools/failures · latency · partials inject/hostile content in retrieved data state/mid-trajectory corruption generator/ adversarial search over conditions escalate profile/ output degradation slope per class confident-error rate under shift abstention appropriateness a profile, not a scalar
CORTEX/FIG-01 · The suite as its repository reads: baseline, four perturbation classes, a generator that escalates instead of replaying, and a profile as the only output.
04CART · The adversarial layer
Build stage · Independent research

CART

Continuous Adversarial Red Teaming

A standing adversarial agent for regulated AI. One engine attacks a deployed system continuously and writes every finding into an audit-ready record, built for the Three Lines of Defense structure that governs AI in financial services.

MechanismPlanner · executor · judge · memory loop
CadenceContinuous · not point-in-time PDFs
WedgeThree Lines of Defense · one engine, three SKUs
StatusLoop v1 in build · open-source targets first
Spec sheet
Loopplanner · executor · judge · memory of attack lineages
Targetsopen-source agent stacks first · deployed systems next
Recordtimestamped · reproducible · mapped to controls
Packagingone engine · Three Lines of Defense · three SKUs
PLANNER selects pressure EXECUTOR runs the attack attack TARGET deployed AI system response JUDGE scores the outcome judged MEMORY attack lineages escalate where weak every finding LIVING RECORD timestamped · reproducible mapped to controls continuous assurance for a system that does not hold still one engine · three lines · three SKUs
CART/FIG-01 · The loop attacks the way failures arrive: constantly and adaptively. Memory makes it escalate; the record makes it evidence auditors can consume.
05The architect

Vaishnavi Yeruva.

Published researcher. Production engineer. Eight-plus years building AI systems at the scale of hundreds of millions of users, now building the Vaish.ai stack. Full story, publications, and credentials on the architect page.

Portrait of Vaishnavi Yeruva The Architect · Vaish.ai
06Initialize connection

Working on something that demands real AI reasoning?

Whether you're a researcher, founder, or hiring for frontier AI roles, let's talk. Research conversations and collaborations are always welcome.