;Sutra — Argmax Studio
← All productsStatus · Beta
Beta product / Retrieval experiments

Sutra

Sutra explores how thoughtful retrieval and language models can make complex working knowledge easier to navigate and act on — built openly, evaluated honestly, and refined with the people who use it.

PythonLLMsRAG
Sutra interface preview
SutraInterface preview
TypeRetrieval research instrumentStageOpen beta · design partnersCommitment~1 hour weekly working sessionCostFree during beta
The idea

An honest instrument for an unsolved problem.

Sutra exists to study retrieval quality in the wild — openly, measurably, and together with design partners who feel the problem every week.

Why this exists

Retrieval looks solved. Up close, it rarely is.

Dense contracts, internal codenames, questions that only make sense to twelve people — confident demos fail quietly against real working conditions, and the failures rarely announce themselves. Client engagements kept revealing that gap; we wanted somewhere to study it honestly instead of papering over it.

So Sutra runs as an open instrument: every strategy change ships with its evaluation attached, every failure mode becomes a documented pattern, and whatever proves itself here graduates into Anvaya — and into how we build retrieval for everyone else.

In practice

What asking looks like.

Did the new chunking strategy actually help?

You see the comparison against your own gold-standard questions — not a vibe, a delta, with regressions visible right next to the wins.

What happens below the confidence threshold?

Sutra offers candidate sources instead of an answer. Calibrated abstention is treated as a feature here, and measured like one.

Can we see why it decided that?

Decisions and evaluation results are documented as they happen, so partners see the reasoning behind behavior while it is still being shaped — not after it is frozen.

What it holds to

The commitments underneath.

Context quality is the product

Chunking respects document structure, hybrid matching tunes itself on real questions, freshness stays visible at a glance. The unglamorous parts of retrieval get the engineering attention they deserve.

Honest uncertainty

Below its confidence threshold, Sutra presents candidate sources instead of guessing. Where the system stands is always legible — to the people using it, and to us building it.

Evidence travels with change

Nothing ships on intuition alone. Strategy changes carry their evaluation results with them, so progress can be argued with data rather than taste.

Built in partnership

Design partners spend an hour a week shaping direction and see the reasoning behind every decision while it is still being made. Beta access is collaborative by design.

Who it is for

“Your team works with material machines cannot skim — contracts, research, internal systems with a language of their own. Off-the-shelf search returns documents; what you need are answers that understand why a clause matters. And you would rather help shape the tool than wait for one.”

Sutra takes on a small circle of design partners who feel their retrieval problem weekly. The commitment is roughly an hour of working session time; the influence on what gets built is proportional.

Before you ask

Straight answers.

What does beta mean here?

The core retrieval experience is stable and useful; interfaces and integration surfaces are still moving quickly based on partner feedback. Expect change, expect honesty about it.

Who can join as a design partner?

Teams with a concrete, high-value knowledge-retrieval problem and the willingness to spend an hour a week with us shaping the direction.

Will Sutra become a product or fold into Anvaya?

Both paths are open. Patterns that prove themselves graduate into Anvaya; problems that need a different shape may keep Sutra growing as its own instrument.

What does beta cost?

Nothing during the beta beyond the partnership commitment. We are buying learning, and partners get production-grade attention in return.

Beta today

Put Sutra in front of a real problem.

Start with the demo, or send us the hardest question your team answered badly last week.