The AI Runtime · FDE Lab

Your best work is behind an NDA. Build the version you can show.

The AI Runtime FDE Lab is a public client for forward deployed AI work: realistic briefs, open labels, deterministic evaluators, and reference builds documented failure by failure.

Free · Open labels · No signup · Reproducible scores

  • 6 planned engagements
  • 1 shared client
  • Public golden labels
  • None hidden test sets

The problem

Forward deployed work is invisible by design.

The best deployments happen inside client environments, behind NDAs, under production pressure. The architecture is confidential. The data is protected. The postmortem is privileged. The result is a field with no commons: no shared problems, no comparable results, and no easy way to learn from each other's builds.

So the field keeps rediscovering the same lessons in private: where the model belongs, where deterministic policy must own the decision, what the harness must catch, what auditors will ask, what breaks when a model changes, and what the last few points of quality cost.

Open source solved this for software libraries. It has not solved it for deployment craft, because deployment craft is judgment under constraints. Judgment gets better only against comparable problems. Comparable problems require a shared client.

The shared client

Every engagement is a problem at Meridian, a fictional financial-services and healthcare group with realistic deployment problems: auditability requirements, cost ceilings, escalation paths, compliance officers, and production failure modes. Data is public where possible, fully synthetic where regulation demands it.

Each engagement covers one core skill the field is building right now. The skill is what you practice. The Meridian constraint is what makes the practice transfer.

  1. The brief. A client-shaped problem with business context, constraints, and a measurable acceptance bar.
  2. Public golden labels. The ground truth is open so every result can be audited.
  3. A deterministic evaluator. Same submission, same score, every run.
  4. The reference build. Built in the open by The AI Runtime, documented in a full debrief: architecture, failures, dead ends, and cost per unit of accuracy.

The evaluator turns "here is my approach" into "here is my result." When two builds score 74 and 81 on the same labels, the conversation is about the seven points between them.

How it works

  1. 01
    The brief drops.

    Problem, labels, evaluator, and starter repo go public. The acceptance bar is stated up front.

  2. 02
    The reference build ships in the open.

    The AI Runtime builds against the brief and publishes the debrief: what worked, what broke, what it cost, what changes at 10x scale. Every debrief is a newsletter long-form.

  3. 03
    Build it yourself, if you want.

    Fork the repo, run the evaluator, beat the reference score. File a field report as a GitHub issue: approach, score, what broke, cost. Notable builds get linked below and credited in the newsletter.

No registration. No cohort. No scheduled calls. The evaluator is always open, and every engagement stays live after it ships.

What a build gets you

  • A verified, public artifact. A repo with a deterministic score anyone can reproduce. The shareable proxy for expertise you usually cannot show.
  • A field report in the record. What worked, what failed, what it cost, in a comparable format. Failures included; dead ends are the most valuable part.
  • Credit that flows back to you. Notable builds are linked from this page and credited in The AI Runtime. The Lab makes good work legible; it does not absorb it.
  • Six core skills under production constraints. Structured outputs, LLM-as-Judge calibration, document AI, tool use and MCP, context engineering and agent memory, agent observability.

The engagements

Engagement 01 is live. The remaining engagements show the planned arc and may change as the Lab learns from each cycle.

Engagement 01 Live

The Regulator Is Reading Your Backlog

Skill focus: Structured outputs and classification agents at scale

Meridian's consumer-complaint backlog is a regulatory exposure. Build the triage system: route CFPB complaints to product category and escalation tier, scored against public golden labels. The reference build ships with a cost report per 1,000 complaints.

The debrief question What did the last five points of F1 cost you?

Engagement 02

The Client Asks, "How Do We Know It’s Right?"

Skill focus: LLM-as-Judge and evaluation calibration

Meridian's risk team will not accept “the model is good.” Build an LLM-as-Judge over Engagement 01 outputs, then audit the judge itself: agreement with ground truth, position bias, verbosity bias, drift across model versions, and cases where the judge sounds confident and is wrong.

The debrief question When should a judge be trusted, and how do we know?

Engagement 03

Prior Authorization, With a Compliance Officer in the Room

Skill focus: Document AI and structured extraction

Meridian's health plan is drowning in prior authorization requests. Extract structured requests from fully synthetic clinical documents and route them against payer rules. No PHI. No shortcuts. Scored per field.

The debrief question Where does the model belong in a compliance pipeline, and where must deterministic policy own the decision?

Engagement 04

The Agent Needs to Touch Production

Skill focus: Tool use, MCP, and harness design

Classification was the easy half. Meridian now wants the triage agent to act: pull account context, file escalations, and write back to case systems through MCP-connected tools. Build the tool layer and the harness around it, then score the same agent under two harness configurations, head to head, including across a model upgrade.

The debrief question How much reliability lives in the harness rather than the model?

Engagement 05

The Agent Has to Remember, and Prove It

Skill focus: Context engineering and agent memory

A Meridian complaint lives for weeks across dozens of touches. Decide what the agent remembers, for how long, under what retention rules. Engineer the context so a long-lived case never blows the window or leaks across cases. Then instrument the agent so every run can prove it touched only authorized data and stopped when it should have.

The debrief question What should an auditor be able to reconstruct from an agent run?

Engagement 06

The Incident

Skill focus: Agent observability and LLMOps

The capstone. A degraded system meets a replay of live-shaped traffic. Detect it, diagnose it, remediate it, and file the postmortem in the format regulated teams actually use.

The debrief question What did you detect, what did you miss, and what would have prevented it?

The rules of the lab

  1. The evaluator is the acceptance criterion. Published first, deterministic, non-negotiable.
  2. Golden labels are public. Anyone can audit any result, including the reference build.
  3. Failures are part of the record. Every debrief and field report includes the dead ends.
  4. Synthetic data where regulation demands it. The constraint is the lesson.
  5. Your build is yours. Fork the repo, beat the reference, file your report, keep the credit.

Field reports

Builds from the field, linked as they land. Beat the reference score or document an instructive failure, and your report goes here.

File a field report ↗

Should Meridian Deploy It?

Readiness reports for AI systems, agent runtimes, and deployment surfaces assessed through an FDE lens.

FAQ

How much time does a build take?
The reference builds take real effort; a follow-along build runs roughly 10 to 20 hours depending on how far past the reference score you push. Everything is async. Build on your own schedule; engagements never close.
Do I need to be an FDE?
No. The Lab is for anyone doing forward deployed work under any title: FDEs, SAs who own delivery, consultants, in-house deployment teams, applied AI engineers. The evaluator does not care about anyone's title.
Can I use my company's stack? Do I have to open-source my build?
Use any stack. A public repo is preferred, but implementation notes plus a reproducible evaluator run are accepted when your stack can't be shared. The score and the field report are the non-negotiables.
Is this a leaderboard competition?
No. Scores exist to make comparisons concrete, not to crown a winner. The field reports, including the failures, are the product.
What does it cost?
Nothing. Briefs, labels, evaluators, and reference debriefs are free and public. No signup is required to build. Subscribe to The AI Runtime only if you want briefs and debriefs delivered as they ship.

Start here

  1. Start Engagement 01. Clone the repo, run the evaluator, see where you stand against the reference build.
  2. File a field report when you have a score worth talking about, or a failure worth documenting.
  3. Subscribe to The AI Runtime. Every brief and reference debrief ships there first.

The evaluator does not care about anyone's title. Clone it, score your build, and put a number next to your name.