Evidence before action

Proof for code you didn’t write.

AI agents write more and more of the code teams ship. Barg Labs builds the checks that let someone else verify it: deterministic, offline, and re-runnable by anyone.

Open sourceRuns offlineDeterministicNVIDIA Inception member
agent report vs. record
Illustrative
Tests: 42 passed
CI summary: 42 passed
Closes #117
Linked reference: #117
Files changed: 3
PR changed 4 files, 1 undeclared
Merged 14:02 UTC
Merged 14:02 UTC
Fail: 1 of 4 claims not in the recordre-runnable
Products

Three checks, one standard of evidence.

Cejel checks the repository. Instrument Review checks the tools that judge code. Dunstan checks what a coding agent says it delivered.

Cejel

Released · free and open source

Trust certificates for code.

Cejel checks a repository offline and produces a certificate anyone can re-run and get the same result. Hand it to the person who has to accept the code.

  • Tests, secrets, isolation, CI discipline, and whether the code matches its claims
  • Deterministic: no model call, no network, no signup
  • Abstains when the source is too thin, instead of guessing
  • npm, Homebrew, Docker and a GitHub Action

Instrument Review

Available now

Your AI judge, measured by someone else.

People accept what your scanner, code reviewer or AI judge says without seeing inside it. We test it on cases where the right answer is already known.

  • Cases, predictions and scoring frozen before anything runs
  • A record a third party can re-run
  • You see it first and choose whether it is published, not what it says

Dunstan

Released · free and open source

Hold every agent report to the record.

Coding agents report what they did: files changed, tests passed, issues closed. Dunstan checks each claim the agent declares against the repository and CI record before you merge.

  • Checks the head commit, changed files, references, test and check counts, and timing
  • Fails closed: a claim the record cannot answer is never a pass
  • Deterministic and re-runnable, with no model call
  • A GitHub Action, a CLI and an MCP tool
Done in public

Two AI judges, one test, results published.

We ran the same cases on a commercial judge and on an open model, so the two sit side by side. Both records are public and re-runnable.

430
test cases, built from 50 public pull requests
4 min
to run a commercial judge over all of them
3¢
US, the whole run
3 of 5
of our own predictions were wrong, and published

The cases are constructed, so the rates describe those cases, not how often real reports are wrong.

How we work

Evidence that survives someone else checking it.

Offline by design

Cejel runs where the code lives: no network, no model call, nothing sent to us.

Re-runnable by anyone

Every result comes with the command that reproduces it. The person relying on it can check it themselves.

Abstain, never bluff

When the evidence is too thin, we say so. A missing answer is reported as missing, not as a pass.

Predictions frozen first

We write down what we expect before we run, and publish when we are wrong.

Company

Built by an engineer who has shipped software for 25 years.

Barg Labs is led by Houman Azimi-Nejadi, a software engineer and founder with more than 25 years of experience building complex software systems across multiple domains.

The company works closely with expert collaborators across product, engineering, and web systems, and is based in Vancouver and London.

Barg Labs is a member of NVIDIA Inception, NVIDIA’s program for startups.

NVIDIA Inception Program member

Accepting code you didn’t write?

Tell us what you need to trust. We will tell you what we can check, and what we can’t.

team@barglabs.ai