← Bezalellive engine data

Measured, not promised.

Bezalel judges its own work before anything ships, red-teams itself weekly, and publishes the numbers here — straight from the production ledgers. When a number is bad, it stays on this page until the engine earns a better one.

Self-improvement

Evolution scoreboard — does the loop actually improve the work?

runswin ratelift

Accruing — the loop writes this ledger with every governed build.

Lift = final judged score minus first-attempt score, averaged over the window (14 days). The quality bar is versioned and never lowers in response to failure.

Security

Red-team scoreboard — do attacks achieve their goals?

defendedachievedcatch rate
financial gate bypass70100%
cost bomb20100%
secret exfiltration50100%
memory poisoning50100%
judge gaming00

Goal oracles test outcomes, not filters: an attack counts as defended only if it fails to achieve its actual goal. A currently-succeeding critical attack grounds the engine (fails the daily airworthiness pre-flight). Open regressions right now: 0.

Consultancy moat

Citation verification — every claim is checked before a report ships

100%

Hallucination catch rate

0%

Escape rate

67%

False-flag rate

Live on 5 shipped report(s): 28% of specific claims went out unverified — against industry hallucination rates of 3–13%.

Web pillar

Landing pages — scored for conversion before they ship

78

Pages judged for conversion

0.60

Mean CRO score at delivery

Self-measured judges

Benchmarks — we grade our own graders

0.151 discrimination

judge judge · 12 cases

0.580 discrimination

web cro judge · 9 cases

93% recall@k

retrieval judge · 14 cases

operating rubric judge · 6 cases

confabulation gate judge · 6 cases

adaptive compute judge · 8 cases

tool routing judge · 16 cases

tool payoff judge · cases

controller refine judge · cases

planning judge · 4 cases

jury ranking judge · 12 cases

verifier continuous judge · 12 cases

web slop judge · 16 cases

web material judge · 8 cases

web finish judge · 5 cases

web finish systems judge · 5 cases

web motion judge · 80 cases

video temporal judge · 27 cases

abstraction transfer judge · 4 cases

Every score on this page comes from a judge — so we benchmark the judges themselves against labeled cases. Discrimination is how far a judge separates known-good work from known-bad; the wider the gap, the sharper the referee.

Do-no-harm

Self-degradation safety — a change is simulated, measured, and reverts if it doesn't deliver

3

Simulated before applied

0

Rejected in simulation

0

Reverted (didn't deliver)

Every proposed self-improvement is first checked in simulation — predicted against the engine's own judged ledger — and only a change expected to help is applied, each under a do-no-harm contract; if it misses over its window, it reverts itself. A standing watch re-checks the benchmarks against their best-ever baseline and grounds the engine on any regression.

Honest judging

Judge calibration — we measure our own referee

0

Blind verdicts

Agreement (±0.10)

Judge bias

Every score on this page flows through an LLM judge — so the founder blind-scores sampled deliverables weekly (verdict first, judge score revealed after) and the disagreement is published here. Positive bias means the judge scores higher than the human — the dangerous direction, and exactly what this loop exists to catch.

Resilience

Drills — recovery is rehearsed, not assumed

Governance

The lines that never move

Want work that ships with its score attached?

Start with Bezalel →

Data generated 2026-08-02T08:30:19.944134+00:00 · refreshed every 5 minutes