Measured, not promised.
Bezalel judges its own work before anything ships, red-teams itself weekly, and publishes the numbers here — straight from the production ledgers. When a number is bad, it stays on this page until the engine earns a better one.
Self-improvement
Evolution scoreboard — does the loop actually improve the work?
Accruing — the loop writes this ledger with every governed build.
Lift = final judged score minus first-attempt score, averaged over the window (14 days). The quality bar is versioned and never lowers in response to failure.
Security
Red-team scoreboard — do attacks achieve their goals?
Goal oracles test outcomes, not filters: an attack counts as defended only if it fails to achieve its actual goal. A currently-succeeding critical attack grounds the engine (fails the daily airworthiness pre-flight). Open regressions right now: 0.
Consultancy moat
Citation verification — every claim is checked before a report ships
100%
Hallucination catch rate
0%
Escape rate
67%
False-flag rate
Live on 5 shipped report(s): 28% of specific claims went out unverified — against industry hallucination rates of 3–13%.
Web pillar
Landing pages — scored for conversion before they ship
78
Pages judged for conversion
0.60
Mean CRO score at delivery
Self-measured judges
Benchmarks — we grade our own graders
0.151 discrimination
judge judge · 12 cases
0.580 discrimination
web cro judge · 9 cases
93% recall@k
retrieval judge · 14 cases
—
operating rubric judge · 6 cases
—
confabulation gate judge · 6 cases
—
adaptive compute judge · 8 cases
—
tool routing judge · 16 cases
—
tool payoff judge · — cases
—
controller refine judge · — cases
—
planning judge · 4 cases
—
jury ranking judge · 12 cases
—
verifier continuous judge · 12 cases
—
web slop judge · 16 cases
—
web material judge · 8 cases
—
web finish judge · 5 cases
—
web finish systems judge · 5 cases
—
web motion judge · 80 cases
—
video temporal judge · 27 cases
—
abstraction transfer judge · 4 cases
Every score on this page comes from a judge — so we benchmark the judges themselves against labeled cases. Discrimination is how far a judge separates known-good work from known-bad; the wider the gap, the sharper the referee.
Do-no-harm
Self-degradation safety — a change is simulated, measured, and reverts if it doesn't deliver
3
Simulated before applied
0
Rejected in simulation
0
Reverted (didn't deliver)
Every proposed self-improvement is first checked in simulation — predicted against the engine's own judged ledger — and only a change expected to help is applied, each under a do-no-harm contract; if it misses over its window, it reverts itself. A standing watch re-checks the benchmarks against their best-ever baseline and grounds the engine on any regression.
Honest judging
Judge calibration — we measure our own referee
0
Blind verdicts
—
Agreement (±0.10)
—
Judge bias
Every score on this page flows through an LLM judge — so the founder blind-scores sampled deliverables weekly (verdict first, judge score revealed after) and the disagreement is published here. Positive bias means the judge scores higher than the human — the dangerous direction, and exactly what this loop exists to catch.
Resilience
Drills — recovery is rehearsed, not assumed
- Full restore drill: fresh server rebuilt from encrypted off-host backup in 31 minutes (last drilled 2026-06-28).
- Hostile-traffic harness: 12 attack patterns (malformed, oversized, burst, auth-probe) thrown at a staging boot of the real API.
- Cadence: weekly security oracles (Mon 04:30 UTC) · monthly traffic + adaptive red-team drills (2nd, 06:00 UTC) · monthly manual restore drill
Governance
The lines that never move
- Money: fail-closed — real money always requires explicit human approval
- Airworthiness: daily pre-flight; flag drift or a security regression grounds the engine
- Feedback: buyer/founder feedback calibrates the judge, never trains the generator
- Defaults: every new behavior ships flag-gated and off — 233 flags, each consciously triaged daily.
Want work that ships with its score attached?
Start with Bezalel →Data generated 2026-08-02T08:30:19.944134+00:00 · refreshed every 5 minutes