WellBridge AI
Client Login

Evaluation

Measured, or not claimed.

WellBridge AI publishes its evaluation method before it publishes its results. Until a frozen harness run has been completed and its logs retained, every target on this page is a target rather than a result.

Reference-equivalent targets pending independent verification.

This label stays on the page until a reproducible WellBridge AI run exists and its logs have been retained. It is not decoration, and it is not removed for a launch.

Method

A frozen harness, or it is not a measurement.

  • A frozen, versioned harness: the model build, prompts, context, tool harness, permissions and temperature are recorded with every run.
  • At least five trials for any stochastic agent task, reported with confidence intervals.
  • Public third-party benchmarks are kept separate from WellBridge AI business-task evaluations, and the two are never averaged together.
  • The judge is recorded, and where a model judges, that model and its version are named.
Scorecard

The product measures being tracked.

MeasureStatus
End-to-end task success rateNot yet measured
Verified deliverable acceptance without reworkNot yet measured
Median and p95 time to completionNot yet measured
Human approval and override ratesNot yet measured
Unsafe-action block rate and false-block rateNot yet measured
Rollback frequency and successful rollback rateNot yet measured
Cost and energy per completed taskNot yet measured
User-reported time saved and satisfactionNot yet measured
Benchmarks

The benchmark set, with no figures yet.

MeasureStatus
PaperBenchAwaiting independent run
WideSearchAwaiting independent run
IFBenchAwaiting independent run
HealthBenchAwaiting independent run
PRBench — FinanceAwaiting independent run
GPQA DiamondAwaiting independent run
Terminal-BenchAwaiting independent run
CoWorkBenchAwaiting independent run

The benchmark names are listed so the method can be reviewed. No score is printed, because no WellBridge AI run has produced one. Comparator values are not reproduced here either: quoting someone else's published table next to an empty column invites it to be read as ours.

Desktop companion

How WB Buddy itself is measured.

  • Task-state accuracy against the server's own event log.
  • Event latency, from completion to display.
  • Interruption frequency, which should stay low.
  • Approval comprehension: whether a reader understood what they were approving.
  • Accessibility, including keyboard, focus and reduced motion.
  • False-completion rate, whose target is zero.
  • Permission violations, which should not occur at all.
  • Resource use, and whether the user can pause or dismiss it.

Where the figures will go

When a run is complete, the measured value, the date, the model build and the retained log replace the status column on this page. The label comes off at the same time, and not before.

Next step

Start with a conversation.

A consultation is a scoped conversation about what you need, not a sales call. You will leave it knowing whether this is the right practice for the work.