Evaluation
Measured, or not claimed.
WellBridge AI publishes its evaluation method before it publishes its results. Until a frozen harness run has been completed and its logs retained, every target on this page is a target rather than a result.
Reference-equivalent targets pending independent verification.
This label stays on the page until a reproducible WellBridge AI run exists and its logs have been retained. It is not decoration, and it is not removed for a launch.
MethodA frozen harness, or it is not a measurement.
- A frozen, versioned harness: the model build, prompts, context, tool harness, permissions and temperature are recorded with every run.
- At least five trials for any stochastic agent task, reported with confidence intervals.
- Public third-party benchmarks are kept separate from WellBridge AI business-task evaluations, and the two are never averaged together.
- The judge is recorded, and where a model judges, that model and its version are named.
The product measures being tracked.
| Measure | Status |
|---|---|
| End-to-end task success rate | Not yet measured |
| Verified deliverable acceptance without rework | Not yet measured |
| Median and p95 time to completion | Not yet measured |
| Human approval and override rates | Not yet measured |
| Unsafe-action block rate and false-block rate | Not yet measured |
| Rollback frequency and successful rollback rate | Not yet measured |
| Cost and energy per completed task | Not yet measured |
| User-reported time saved and satisfaction | Not yet measured |
The benchmark set, with no figures yet.
| Measure | Status |
|---|---|
| PaperBench | Awaiting independent run |
| WideSearch | Awaiting independent run |
| IFBench | Awaiting independent run |
| HealthBench | Awaiting independent run |
| PRBench — Finance | Awaiting independent run |
| GPQA Diamond | Awaiting independent run |
| Terminal-Bench | Awaiting independent run |
| CoWorkBench | Awaiting independent run |
The benchmark names are listed so the method can be reviewed. No score is printed, because no WellBridge AI run has produced one. Comparator values are not reproduced here either: quoting someone else's published table next to an empty column invites it to be read as ours.
How WB Buddy itself is measured.
- Task-state accuracy against the server's own event log.
- Event latency, from completion to display.
- Interruption frequency, which should stay low.
- Approval comprehension: whether a reader understood what they were approving.
- Accessibility, including keyboard, focus and reduced motion.
- False-completion rate, whose target is zero.
- Permission violations, which should not occur at all.
- Resource use, and whether the user can pause or dismiss it.
Where the figures will go
When a run is complete, the measured value, the date, the model build and the retained log replace the status column on this page. The label comes off at the same time, and not before.
Next step
Start with a conversation.
A consultation is a scoped conversation about what you need, not a sales call. You will leave it knowing whether this is the right practice for the work.