How a campaign runs
How a campaign runs.
We use this compute to run evaluations on new model releases, and to work with the labs to troubleshoot safety issues as they arise. Should an issue arise during an evaluation, we make resources available to the lab to close the problem and solve it for good.
- 1 · Evaluate
The launch evaluation runs
Before release, against an endpoint the lab hosts; after release, on the open weights. Ten to fifteen benchmarks across two or three checkpoints, branching into targeted red teaming when failure points are detected.
- Passes
Every job on the ledger
The job is on the public ledger: which model, what category of test, how much compute. We do not publish evaluation results; we provide the foundation for people to run them.
FailsReproduce and diagnose
We help reproduce and diagnose the finding. The lab chooses what happens next.
- 2 · The lab chooses
- 3 · Retest, record
Retest for free after a lab responds
As a partner we reproduce and diagnose findings and retest for free after a lab responds. Every job lands on the ledger.
Write to us with the model and what you would run: info@pacificcompute.org, or use the request form.