How a campaign runs

How a campaign runs.

We use this compute to run evaluations on new model releases, and to work with the labs to troubleshoot safety issues as they arise. Should an issue arise during an evaluation, we make resources available to the lab to close the problem and solve it for good.

  1. 1 · Evaluate

    The launch evaluation runs

    Before release, against an endpoint the lab hosts; after release, on the open weights. Ten to fifteen benchmarks across two or three checkpoints, branching into targeted red teaming when failure points are detected.

  2. Passes

    Every job on the ledger

    The job is on the public ledger: which model, what category of test, how much compute. We do not publish evaluation results; we provide the foundation for people to run them.

    Fails

    Reproduce and diagnose

    We help reproduce and diagnose the finding. The lab chooses what happens next.

  3. 2 · The lab chooses

    Retest

    Repeat the affected evaluations, on the same setup, after the lab has made a change on its own systems.

    Safety fine-tuning before launch

    Where a lab wants help fixing what we found. It runs as a separately approved job and lands on the ledger where the other participants can see it.

    Diagnose deeper

    Interpretability and diagnosis on the open weights. Research projects run at or below cost.

    Ship as is

    The lab's choice. Every job lands on the ledger.

  4. 3 · Retest, record

    Retest for free after a lab responds

    As a partner we reproduce and diagnose findings and retest for free after a lab responds. Every job lands on the ledger.

Write to us with the model and what you would run: info@pacificcompute.org, or use the request form.