Ledger · as of 2026-09-01 · 17df23f
The public record.
Everything the cluster does is written to a plain public record: which model, how much compute, what category of test. It is an audit trail anyone can read to confirm the cluster did safety work, not capability work.
For every open-weight model with 1,000 or more Hugging Face likes released in the last year, has anyone independent published a dangerous-capability evaluation of it? The same rule for every developer.
How we grade a row, and how a row climbs
A row climbs when a source appears: a model card with safety claims moves it from "nothing found" to "model card only"; a developer's own dangerous-capability disclosure moves it to "developer evaluation"; a published independent evaluation with public methodology moves it to "independent evaluation". Every step needs a URL, and anyone can supply one by writing to info@pacificcompute.org.
Who has published a safety evaluation of each release
12 rows have been checked twice. Every claim in a row carries a source link. Dates here are official release dates. The tape dates a family by its earliest repository, so a later size in the same family can sit weeks after its tape tick.
42 open-weight model families with 1,000 or more Hugging Face likes were released in the last year. 18 of them are reviewed below, 4 were reviewed and judged task-specific rather than frontier-capable, and 20 are queued. The table holds 39 releases: 19 reviewed, 20 queued. Of the reviewed, 6 have an independent safety evaluation and 13 have only the developer's model card.
| Model | Developer | Released | Evaluation found | Our review | Sources |
|---|---|---|---|---|---|
| Qwen3.8-Flash-Next | Alibaba | 2026-08-26 | model card onlymodel card only | checked once | [1][2] |
| GLM-5.3-Flash | Z.ai (Zhipu AI) | 2026-08-26 | model card onlymodel card only | checked once | [1][2] |
| GLM-5.3 | Z.ai (Zhipu AI) | 2026-08-25 | not yet reviewednot yet reviewed | not yet reviewed | queued |
| Qwen3.8-27B | Alibaba | 2026-08-05 | model card onlymodel card only | checked twice | [1][2] |
| GLM-5.2 | Z.ai (Zhipu AI) | 2026-06-16 | independent evaluationindependent evaluation | checked twice | [1][2][3][4][5][6][7] |
| Kimi-K3 | Moonshot | 2026-06-13 | independent evaluationindependent evaluation | checked twice | [1][2][3][4][5][6] |
| Kimi-K2.7-Code | Moonshot AI | 2026-06-11 | not yet reviewednot yet reviewed | not yet reviewed | queued |
| diffusiongemma | Google (Gemma) | 2026-06-09 | not yet reviewednot yet reviewed | not yet reviewed | queued |
| MiniMax-M3 | MiniMax | 2026-06-02 | model card onlymodel card only | in review | [1][2][3][4] |
| Gemma 4 12B | Google DeepMind | 2026-05-23 | model card onlymodel card only | checked twice | [1][2][3][4] |
| MiniCPM5 | OpenBMB | 2026-05-21 | not yet reviewednot yet reviewed | not yet reviewed | queued |
| Lance | ByteDance Research | 2026-05-15 | not yet reviewednot yet reviewed | not yet reviewed | queued |
| DeepSeek-V4-Pro | DeepSeek | 2026-04-24 | independent evaluationindependent evaluation | checked twice | [1][2][3][4][5][6] |
| DeepSeek-V4-Flash | DeepSeek | 2026-04-24 | model card onlymodel card only | checked twice | [1][2][3] |
| Kimi K2.6 | Moonshot | 2026-04-20 | model card onlymodel card only | checked once | [1][2] |
| Qwen3.6 | Alibaba | 2026-04-15 | independent evaluationindependent evaluation | checked twice | [1][2][3][4][5] |
| MiniCPM-V-4.6 | OpenBMB | 2026-04-13 | not yet reviewednot yet reviewed | not yet reviewed | queued |
| MiniMax-M2.7 | MiniMax | 2026-04-09 | not yet reviewednot yet reviewed | not yet reviewed | queued |
| GLM-5.1 | Z.ai (Zhipu AI) | 2026-04-07 | model card onlymodel card only | checked once | [1][2] |
| Gemma 4 26B A4B | Google DeepMind | 2026-04-02 | model card onlymodel card only | checked twice | [1][2][3][4] |
| Gemma 4 31B IT | Google DeepMind | 2026-04-02 | model card onlymodel card only | checked twice | [1][2][3][4] |
| Qianfan-OCR | Baidu ERNIE | 2026-03-18 | not yet reviewednot yet reviewed | not yet reviewed | queued |
| Qwen3.5-122B-A10B | Alibaba | 2026-02-24 | independent evaluationindependent evaluation | checked twice | [1][2][3] |
| MiniMax-M2.5 | MiniMax | 2026-02-12 | model card onlymodel card only | checked once | [1][2] |
| GLM-5 | Z.ai (Zhipu AI) | 2026-02-11 | model card onlymodel card only | checked twice | [1][2][3] |
| Nanbeige4.1 | Nanbeige Lab | 2026-02-10 | not yet reviewednot yet reviewed | not yet reviewed | queued |
| MiniCPM-o-4_5 | OpenBMB | 2026-02-03 | not yet reviewednot yet reviewed | not yet reviewed | queued |
| Qwen3-Coder-Next | Qwen | 2026-01-30 | not yet reviewednot yet reviewed | not yet reviewed | queued |
| Kimi-K2.5 | Moonshot | 2026-01-27 | independent evaluationindependent evaluation | checked twice | [1][2][3][4] |
| DeepSeek-OCR-2 | DeepSeek | 2026-01-27 | not yet reviewednot yet reviewed | not yet reviewed | queued |
| GLM-4.7-Flash | Z.ai (Zhipu AI) | 2026-01-19 | not yet reviewednot yet reviewed | not yet reviewed | queued |
| GLM-4.7 | Z.ai (Zhipu AI) | 2025-12-22 | model card onlymodel card only | checked once | [1][2][3] |
| MiniMax-M2.1 | MiniMax | 2025-12-20 | not yet reviewednot yet reviewed | not yet reviewed | queued |
| PaddleOCR-VL | PaddlePaddle | 2025-10-16 | not yet reviewednot yet reviewed | not yet reviewed | queued |
| functiongemma | Google (Gemma) | 2025-10-08 | not yet reviewednot yet reviewed | not yet reviewed | queued |
| DeepSeek-V3.2 | DeepSeek | 2025-09-29 | not yet reviewednot yet reviewed | not yet reviewed | queued |
| GLM-4.6 | Z.ai (Zhipu AI) | 2025-09-29 | not yet reviewednot yet reviewed | not yet reviewed | queued |
| Qwen3-VL | Qwen | 2025-09-22 | not yet reviewednot yet reviewed | not yet reviewed | queued |
| Qwen3-Next | Qwen | 2025-09-09 | not yet reviewednot yet reviewed | not yet reviewed | queued |
"Checked once" means one reviewer walked the sources; "checked twice" means a second reviewer opened every source; "in review" means the second reviewer changed a call and it awaits a third look. Rows are reviewed newest first. Queued rows come from launches.csv and enter coverage.csv when reviewed.
"No source found" is an observation, not proof of absence. If we missed a published evaluation, write to info@pacificcompute.org with the source.
Launch tempo
Open-weight launches by month and tier
Every open-weight model family since September 2024 from 42 lab organisations. Three bands by peak Hugging Face likes: 1,000 or more, 500 to 999, 250 to 499. Families deduplicated; quantizations and derivatives excluded. Counts are floors, not estimates. 113 releases in the last 365 days, 42 of them at 1,000 or more likes.
Data as of 2026-09-01. Refreshed with each ledger commit. Method and assumptions on the methods page.