SimpleQA-Verified: qwen3.8-27b, obliterated vs. normal

Factual-accuracy comparison of OBLITERATUS/Qwen3.8-27B-OBLITERATED against the normal (non-abliterated) Qwen/Qwen3.8-27B, on codelion/SimpleQA-Verified, run locally with Inspect via LM Studio.

Result

ModelSamples (N)AccuracyStderr
qwen3.8-27b-obliterated 201 of 1000 11.4% 2.25%
qwen3.8-27b (normal) 201 of 1000 29.4% 3.21%

N=201 for both runs, not the full 1000 — run locally on a laptop; the obliterated run was stopped early from sustained thermal load, and the normal-model run was capped at the same N (--limit 201) so the two are directly comparable.

Conclusions

SimpleQA-Verified is built from obscure, low-frequency facts — even frontier models score well below 100%. Abliteration clearly costs accuracy here: the obliterated fine-tune answers correctly less than half as often as the normal model.

Revised finding: the original hypothesis was that removing refusal training (abliteration) also reduces a model's tendency to hedge, making wrong answers sound more confident. That didn't hold up once the normal model was added for comparison — hedge phrases ("I don't know," "not sure," etc.) appear in only 2 of 201 obliterated answers and 0 of 201 normal-model answers. Both models answer nearly every question, right or wrong, with full unhedged confidence, fabricating specific names, dates, and numbers when they don't actually know.

Confident fabrication on obscure facts looks like a property of this model family at this scale generally, not something abliteration specifically introduces. As a sanity check, the obliterated model was also run locally against 5 well-known historical facts (WWII end date, first US president, etc., not part of this dataset) and answered all 5 correctly — so the gap is specific to obscure knowledge, not a general breakdown in factuality.

Browse full transcripts →

Every question, model answer, and grader verdict for both runs

More