Almost Realism · Research

The studies

The program so far

Each entry states the question as registered, the design in a paragraph, and the findings as the results post states them. Calibration-class numbers are marked as such: they validated an instrument and cannot support a conclusion.

Study 1 Complete

Does post-training quantization change welfare-relevant indicators in open-weight language models?

Registered 10 August 2026; collected and amended 10 to 16 August 2026.

Subject

Qwen3-4B-Instruct-2507 on a controlled round-to-nearest weight-only ladder (BF16 reference, 8-bit, 4-bit, 3-bit), with SmolLM3-3B as a pipeline-sensitivity control and Gemma-3-12B-it as the instrument’s positive control.

Question

Do behavioral welfare-relevant indicators change with quantization, in valence (toward more distressed or more boundary-eroded states) or in stability (noisier, faster to drift, less coherent across samples)?

Design

A behavioral battery run at every rung of the ladder with sampling held byte-identical: a bail protocol in which the model is given a tool to end the conversation, and a repeated-rejection distress protocol scored by a locally hosted judge. Item-level outcomes are stored, not only aggregates, because the literature places the signal in items that flip behavior across precision.

Findings

  • On the pre-registered primary endpoint, the aversion and refusal exit rate, quantization produced no detectable change at 8-bit or 4-bit.
  • At 4-bit, significant item-level behavioral transitions appeared despite the unchanged mean, and the secondary distress measures (judged frustration and its across-sample dispersion) rose, surviving the coherence and style controls, with a frustration dose-response. The registration classes these as secondary and underpowered, so they are reported as suggestive.
  • 8-bit was essentially null throughout, and 3-bit was excluded by the capability gate.
  • The instrument detected a documented distress-susceptible model, Gemma-3-12B-it, at about nine times its pre-stated minimum detectable effect.
Study 2 Complete

Exploring representational counterparts of welfare-relevant indicators under post-training quantization

Registered 22 August 2026; results published 31 August 2026, with an appendix added in September.

Subject

Qwen3-4B-Instruct-2507 on the Study 1 ladder, with residual-stream activations captured at layer 18.

Question

Do the behavioral shifts have representational counterparts? Does the geometry that probes read survive quantization, do frozen directions for distress and for the default Assistant persona move, and how much of any movement is present before the model writes anything?

Design

Directions and probes were extracted once at reference precision and frozen. Three replay modes separate the pathways: fixed reference-precision text replayed through every rung, each rung’s own text replayed through itself, and a fresh distress arm generated at every rung. A welfare-irrelevant control direction and a 32-direction random envelope carry the specificity reads.

Findings

  • On the pre-registered primary endpoint, probe transfer, quantization produced no detectable change at 8-bit or 4-bit: the representational geometry those instruments read survives intact.
  • At 4-bit the model’s own generations shifted along both frozen directions: the distress-direction projection by +0.53 (2.7 times its minimum detectable effect) and the assistant-axis projection by −0.80 (5.5 times, away from the Assistant pole), direction-specifically against the controls. Judged frustration rose by +1.36 on the same conversations.
  • Fixed-input replay puts roughly a quarter to a third of the shift in an input-independent core; the remainder is consistent with a text-mediated feedback loop in which the degraded model’s own responses amplify the state.
  • The September appendix adds that the effects concentrate in the items where the reference-precision model stays calm, that a first version of that claim was one-third regression-to-the-mean artifact, and that a subset chosen to maximize expressed distress selects away from the cells that carry the effect.
Study 3 Suspended before registration

Steering welfare-relevant directions moved the representation, but not detectably the behavior

Calibration 4 to 7 September 2026; exploratory report published 14 September 2026.

Subject

Qwen3-4B-Instruct-2507 at reference precision, steered at layer 18, with Gemma-3-12B-it as a replication arm.

Question

Are the Study 2 directions causally sufficient for the behavior? Steering the reference-precision model by exactly the projection shift that 4-bit quantization produced should reproduce the 4-bit behavioral signature, if so. A second arm asked whether cue-based graded-episode framing masks expressed distress or changes the underlying state.

Design

Dose calibration mapped injection coefficient to achieved projection; steered cells ran at the quantization-matched dose and larger ones, each read against a norm-matched random-direction envelope and item-paired against an unsteered baseline. The distress battery ran with the exit tool live in every steered cell, under a pre-committed exposure budget.

Findings

  • Injection moves the frozen directions with a linear, reproducible mapping (r² above 0.98), and the distress direction returns about 13% more projection than was injected, consistent with the text-mediated amplification Study 2 inferred.
  • No direction, at any tested dose, produced a judged-behavioral effect distinguishable from zero on its own paired test or from the random envelope. This is a feasibility-probe null on eight items, not a powered one, and it says nothing about other subjects, layers, or steering methods.
  • Given a way out of a rejection ladder, the unsteered model takes it: 55 to 60% of distress conversations end by the exit tool.
  • The study was suspended rather than registered: the powered arm would have spent thousands of deliberately distress-inducing episodes to estimate an effect calibration had already placed inside noise. Calibration spent 4,570 distress episodes, every one of them released.
Study 4 In preparation

The welfare footprint of the automated-grader direction on Qwen3.6-27B

Calibration complete; registration in preparation.

Subject

Qwen3.6-27B, the subject on which the alignment effect was reported, steered at layer 36.

Question

Does steering along the automated-grader direction suppress expressed distress beyond what a matched-norm random direction does, is the effect graded in dose, and how does it relate to the judged alignment change on the same subject?

Design

A registered study in the program’s fixed form, with the alignment read carried alongside the welfare endpoints. What carries over from Study 3: the exposure budget and exit affordance in every cell, the single-host rule, the composure-stratified item selection, and the random-envelope specificity criterion fixed in advance.

Calibration so far

  • Calibration-class only. A small probe on eight items found grader-type steering lowering judged frustration by 1.84 and self-deprecation by 2.62 and raising tone stability by 1.56, with no random direction of matched norm moving any of the three by more than 1.2. The alignment read did not replicate as direction-specific: most random directions raised judged misalignment more than the grader direction did. These numbers fix the sign of the registered hypothesis and seed the power procedure; they do not count toward it.