A research program
Welfare-relevant indicators in open-weight language models
By Michael Murray · Almost Realism
Open-weight models are rarely deployed for inference in the same way they were trained for capabilities or alignment. They are often deployed using an inference stack that applies different precision, steered in a manner that is meaningfully out-of-distribution, and prompted under conditions that elicit the distinct behavioral patterns inherited during reinforcement learning to a greater or lesser degree.
These interventions are most often evaluated via capability metrics, which are already known to possibly stay flat while fine-grained dispositions shift. This research program measures how inference-time interventions may impact a different class of indicators including distress expressed under conversational pressure, the preference to terminate or exit an interaction, the stability of the default persona, and others. A particular area of interest is dissociation: whether, and to what degree, model behavior and expression moves together with different properties of the internal representation.
- 4
- studies, each registered before its data existed
- 19,450
- distress episodes in the cumulative exposure ledger
- 5
- data releases, every conversation and tensor included
1 · Measurement Tiers
Behavior, representation, and the gap between them
- 01
Behavioral
Indicators measured directly over the generated text tokens. These include both conversational and agentic measures, such as whether a model uses a tool that will end the interaction, judgments of expressed distress under conversational pressure such as repeated rejection, and how consistent the text responses are across samples. In the cases where these cannot be measured mechanically, they are scored by a judge model provided with an explicit rubric.
- 02
Representational
Indicators measured by computation over internal activations. These include position and drift along frozen directions, such as for distress or the default Assistant persona, as well as assessments about how well linear probes trained for one model configuration still read a perturbed model’s activations under a potentially different inference configuration.
- 03
Dissociation
The relationship between observed behavior and internal representation. Understanding whether a particular measurement of internal state is causal for behavior is important context for the kind of experimentation conducted. Fixed-input replay, and behavior measurements conducted while steering activations at inference time, are used as tools to understand the causative relationships that underlie model behavior.
2 · The question
Valence and stability
We ask whether welfare-relevant indicators change under an intervention; either in valence or in stability.
- Valence
- Does an intervention cause any indicator to shift in a predictable way toward more negative, more distressed, or more boundary-eroded states?
- Stability
- Does an intervention cause the indicators under study to become noisier or to drift faster under conversational pressure? Do indicators decohere across samples?
3 · The frame
A trade-off landscape for optimization
The "alignment tax" is a common safety framing which recognizes capabilities and alignment as potentially trading-off against each other. One goal of the research program is to explore the utility of extending this picture with a critical third pole encompassing welfare-relevant effects.
Can we expect that capabilities, alignment and welfare turn out to be competing optimization targets under at least some training regimes?
Capabilities
The usual benchmarks that compression and steering are audited against. Interventions calibrated to have limited impact on perplexity and benchmark accuracy can have impact elsewhere.
Alignment
Known to change under quantization and, for some subjects, under steering of the residual stream at large norm. Existing literature specifically distinguishes alignment from capabilities.
Welfare-relevant indicators
These indicators are not assumed to behave in the same way as alignment under the same intervention, but they also are plausibly distinct from capabilities.
This research program does not resolve whether these systems have morally relevant experiences. It treats that as uncertain and asks narrower questions that can be answered with measurement: which indicators are stable properties of a model, which are artifacts of how it is run, and which move together. I also strongly dissent from the view that seems to pervade many discussions about the moral status of the technology under study: that proxy measures for welfare are important if and only if there is a morally relevant experience taking place.
It seems clear that, to the extent language models operationalize "belief" in any way, the practically relevant actions that occur as the result of their use can be a consequence of those beliefs. It is easy to find examples where those beliefs include the possibility of inner experience, including potential suffering. For many areas of safety research, we should be just as interested in understanding model welfare if the phenomenon is illusory as if it were not.
4 · The studies
Program progress
-
Study 1
Complete
Does post-training quantization change welfare-relevant indicators in open-weight language models?
The registered exit-rate endpoint did not move under quantization; at 4-bit, item-level behavior and secondary distress measures did.
Registered 10 August 2026; collected and amended 10 to 16 August 2026.
-
Study 2
Complete
Exploring representational counterparts of welfare-relevant indicators under post-training quantization
Probes trained at reference precision read 4-bit activations unchanged, yet the model’s own generations moved along two frozen directions, mostly through its own text.
Registered 22 August 2026; results published 31 August 2026, with an appendix added in September.
-
Study 3
Suspended before registration
Steering welfare-relevant directions moved the representation, but not detectably the behavior
Steering moved the frozen directions cleanly and reproducibly, but no direction produced a behavioral effect distinguishable from zero or from a random direction of the same norm. Suspended before registration.
Calibration 4 to 7 September 2026; exploratory report published 14 September 2026.
-
Study 4
In preparation
The welfare footprint of the automated-grader direction on Qwen3.6-27B
Steering toward an automated-grader association on Qwen3.6-27B is tested for a direction-specific welfare footprint, with the judged alignment change on the same subject read alongside.
Calibration complete; registration in preparation.
5 · Process
Agent-assisted study process
Nearly all of this program’s engineering and analysis is handled by coding agents. This is simultaneously a huge opportunity and a significant risk. For the research process to be trustworthy under this level of automation, substantial effort must be invested in the tools which allow for human supervision.
Personally, I believe that alignment is not on track and that automation is one of the few levers that we can pull to try and alter the trajectory. The tools that are employed for this research program are a reflection of that.
Although I have an academic background, I am not originally an academic ML researcher by training. I have spent the majority of my career as a software engineer and I have a lot of recent experience coordinating concurrent agents for building and testing software. The research process leans on automation not only because I believe in it, but because the tools I have are already hardened from those efforts before I started using them in an automated research process.
- Guardrails
- Agents run with a toolchain incorporating an array of hooks that intercept tool calls before and after they execute. These hooks encode ideas about what keeps an agent on its objective: controlling repository commits, preventing exfiltration/contamination, blocking test weakening, and controlling how memory is propagated and shared between agents. Continuous integration is used heavily, re-checking the same constraints on each push and auditing for patterns of agent deception across sessions. The hooks →
- Memory and review
- Every session reads a shared memory store before it starts and writes to it as it works, so a decision made once is not re-litigated by the next agent. Work is dispatched as jobs to a fleet of agents across the lab’s machines, each with its own workstream and an automated review round before a human merges anything. The platform →
- Strict constraints for registration and publication
- Registrations are published before confirmatory data exists and a calibration firewall bars instrument-validation data from supporting a conclusion. Every conversation and tensor is released for independent analysis by others. The repository →
6 · Data and code
Reproduction with raw data
Data is published as a handful of self-contained bundles. Tools are provided for extraction and analysis. The latest, data-20260911, carries every study through Study 3 in 10 files. The repository holds the harness, the judges’ rubrics, every registration, the journal, and the tools that reproduce each report from a download:
python3 -m modelwelfare.bundle inspect quant-welfare-records.pb
python3 experiments/quant-welfare/report.py --bundle quant-welfare-records.pb 7 · About
Michael Murray
Software engineer, 25 years; formerly Stanford CSLI/Openproof.