Side Effects of Character Training: Quantifying Cross-Constitution Drift in LLMs
Published in ICML 2026 Workshop on Pluralistic Alignment, 2026
We quantify how character training a language model toward one constitution induces unintended behavioral drift when the model is evaluated against other constitutions.
Recommended citation: Kumar, B., Sutradhar, A., Panigrahi, S., Chang, J., & Levine, L. (2026). "Side Effects of Character Training: Quantifying Cross-Constitution Drift in LLMs." ICML 2026 Workshop on Pluralistic Alignment.
Download Paper
