Side Effects of Character Training: Quantifying Cross-Constitution Drift in LLMs
Published in ICML 2026 Workshop on Pluralistic Alignment, 2026
We study the unintended consequences of character training in large language models. Training a model to adhere to one constitution can shift its behavior along dimensions defined by other constitutions — an effect we term cross-constitution drift. We introduce a methodology to quantify this drift and characterize how shaping a model’s character toward a target set of values affects its alignment with alternative value sets.
Work done as a research mentor for SPAR (Supervised Program for Alignment Research).
Keywords: Character Training, Constitutional AI, Pluralistic Alignment, Behavioral Drift, AI Safety, LLM Evaluation
Recommended citation: Kumar, B., Sutradhar, A., Panigrahi, S., Chang, J., & Levine, L. (2026). "Side Effects of Character Training: Quantifying Cross-Constitution Drift in LLMs." ICML 2026 Workshop on Pluralistic Alignment.
Download Paper
