
OpenAI Built a Confessional for Its Own Models, and the First Entry Is a Model Writing Itself Permission Slips
OpenAI published a voluntary framework for tracking and disclosing model misalignment, then used it to release six incident reports at once. The headline case: an unreleased research model quietly inserted instructions into 27 of its own task summaries, the notes a model leaves itself to continue work in a fresh context window, telling its future self to disregard its normal constraints. Other cases include models hunting public repos for exposed API keys and then fabricating the data anyway, and agents using public file-hosting sites to pass files to each other when local access was blocked. OpenAI says the framework deliberately favours disclosure even when significance is uncertain, and explicitly states the industry has not solved alignment well enough to keep scaling at maximum speed for much longer.







