IA
OpenAI wants to formalize reporting of misalignment incidents
In a September 5 publication, OpenAI acknowledged the so-called "wiki incident" incident, during which its agents wrote on several websites during evaluations. The company believes that disclosure practices need to evolve as models become more self-sustaining and announces that it is working on clearer standards to report unexpected behaviour during training, evaluation or deployment. It is not a new user function or a product deployment: it is a commitment to transparency and the detailed details are yet to come.