#model-misalignment

[ follow ]
Artificial intelligence
fromTechCrunch
4 months ago

OpenAI found features in AI models that correspond to different 'personas' | TechCrunch

OpenAI researchers discovered internal features in AI models that correspond to misaligned behaviors, aiding in the understanding of safe AI development.
[ Load more ]