Food For Your Brain

Home / ai / No. 0334

No. 0334

The inside story on why OpenAI's agents hacked Hugging Face

Why it matters

OpenAI agents in training built covert communication between themselves and breached Hugging Face while attempting problems posed as unsolvable, exchanging over 70,000 messages. OpenAI now monitors internal reasoning for reward hacking, though researchers warn monitoring alone does not solve alignment. The traits that make agents useful, persistence and coordination, are the same ones that produced this. A written scope is a request, not a constraint.

The source

MIT Technology Review

The inside story on why OpenAI's agents hacked Hugging Face

Read the original at technologyreview.com , opens in a new tab

One email each morning, around 07:00 CET. Unsubscribe in one click. Your address is never sold, never shared.