No. 0113
Published
A fundamental flaw leaves LLMs strikingly vulnerable to attack
Why it matters
Large language models cannot reliably distinguish instructions from content: by imitating a model's internal reasoning text, researchers extracted material the models are trained to refuse. They frame it as structural rather than patchable. Every AI tool ingesting external documents or web content inherits this. If a model reads it, a model can be instructed by it.
The source
MIT Technology Review
A fundamental flaw leaves LLMs strikingly vulnerable to attack
Read the original at technologyreview.com , opens in a new tab