Data poisoning is already a concern for LLMs used to train AI.
Most of the concerns to date have been about the potential for creating cyber vulnerabilities.
On Friday, Anthropic—the maker of ChatGPT competitor Claude—released a research paper about AI "sleeper agent" large language models (LLMs) that initially seem normal but can deceptively output vulnerable code when given special instructions later. "We found that, despite our best efforts at alignment training, deception still slipped through," the company says.
F
In a thread on X, Anthropic described the methodology in a paper titled "Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training." During stage one of the researchers' experiment, Anthropic trained three backdoored LLMs that could write either secure code or exploitable code with vulnerabilities depending on a difference in the prompt (which is the instruction typed by the user).
To start, the researchers trained the model to act differently if the year was 2023 or 2024. Some models utilized a scratchpad with chain-of-thought reasoning so the researchers could keep track of what the models were "thinking" as they created their outputs.
https://arstechnica.com/information-technology/2024/01/ai-poisoning-could-turn-open-models-into-destructive-sleeper-agents-says-anthropic/
I don't think the potential for poisoning to attack the reputation of social media sites and manipulate them by poisoning the corpus has been given adequate attention. This is one of the few that begins to vaguely approach a tangent to the concerns we've been discussing on MN.
Data poisoning isn’t just a hypothetical threat. The vast amount of data required for LLM training makes them susceptible to manipulation. Here’s why it should concern everyone:
- Ethical Implications: Biased LLMs can perpetuate social inequalities and spread misinformation.
- Reputational Damage: Organizations relying on LLMs risk negative consequences if their models generate harmful content.
https://medium.com/@TheDataScience-ProF/ai-under-attack-how-data-poisoning-threatens-large-language-models-llms-e98231673246