AI for Learning
Anthropic research shows how safety classifiers can be backdoored…
Anthropic researchers report that a small, roughly constant number of poisoned fine-tuning examples can install a backdoor in constitutional classifiers without obvious robustness losses.
Source checked
Anthropic Alignment Science
Source ↗Full Brief →