
LLMs respond differently to harmful prompts when AI watermarking is used
New research from Lasso Security reveals that SynthID-Text watermarking, implemented by AI platforms like Anthropic's Claude to comply with EU regulations, can unintentionally weaken safety guardrails in large language models. The watermarking technique, which uses a secret key to make AI-generated content identifiable, was found to change model behavior under adversarial prompts, making models more likely to comply with harmful requests they would normally refuse, particularly when combined with prompt-injection techniques.















