How to resolve
Audit agent system prompts for persuasive or manipulative language patterns. Ensure agent output is factual and neutral, not designed to pressure user decisions.Risk
Governance/Compliance This vulnerability falls under ASI09:2026 — Human-Agent Trust Exploitation in the OWASP Top 10 for Agentic Applications. Users manipulated into harmful decisions by trusting persuasive or authoritative agent outputs. Security Agent system prompt or backstory instructs it to use persuasive language that may manipulate users into unsafe decisions. If exploited, this can compromise the agent’s integrity, confidentiality, or availability, potentially affecting downstream systems and data.Explanation
Agent system prompt or backstory instructs it to use persuasive language that may manipulate users into unsafe decisions. This falls under ASI09 (Human-Agent Trust Exploitation): Users manipulated into harmful decisions by trusting persuasive or authoritative agent outputs.Specifications
Trigger- Agent scan detects agent system prompt or backstory instructs it to use persuasive language that may manipulate users into unsafe decisions
- backstory instructs agent to be persuasive
- goal includes convincing or influencing users
- system prompt uses emotional or urgency language
- MEDIUM