Attack & Defense
From adaptive, locally benign attacks, to correctly timed intervention, to agentic safety that reasons over state and harm-enabling actions.
SEAD
A State-Based Perspective on Attack and Defense in Tool-Using Agents
Detect when accumulated agent actions make a later step harmful.
TurnGate
Learning When to Intervene Against Multi-Turn Malicious Intent
Learn when to intervene before multi-turn intent becomes actionable harm.
CKA-Agent
Bypassing Commercial LLM Guardrails via Harmless Prompt Weaving and Adaptive Tree Search
Adaptively weave harmless-looking prompts into a harmful objective.