FAGEN Failure Modes in Agentic AI
Tutorial EMNLP 2026 · November 2026

Failure Modes in Agentic AI

Reproducible Triggers, Trace Diagnostics, and Verified Fixes

Abstract

Foundation-model agents now run in search, browsing, coding, and embodied settings, where quality hinges on long-horizon interaction: small errors compound across tool calls and memory, and single-turn evaluation misses them. This tutorial gives an operational view of agent failure modes, how to trigger them reproducibly, diagnose them from traces, and verify that fixes hold.

Tutorial Outline

Part 1 35 min

Agent Environments and Infrastructure

Failure patterns

  • Brittle tool use under interface shifts
  • Unstable memory over long episodes
  • Unfaithful chain-of-thought traces

Diagnostics

  • Trajectory schemas for actions, tools, and memory
  • Minimal closed-loop scaffolds as test stands
  • Process metrics beyond terminal success

Fixes

  • Tool and memory interface contracts
  • Standardized logging and trace tooling
  • Runnable baselines others can replay
Part 2 50 min

Training Dynamics

Failure patterns

  • Echo Trap: reasoning decays into templates
  • Entropy and mode collapse under KL
  • Reward hacking and overoptimization

Diagnostics

  • Reward variance, entropy, gradient norms
  • Mutual-information and SNR-aware filtering
  • Policy-gradient decomposition for instability

Fixes

  • Variance-aware trajectory filtering
  • KL choices that preserve diversity
  • Large-scale recipes like DAPO
Part 3 45 min

Evaluation and Trace Diagnostics

Failure patterns

  • Benchmark gains that vanish in deployment
  • Hallucination and knowledge overshadowing
  • Out-of-sync recovery failures

Diagnostics

  • Process-level, behavioral testing
  • Contrastive and counterfactual evaluation
  • Asynchronous, adversarial environments

Fixes

  • Standardized trace logging and reproductions
  • Conflict-aware reviewing and meta-eval hygiene
  • Comparable protocols across papers
Part 4 35 min

Test-Time Scaling and Adaptation

Failure patterns

  • Budget unawareness: overspent tokens and calls
  • Uncontrolled exploration in rollouts
  • Diversity collapse from format constraints

Diagnostics

  • Cost-aware trace analysis (BAGEN)
  • Uncertainty-guided exploration control (T2PO)
  • Diversity-aware optimization and format audits

Fixes

  • Context management with long-context models
  • Inference scaffolds with recovery and abstention
  • Evolving rubrics for long-form tasks

Schedule

Time Session Presenter
0:00 - 0:35
Part 1: Agent Environments and Infrastructure
Agent loop anatomy, canonical scaffolds, a taxonomy of failure surfaces, and what a reproducible trigger looks like.
Manling Li
0:35 - 1:25
Part 2: Training Dynamics
Reward design and hacking, optimization-induced collapse, and stabilization recipes. Safety: reward hacking under constraints.
Zoey Sha Li
1:25 - 1:55
Coffee break
1:55 - 2:40
Part 3: Evaluation and Trace Diagnostics
Why static benchmarks miss failures, closed-loop and asynchronous evaluation, and trace-level diagnostics. Safety: attack-surface evaluation.
Zihan Wang
2:40 - 3:15
Part 4: Test-Time Scaling and Adaptation
When more inference hurts, budget-aware agents, exploration control, and context management. Safety: guardrails under long interaction.
Hejie Cui
3:15 - 3:30
Synthesis and Q&A
Discussion across the four axes; release of reproductions and trace examples.
All presenters

Reading List

Part 1

Agent Environments and Infrastructure

Part 3

Evaluation and Trace Diagnostics

Part 4

Test-Time Scaling, Adaptation, and Safety

Presenters