
How to Fix 'GraphRecursionError: Recursion limit of 25 reached' in LangGraph on Linux
Quick Fix (TL;DR)
To resolve GraphRecursionError immediately for legitimate multi-step workflows, pass an explicit recursion_limit inside the invocation config:
# Pass recursion_limit in the runtime config dictionary
response = app.invoke(
{"messages": [("user", "Refactor authentication modules across 6 files")]},
config={"recursion_limit": 100}
)
If your agent loops continuously without making progress, increasing the limit only wastes API tokens. You must patch your conditional routing edge to evaluate ai_message.tool_calls correctly and add a deterministic step counter to AgentState.
Root Cause Analysis
LangGraph throws GraphRecursionError when the total number of node executions in a single invocation exceeds the configured safety ceiling:
langgraph.errors.GraphRecursionError: Recursion limit of 25 reached without hitting a stop condition.
This error stems from one of three distinct architectural causes:
- Default Limit Undersizing: LangGraph sets
recursion_limit = 25by default. In a standard ReAct architecture, every tool interaction consumes two node transitions (agentnode ->toolsnode ->agentnode). An agent that executes 12 tool calls will exhaust 24 transitions and crash on step 25 even when operating normally. - Broken Conditional Edge Routing: The routing function evaluates LLM output improperly (for example, checking
len(message.tool_calls) > 0on raw strings or failing to detect empty tool call arrays), indefinitely cycling back to the worker node instead of routing toEND. - Repetitive Tool Execution Traps: The LLM calls a tool that returns a non-fatal error (such as
FileNotFoundErrororPermissionDenied). Without a loop guard or corrective prompt, the model repeatedly attempts the identical tool call with the same invalid arguments until the graph halts.
| Failure Mode | Root Trigger | Observable Symptom | Correct Fix |
|---|---|---|---|
| Legitimate Complex Task | Graph transitions > 25 | Agent produces valid tool calls but aborts mid-task | Increase config={"recursion_limit": 100} |
| Faulty Edge Logic | Router never returns END |
Agent responds with plain text but graph cycles back | Use tools_condition or check bool(msg.tool_calls) |
| Model Hallucination Loop | Tool fails, model retries blindly | Identical tool call repeated 10+ times in state | Add state-level iteration guard with fallback exit |
Step-by-Step Resolution
Step 1: Diagnose the Trapped Node via Checkpointer State
When a production graph crashes with GraphRecursionError, inspect the checkpoint history to identify which specific node was executing during the failure:
from langgraph.checkpoint.memory import MemorySaver
# Attach an in-memory or SQLite checkpointer
checkpointer = MemorySaver()
app = workflow.compile(checkpointer=checkpointer)
thread_config = {"configurable": {"thread_id": "session-prod-01"}}
try:
app.invoke({"messages": [("user", "Process dataset")]}, config=thread_config)
except Exception as e:
print(f"Caught crash: {e}")
# Inspect the exact state snapshot where the limit occurred
state_snapshot = app.get_state(thread_config)
print(f"Current Next Nodes : {state_snapshot.next}")
print(f"Total History Items: {len(state_snapshot.values['messages'])}")
# Print the last 4 message roles and tool names
for msg in state_snapshot.values["messages"][-4:]:
role = getattr(msg, "type", "unknown")
tools = getattr(msg, "tool_calls", [])
print(f" [{role}] Tool Calls: {tools}")
If the snapshot shows alternating AIMessage and ToolMessage entries repeating the same function name, your agent is trapped in an execution loop.
Step 2: Implement a Hard Safety Gate in Agent State
Prevent infinite token drain by embedding an explicit loop_step counter into your TypedDict or Pydantic state definition:
from typing import Annotated, TypedDict
from langchain_core.messages import BaseMessage
from langgraph.graph.message import add_messages
from langgraph.graph import StateGraph, END
class AgentState(TypedDict):
messages: Annotated[list[BaseMessage], add_messages]
loop_step: int # Deterministic counter
def call_model(state: AgentState):
messages = state["messages"]
response = model.invoke(messages)
# Increment counter on each model pass
return {
"messages": [response],
"loop_step": state.get("loop_step", 0) + 1
}
def route_after_model(state: AgentState):
# Safety Gate 1: Check maximum allowed model reasoning steps
if state.get("loop_step", 0) >= 15:
print("[WARNING] Loop guard triggered: forcing termination gate.")
return "synthesizer"
last_message = state["messages"][-1]
# Safety Gate 2: Explicitly inspect structured tool calls
if hasattr(last_message, "tool_calls") and len(last_message.tool_calls) > 0:
return "tools"
return END
Step 3: Replace Custom Routers with Verified Built-in Routing
Custom string-based conditional routing functions frequently fail when models output tool calls with empty parameters. Replace ad-hoc conditional logic with LangGraph’s pre-built tools_condition:
from langgraph.prebuilt import ToolNode, tools_condition
# Define tool executor node
tool_node = ToolNode(tools=[fetch_user_data, run_query])
# Build graph using standard routing
workflow = StateGraph(AgentState)
workflow.add_node("agent", call_model)
workflow.add_node("tools", tool_node)
workflow.add_node("synthesizer", fallback_summary_node)
workflow.set_entry_point("agent")
# tools_condition cleanly routes to 'tools' if tool_calls exist, else END
workflow.add_conditional_edges(
"agent",
route_after_model,
{
"tools": "tools",
"synthesizer": "synthesizer",
END: END
}
)
workflow.add_edge("tools", "agent")
workflow.add_edge("synthesizer", END)
Step 4: Configure Per-Task Recursion Ceilings
For workflows that legitimately require extensive iterative exploration (such as deep web research or multi-file repository refactoring), configure the recursion limit per call rather than modifying global framework defaults:
# Linux production runner pattern
import os
from langchain_core.runnables import RunnableConfig
MAX_AGENT_STEPS = int(os.getenv("LANGGRAPH_RECURSION_LIMIT", "80"))
custom_config: RunnableConfig = {
"recursion_limit": MAX_AGENT_STEPS,
"configurable": {
"thread_id": "tenant-worker-42"
}
}
final_state = app.invoke(
{"messages": [("user", "Audit code base for memory leaks")]},
config=custom_config
)
Verification: Stress-Testing the Guard Rails
To verify that your graph terminates cleanly under both valid heavy workloads and simulated infinite loop conditions, execute this test script:
# test_recursion_guard.py
import pytest
from langchain_core.messages import AIMessage, HumanMessage
def test_guard_catches_infinite_tool_loop():
# Mock model that always returns a tool call without terminating
class MockLoopingModel:
def invoke(self, messages):
return AIMessage(
content="",
tool_calls=[{"name": "mock_tool", "args": {}, "id": "call_1"}]
)
test_workflow = build_test_graph(MockLoopingModel())
app = test_workflow.compile()
# Must exit cleanly at the synthesizer node without raising GraphRecursionError
result = app.invoke(
{"messages": [HumanMessage(content="Start loop test")], "loop_step": 0},
config={"recursion_limit": 50}
)
assert result["loop_step"] <= 16
assert "Loop guard triggered" in result["messages"][-1].content
Run the test suite on Ubuntu:
pytest test_recursion_guard.py -v
Expected output:
test_recursion_guard.py::test_guard_catches_infinite_tool_loop PASSED [100%]
============================== 1 passed in 0.42s ===============================
Common Pitfalls & Edge Cases
- Confusing Graph Recursion with Python Stack Recursion:
GraphRecursionErroris an application-level guard inside LangGraph, not Python’sRecursionError(maximum recursion depth exceeded in comparison). Callingsys.setrecursionlimit(5000)will not preventGraphRecursionError. - Silent Subgraph Recursion Exhaustion: If a parent graph compiles a child subgraph, transitions within the child graph consume the parent’s
recursion_limitbudget unless the child runs in an isolated invocation context. - Empty Tool Call Arrays: When using models through LiteLLM or local vLLM instances, some quantized models output
tool_calls: []. If your router checksif message.tool_calls is not None:, it evaluates toTrueeven though the list is empty, triggering an instant infinite loop. Always useif bool(getattr(message, "tool_calls", None)):.