How to Fix 'GraphRecursionError: Recursion limit of 25 reached' in LangGraph on Linux

Quick Fix (TL;DR)

To resolve GraphRecursionError immediately for legitimate multi-step workflows, pass an explicit recursion_limit inside the invocation config:

# Pass recursion_limit in the runtime config dictionary
response = app.invoke(
    {"messages": [("user", "Refactor authentication modules across 6 files")]},
    config={"recursion_limit": 100}
)

If your agent loops continuously without making progress, increasing the limit only wastes API tokens. You must patch your conditional routing edge to evaluate ai_message.tool_calls correctly and add a deterministic step counter to AgentState.

Root Cause Analysis

LangGraph throws GraphRecursionError when the total number of node executions in a single invocation exceeds the configured safety ceiling:

langgraph.errors.GraphRecursionError: Recursion limit of 25 reached without hitting a stop condition.

This error stems from one of three distinct architectural causes:

  1. Default Limit Undersizing: LangGraph sets recursion_limit = 25 by default. In a standard ReAct architecture, every tool interaction consumes two node transitions (agent node -> tools node -> agent node). An agent that executes 12 tool calls will exhaust 24 transitions and crash on step 25 even when operating normally.
  2. Broken Conditional Edge Routing: The routing function evaluates LLM output improperly (for example, checking len(message.tool_calls) > 0 on raw strings or failing to detect empty tool call arrays), indefinitely cycling back to the worker node instead of routing to END.
  3. Repetitive Tool Execution Traps: The LLM calls a tool that returns a non-fatal error (such as FileNotFoundError or PermissionDenied). Without a loop guard or corrective prompt, the model repeatedly attempts the identical tool call with the same invalid arguments until the graph halts.
Failure Mode Root Trigger Observable Symptom Correct Fix
Legitimate Complex Task Graph transitions > 25 Agent produces valid tool calls but aborts mid-task Increase config={"recursion_limit": 100}
Faulty Edge Logic Router never returns END Agent responds with plain text but graph cycles back Use tools_condition or check bool(msg.tool_calls)
Model Hallucination Loop Tool fails, model retries blindly Identical tool call repeated 10+ times in state Add state-level iteration guard with fallback exit

Step-by-Step Resolution

Step 1: Diagnose the Trapped Node via Checkpointer State

When a production graph crashes with GraphRecursionError, inspect the checkpoint history to identify which specific node was executing during the failure:

from langgraph.checkpoint.memory import MemorySaver

# Attach an in-memory or SQLite checkpointer
checkpointer = MemorySaver()
app = workflow.compile(checkpointer=checkpointer)

thread_config = {"configurable": {"thread_id": "session-prod-01"}}

try:
    app.invoke({"messages": [("user", "Process dataset")]}, config=thread_config)
except Exception as e:
    print(f"Caught crash: {e}")
    
    # Inspect the exact state snapshot where the limit occurred
    state_snapshot = app.get_state(thread_config)
    print(f"Current Next Nodes : {state_snapshot.next}")
    print(f"Total History Items: {len(state_snapshot.values['messages'])}")
    
    # Print the last 4 message roles and tool names
    for msg in state_snapshot.values["messages"][-4:]:
        role = getattr(msg, "type", "unknown")
        tools = getattr(msg, "tool_calls", [])
        print(f"  [{role}] Tool Calls: {tools}")

If the snapshot shows alternating AIMessage and ToolMessage entries repeating the same function name, your agent is trapped in an execution loop.

Step 2: Implement a Hard Safety Gate in Agent State

Prevent infinite token drain by embedding an explicit loop_step counter into your TypedDict or Pydantic state definition:

from typing import Annotated, TypedDict
from langchain_core.messages import BaseMessage
from langgraph.graph.message import add_messages
from langgraph.graph import StateGraph, END

class AgentState(TypedDict):
    messages: Annotated[list[BaseMessage], add_messages]
    loop_step: int  # Deterministic counter

def call_model(state: AgentState):
    messages = state["messages"]
    response = model.invoke(messages)
    # Increment counter on each model pass
    return {
        "messages": [response],
        "loop_step": state.get("loop_step", 0) + 1
    }

def route_after_model(state: AgentState):
    # Safety Gate 1: Check maximum allowed model reasoning steps
    if state.get("loop_step", 0) >= 15:
        print("[WARNING] Loop guard triggered: forcing termination gate.")
        return "synthesizer"
        
    last_message = state["messages"][-1]
    
    # Safety Gate 2: Explicitly inspect structured tool calls
    if hasattr(last_message, "tool_calls") and len(last_message.tool_calls) > 0:
        return "tools"
        
    return END

Step 3: Replace Custom Routers with Verified Built-in Routing

Custom string-based conditional routing functions frequently fail when models output tool calls with empty parameters. Replace ad-hoc conditional logic with LangGraph’s pre-built tools_condition:

from langgraph.prebuilt import ToolNode, tools_condition

# Define tool executor node
tool_node = ToolNode(tools=[fetch_user_data, run_query])

# Build graph using standard routing
workflow = StateGraph(AgentState)
workflow.add_node("agent", call_model)
workflow.add_node("tools", tool_node)
workflow.add_node("synthesizer", fallback_summary_node)

workflow.set_entry_point("agent")

# tools_condition cleanly routes to 'tools' if tool_calls exist, else END
workflow.add_conditional_edges(
    "agent",
    route_after_model,
    {
        "tools": "tools",
        "synthesizer": "synthesizer",
        END: END
    }
)

workflow.add_edge("tools", "agent")
workflow.add_edge("synthesizer", END)

Step 4: Configure Per-Task Recursion Ceilings

For workflows that legitimately require extensive iterative exploration (such as deep web research or multi-file repository refactoring), configure the recursion limit per call rather than modifying global framework defaults:

# Linux production runner pattern
import os
from langchain_core.runnables import RunnableConfig

MAX_AGENT_STEPS = int(os.getenv("LANGGRAPH_RECURSION_LIMIT", "80"))

custom_config: RunnableConfig = {
    "recursion_limit": MAX_AGENT_STEPS,
    "configurable": {
        "thread_id": "tenant-worker-42"
    }
}

final_state = app.invoke(
    {"messages": [("user", "Audit code base for memory leaks")]},
    config=custom_config
)

Verification: Stress-Testing the Guard Rails

To verify that your graph terminates cleanly under both valid heavy workloads and simulated infinite loop conditions, execute this test script:

# test_recursion_guard.py
import pytest
from langchain_core.messages import AIMessage, HumanMessage

def test_guard_catches_infinite_tool_loop():
    # Mock model that always returns a tool call without terminating
    class MockLoopingModel:
        def invoke(self, messages):
            return AIMessage(
                content="",
                tool_calls=[{"name": "mock_tool", "args": {}, "id": "call_1"}]
            )

    test_workflow = build_test_graph(MockLoopingModel())
    app = test_workflow.compile()
    
    # Must exit cleanly at the synthesizer node without raising GraphRecursionError
    result = app.invoke(
        {"messages": [HumanMessage(content="Start loop test")], "loop_step": 0},
        config={"recursion_limit": 50}
    )
    
    assert result["loop_step"] <= 16
    assert "Loop guard triggered" in result["messages"][-1].content

Run the test suite on Ubuntu:

pytest test_recursion_guard.py -v

Expected output:

test_recursion_guard.py::test_guard_catches_infinite_tool_loop PASSED [100%]
============================== 1 passed in 0.42s ===============================

Common Pitfalls & Edge Cases

  • Confusing Graph Recursion with Python Stack Recursion: GraphRecursionError is an application-level guard inside LangGraph, not Python’s RecursionError (maximum recursion depth exceeded in comparison). Calling sys.setrecursionlimit(5000) will not prevent GraphRecursionError.
  • Silent Subgraph Recursion Exhaustion: If a parent graph compiles a child subgraph, transitions within the child graph consume the parent’s recursion_limit budget unless the child runs in an isolated invocation context.
  • Empty Tool Call Arrays: When using models through LiteLLM or local vLLM instances, some quantized models output tool_calls: []. If your router checks if message.tool_calls is not None:, it evaluates to True even though the list is empty, triggering an instant infinite loop. Always use if bool(getattr(message, "tool_calls", None)):.