How to Fix 'RuntimeError: Address already in use' in PyTorch Distributed on Linux

Quick Fix (TL;DR)

To clear the port collision immediately, kill the orphaned process holding the default master port:

# 1. Terminate any lingering process on default PyTorch distributed port 29500
sudo fuser -k 29500/tcp

# 2. Alternatively, launch torchrun with a dynamic ephemeral port
torchrun --nproc_per_node=4 --master_port=$(shuf -i 30000-60000 -n 1) train.py

PyTorch’s c10d::TCPStore binds to port 29500 by default on rank 0 to coordinate distributed workers. When a previous training session aborts without a clean teardown, child worker processes remain alive in the background holding the socket open.

Root Cause Analysis

When launching multi-GPU training with torchrun, accelerate, or torch.distributed.launch, PyTorch initializes an inter-process coordination store called TCPStore. Rank 0 attempts to bind a listening TCP socket to 0.0.0.0:29500.

If the port is occupied, PyTorch halts with this failure:

[E socket.cpp:926] [c10d] The server socket has failed to bind on [0.0.0.0]:29500 (errno: 98 - Address already in use).
Traceback (most recent call last):
  File "train.py", line 42, in <module>
    dist.init_process_group(backend="nccl")
RuntimeError: Address already in use

This error stems from three primary operational scenarios on Linux:

  1. Orphaned Worker Processes (Zombies): If you interrupt training with Ctrl+C or a job manager kills the parent torchrun process with SIGKILL, the child Python worker ranks do not receive termination signals. They remain running in the background, keeping the socket connection in an ESTABLISHED or CLOSE_WAIT state.
  2. Concurrent Multi-Job Collisions: Two independent training runs or multiple developers executing on the same multi-GPU node default to master port 29500 simultaneously.
  3. Dual-Stack IPv4/IPv6 Binding Conflicts: In containerized Docker or Kubernetes environments, localhost resolves to both 127.0.0.1 and ::1. If MASTER_ADDR is unset, TCPStore may attempt dual binding and collide with itself.
Scenario Root Cause Diagnosis Command Recommended Fix
Aborted Previous Run Orphaned rank 0 process still alive ss -tulpn | grep 29500 Kill process via fuser -k 29500/tcp
Shared GPU Server Another user running torchrun lsof -i :29500 Use random --master_port
Container / K8s Pod IPv4 vs IPv6 resolution conflict ping -c 1 localhost Export MASTER_ADDR=127.0.0.1

Step-by-Step Resolution

Step 1: Identify the Process Holding Port 29500

Inspect active socket listeners on your Linux machine using ss or lsof:

# Using ss (iproute2)
ss -tulpn | grep 29500

# Expected output:
# tcp   LISTEN 0      128          0.0.0.0:29500      0.0.0.0:*    users:(("python",pid=84291,fd=7))

Or check with lsof:

sudo lsof -i :29500

# Expected output:
# COMMAND     PID    USER   FD   TYPE DEVICE SIZE/OFF NODE NAME
# python    84291 beomjin    7u  IPv4 928374      0t0  TCP *:29500 (LISTEN)

Note the PID (84291 in this example).

Step 2: Safely Terminate Zombie Worker Processes

Terminate the specific blocking process:

# Kill the single blocking PID
kill -9 84291

If multiple multi-GPU worker processes are desynchronized across different GPUs, kill all detached distributed Python workers at once:

# Terminate all orphaned distributed PyTorch workers safely
pkill -9 -f "torch.distributed"
pkill -9 -f "torchrun"

Verify that the port has been released:

ss -tulpn | grep 29500
# Output should be empty

Step 3: Configure Dynamic Ephemeral Ports in Production Scripts

Never hardcode port 29500 in automated scripts. In PyTorch 2.4 and newer, you can pass --master_port=0 to instruct the launcher to discover an open ephemeral port automatically:

# Native automatic port selection (PyTorch >= 2.4)
torchrun --nproc_per_node=4 --master_port=0 train.py

For older PyTorch releases, generate an open ephemeral port directly in bash or Python:

# Bash: Random high-range port allocation
MASTER_PORT=$(python3 -c 'import socket; s=socket.socket(); s.bind(("", 0)); print(s.getsockname()[1]); s.close()')
echo "Selected Master Port: $MASTER_PORT"

torchrun --nproc_per_node=4 --master_port=$MASTER_PORT train.py

Step 4: Prevent IPv6 Dual-Stack Collisions in Containers

Inside Docker containers, ensure that PyTorch binds explicitly to IPv4 by setting environment variables in your launch script or Dockerfile:

# Explicitly set IPv4 localhost and master address
export MASTER_ADDR=127.0.0.1
export GLOO_SOCKET_IFNAME=lo
export NCCL_SOCKET_IFNAME=lo

# Run distributed workload
torchrun --nproc_per_node=4 --master_addr=127.0.0.1 --master_port=29555 train.py

Python Implementation: Automatic Free Port Discovery

If your pipeline initializes distributed process groups directly from Python code rather than CLI wrappers, wrap the socket allocation logic into a helper function before calling init_process_group:

import os
import socket
import torch
import torch.distributed as dist

def get_available_port() -> int:
    """Bind to port 0 to let the OS assign an available ephemeral port."""
    with socket.socket(socket.AF_INET, socket.SOCK_STREAM) as s:
        s.bind(('', 0))
        s.setsockopt(socket.SOL_SOCKET, socket.SO_REUSEADDR, 1)
        return s.getsockname()[1]

def init_distributed_node():
    # Only Rank 0 assigns and broadcasts the port if not set by launcher
    if "MASTER_PORT" not in os.environ:
        master_port = str(get_available_port())
        os.environ["MASTER_PORT"] = master_port
        print(f"[Rank Init] Allocated dynamic master port: {master_port}")

    if "MASTER_ADDR" not in os.environ:
        os.environ["MASTER_ADDR"] = "127.0.0.1"

    dist.init_process_group(backend="nccl")
    local_rank = int(os.environ["LOCAL_RANK"])
    torch.cuda.set_device(local_rank)
    print(f"[Rank {local_rank}] Connected to TCPStore successfully.")

if __name__ == "__main__":
    init_distributed_node()

Verification: Test Socket Binding Under Load

Verify that port allocation works without conflicts by simulating rapid consecutive launches:

# Run a quick 2-GPU test job
torchrun --nproc_per_node=2 --master_port=0 -c "
import torch
import torch.distributed as dist
dist.init_process_group(backend='nccl')
print(f'Rank {dist.get_rank()} successfully initialized.')
dist.destroy_process_group()
"

Expected terminal output:

Rank 0 successfully initialized.
Rank 1 successfully initialized.

If both ranks exit cleanly without an Address already in use trace, socket cleanup and dynamic port binding are functioning properly.

Common Pitfalls & Edge Cases

  • TCP TIME_WAIT Sockets: When a process closes a TCP socket, Linux keeps the connection in TIME_WAIT for 60 seconds (defined by tcp_fin_timeout). If you restart a job within seconds on the same port, the bind will fail. Using --master_port=0 bypasses this delay by selecting a fresh port.
  • Slurm Multi-Node Deployments: On Slurm clusters, all nodes must connect to the head node’s IP address. Do not use 127.0.0.1 on multi-node jobs. Instead, pass the allocation master hostname:
    MASTER_ADDR=$(scontrol show hostnames "$SLURM_JOB_NODELIST" | head -n 1)
  • Shared /dev/shm IPC exhaustion: Multi-GPU DDP crashes can sometimes leave lingering shared memory segments in /dev/shm. If training fails immediately after freeing the port, clean stale IPC memory with rm -rf /dev/shm/torch*.