Automatic Retries

Network glitches, momentary rate limits (HTTP 429), and remote service restarts (HTTP 503) are common when tools connect to third-party APIs. Failing an entire agentic workflow due to a single transient glitch degrades the user experience.

mcponce provides a built-in Exponential Backoff Retry Engine with customizable delays, factors, and selective error filters.


Basic Configuration #

You can enable retries by specifying the number of attempts or passing a detailed ToolRetryConfig:

app.tool({
name: 'fetch_weather',
description: 'Fetches weather with 3 automatic retries',
// Will retry up to 3 times with default exponential backoff (100ms, 200ms, 400ms)
retry: 3,
handler: async ({ city }) => {
const res = await fetch(`https://api.weather.com/v1?city=${city}`);
if (!res.ok) throw new Error(`Weather API failed: ${res.status}`);
return res.json();
}
});

Configuration Options #

OptionTypeDefaultDescription
attemptsnumberRequiredMaximum number of retry attempts after the first failure.
backoffMsnumber100Initial delay in milliseconds before the first retry.
factornumber2Multiplier applied to the backoff delay on each subsequent attempt.
maxBackoffMsnumber5000Upper limit for the backoff delay.
retryIf(err) => boolean() => truePredicate function. If it returns false, retrying stops immediately.

Overriding Retries per Invocation #

When invoking tools programmatically, you can override or disable retries:

TYPESCRIPT
// Disable retries for an idempotent check
const health = await app.callTool('execute_remote_sql', { query: 'SELECT 1' }, {
  retry: false
});

// Or customize attempts for a critical transaction
await app.callTool('execute_remote_sql', { query: 'COMMIT' }, {
  retry: { attempts: 5, backoffMs: 500 }
});

Telemetry & Metrics #

Every retry event is tracked in the telemetry pipeline:

  • In app.getAnalytics(), each tool lists its total retries and failedInvocations.
  • The Prometheus exporter emits mcp_tool_retries_total to track transient failure rates in Grafana dashboards.
Updated