Skip to main contentSkip to navigationSkip to footer
New: We launched Praxismith - practical courses on working with AI and production AI agents.Explore Praxismith
Eunix Tech - Software Engineering Company
Debugging Claude Tool Use in Production: A Practical Guide (Formerly Custom Actions)

Debugging Claude Tool Use in Production: A Practical Guide (Formerly Custom Actions)

Rajesh DhimanUpdated 13 min readAI Development

Fix Claude tool use bugs in production: bad arguments, missed tool calls, tool_use_id errors, loops, 429 and 529 retries, logging and evals.

We first wrote this post in January 2024 about "custom actions" in a Claude 3 app. That term comes from GPT Actions and OpenAPI schemas. Claude does not work that way. In the Claude API, the feature is called tool use (also called function calling). If you searched for "Claude custom actions", this is the page you want. The problems are the same as before: bad arguments, missed calls, broken conversations and rate limits. The names and the fixes are new.

All code is TypeScript with the official @anthropic-ai/sdk package. All API details come from Anthropic's documentation (see Sources).

How Claude tool use works

Tool use is a loop. Your app sends tools, Claude asks to call one, your app runs it and sends back the result.

  1. You send a request with a tools array. Each tool has a name, a description and an input_schema (a JSON Schema object) [2].
  2. If Claude wants to use a tool, the response has stop_reason: "tool_use" and one or more tool_use content blocks. Each block has an id, a name and an input object [3].
  3. Your code runs the tool.
  4. You send a new user message with a tool_result block for each call. Each result has the tool_use_id, the content and, if the tool failed, is_error: true [3].
  5. Claude reads the result and either calls another tool or gives the final answer.

These are client tools that run in your code. Anthropic also has server tools, such as web search, that run on Anthropic's side [1]. This guide is about client tools. Here is the tool we use in the examples:

import Anthropic from "@anthropic-ai/sdk";

const tools: Anthropic.Tool[] = [
  {
    name: "issue_refund",
    description:
      "Issue a refund for one customer order. Use this only after the user clearly asks for a refund " +
      "and gives an order ID. Do not use it to check order status. It returns the refund ID and the " +
      "amount refunded. It fails if the order is older than 90 days or already refunded.",
    strict: true,
    input_schema: {
      type: "object",
      properties: {
        order_id: { type: "string", description: "Order ID in the form ORD-123456." },
        amount: { type: "number", description: "Refund amount. More than 0 and not more than the order total." },
        reason: { type: "string", enum: ["damaged", "late", "wrong_item", "other"] }
      },
      required: ["order_id", "amount", "reason"],
      additionalProperties: false
    }
  }
];

Problem 1: Claude sends arguments that do not match your schema

Short answer: turn on strict tool use, and still validate every input in your own code.

Without strict mode, Claude may send a wrong type, like "2" instead of 2, or leave out a required field [4]. When you set strict: true on a tool, the API constrains the model's output so the input always follows your input_schema, and the tool name is always valid [4]. For strict tools, set additionalProperties: false on objects [5].

But strict mode uses a JSON Schema subset that does not support limits like minimum, maximum, minLength, maxLength and pattern [5]. So it can guarantee that amount is a number, but not that it is more than zero. It also cannot check business rules, like "the order must exist". You still need validation in code. Here is one with zod:

import { z } from "zod";

const RefundInput = z.object({
  order_id: z.string().regex(/^ORD-\d{6}$/),
  amount: z.number().positive(),
  reason: z.enum(["damaged", "late", "wrong_item", "other"])
});

async function issueRefund(input: unknown) {
  const parsed = RefundInput.safeParse(input);
  if (!parsed.success) {
    const problems = parsed.error.issues
      .map((issue) => `${issue.path.join(".")}: ${issue.message}`)
      .join("; ");
    throw new Error(`Invalid input: ${problems}. Fix these fields and call issue_refund again.`);
  }
  // parsed.data is now fully typed and checked
  return refundService.create(parsed.data);
}

When validation fails, do not crash. Send the error back as a tool result (see Problem 4). With an invalid or incomplete call, Claude will retry 2 to 3 times with corrections before it apologizes to the user [3].

Problem 2: Why is Claude not calling my tool (or calling it at the wrong time)?

Short answer: fix the tool description first. Then use the system prompt. Use tool_choice last, and check that your model supports the option you pick.

Anthropic calls detailed descriptions "by far the most important factor in tool performance" [2]. A good description says:

  • what the tool does,
  • when to use it, and when not to use it,
  • what each parameter means,
  • what the tool does not return.

Anthropic suggests at least 3 to 4 sentences per tool [2]. For complex or format-sensitive inputs, you can also add input_examples, which must match your schema [2]. Two more tips: group related actions into fewer tools, and prefix names by service, like github_list_prs [2].

With the default tool_choice of auto, Claude decides on each turn if it should call a tool. The system prompt moves this line. "Use the tools to investigate before responding." increases tool use. "Use your judgment about whether to call a tool or respond directly." makes it more careful [1].

tool_choice on the new models

tool_choice has four types: auto (Claude decides, the default), any (must use some tool), tool (must use one named tool) and none (no tools) [2].

Important: Claude Opus 5.5, Claude Sonnet 5.5 and Claude Fable 5.1 do not support forced tool use. Sending any or tool to them returns a 400 error [2][8]:

tool_choice: type "tool" and "any" are not supported for this model.

If forced tool calls started failing after a model upgrade, this is why. Use auto with strict tools and clear prompting, or structured outputs if you need the whole response as fixed JSON [2].

One more case: if a required value is missing from the user message, Claude Opus is more likely to ask for it, while Claude Sonnet may guess a value [1]. If a guess is dangerous (an order ID, an amount), check it in code and return an error that tells Claude to ask the user.

Problem 3: "tool_use ids were found without tool_result blocks immediately after"

Short answer: every tool_use block needs a matching tool_result in the very next user message, and the results must come first.

This is the most common 400 error in tool use code. The rules are [3]:

  • The tool results must come right after the assistant message that had the tool calls. No other message can sit between them.
  • In that user message, all tool_result blocks must come first. Any text must come after them.
  • Each result's tool_use_id must equal the id of the tool_use block it answers.

Common causes we see: the code answers only the first tool call, adds "Here are the results:" before the results, saves history in the wrong order after a retry, or filters the assistant content before sending it back.

The last one matters more on current models. With tool use, every thinking and redacted_thinking block must go back exactly as received. If you edit, reorder or remove them, the API returns a 400 error [8]. The safe rule: push response.content back into the history as it is.

Problem 4: How do I return tool errors so Claude can recover?

Short answer: send a normal tool_result with is_error: true and a clear message. Do not throw the error out of your agent loop.

Example from the docs [3]:

{
  "type": "tool_result",
  "tool_use_id": "toolu_01A09q90qw90lq917835lq9",
  "content": "ConnectionError: the weather service API is not available (HTTP 500)",
  "is_error": true
}

Claude will then explain the problem or try another way. Avoid generic text like "failed". Say what went wrong and what to do next, for example "Rate limit exceeded. Retry after 60 seconds." [3]. This function handles unknown tools, validation errors, thrown errors and timeouts in one place:

type Handler = (input: unknown) => Promise<unknown>;

const handlers: Record<string, Handler> = {
  issue_refund: issueRefund
};

function withTimeout<T>(promise: Promise<T>, ms: number): Promise<T> {
  return new Promise((resolve, reject) => {
    const timer = setTimeout(
      () => reject(new Error(`Timed out after ${ms} ms. Tell the user the service is slow.`)),
      ms
    );
    promise.then(
      (value) => { clearTimeout(timer); resolve(value); },
      (error) => { clearTimeout(timer); reject(error); }
    );
  });
}

async function runTool(call: Anthropic.ToolUseBlock): Promise<Anthropic.ToolResultBlockParam> {
  const handler = handlers[call.name];
  if (!handler) {
    return {
      type: "tool_result",
      tool_use_id: call.id,
      is_error: true,
      content: `Unknown tool "${call.name}". Available tools: ${Object.keys(handlers).join(", ")}.`
    };
  }
  try {
    const output = await withTimeout(handler(call.input), 10_000);
    return { type: "tool_result", tool_use_id: call.id, content: JSON.stringify(output) };
  } catch (error) {
    const message = error instanceof Error ? error.message : String(error);
    return { type: "tool_result", tool_use_id: call.id, is_error: true, content: message };
  }
}

The timer stops your agent from waiting forever, but it does not cancel the slow work. If your tool calls an HTTP API, also pass an abort signal to fetch.

Keep tool results small: return only high-signal fields and stable IDs [2]. And treat them as untrusted. Web pages, emails and third-party APIs can contain text that tries to give Claude new instructions, so keep that content inside tool_result blocks, not in the system prompt [3].

Problem 5: How do I handle parallel tool calls?

Short answer: one assistant turn can have several tool_use blocks. Answer all of them together in one user message.

Parallel calls are on by default [6]. You can run the calls at the same time with Promise.all or one after another. Read-only tools are usually safe to run together. Tools that write data or depend on each other are often better one by one [6].

Rules [6]:

  • Return one tool_result for each tool_use block, all in the same user message.
  • If you skip a call, still return a result with is_error: true and a short reason.
  • Separate user messages per result "teach" Claude to stop making parallel calls.

To turn parallel calls off, set disable_parallel_tool_use: true inside tool_choice, not at the top level. With auto, Claude then calls at most one tool per response [6].

tool_choice: { type: "auto", disable_parallel_tool_use: true }

Problem 6: Infinite tool loops and missing step limits

Short answer: always put a hard limit on the number of steps, and check stop_reason on every turn.

A model can keep calling tools, for example when a tool keeps failing. Your code must stop this. Here is a full agent loop with a step limit:

const client = new Anthropic({ maxRetries: 4, timeout: 60_000 });
const MODEL = "claude-sonnet-5-5";
const MAX_STEPS = 8;

async function runAgent(userText: string): Promise<Anthropic.Message> {
  const messages: Anthropic.MessageParam[] = [{ role: "user", content: userText }];

  for (let step = 1; step <= MAX_STEPS; step++) {
    const response = await client.messages.create({
      model: MODEL,
      max_tokens: 4096,
      tools,
      messages
    });
    logTurn(step, response);

    if (response.stop_reason === "max_tokens") {
      throw new Error("Response was cut off. Raise max_tokens and retry.");
    }
    if (response.stop_reason !== "tool_use") {
      return response; // end_turn, refusal, and so on
    }

    // Send the assistant content back exactly as received
    messages.push({ role: "assistant", content: response.content });

    const calls = response.content.filter(
      (block): block is Anthropic.ToolUseBlock => block.type === "tool_use"
    );
    const results = await Promise.all(calls.map(runTool));
    messages.push({ role: "user", content: results });
  }

  throw new Error(`Stopped after ${MAX_STEPS} steps without a final answer.`);
}

Why the max_tokens check? If the output is cut off while Claude writes a tool call, the tool_use block is incomplete. Retry with a higher max_tokens [7]. Other values: end_turn (done), refusal (Claude declined, read stop_details), pause_turn (a server tool loop paused) and model_context_window_exceeded (treat as cut off) [7].

Sometimes Claude returns an almost empty end_turn response. The docs link this to adding text right after tool results. Stop adding that text. If it still happens, send a new user message like "Please continue" [7].

If you do not need custom control, the SDK's tool runner (beta) runs this loop for you: client.beta.messages.toolRunner() with tools from betaZodTool(). Thrown errors go back to Claude with is_error: true, and maxIterations limits the loop [10][11]. We still write the loop by hand when we need approval steps or custom logging, as the docs suggest [11].

Problem 7: Rate limits (429) and overloaded errors (529)

Short answer: let the SDK retry, set a sensible maxRetries, and make sure your app handles the errors that are left.

The errors you will see most [8]:

  • 429 rate_limit_error: you hit a rate limit. The response has a retry-after header that says how many seconds to wait [9].
  • 529 overloaded_error: the API is busy across all users.
  • 500 api_error: an internal error. Retry with exponential backoff.

The TypeScript SDK retries connection errors, 408, 409, 429 and 500+ errors 2 times by default, with exponential backoff [10], and honours retry-after [8]. Change this with maxRetries on the client or per request [10].

Two special cases:

  • Spend cap. At your monthly spend cap, the API returns 429 with no retry-after header and error.details.error_code set to enforced_spend_limit_reached. Retrying will not help [9].
  • Acceleration limits. A sudden jump in traffic can cause 429 errors. Ramp up slowly [8].

Limits are per model, in requests, input tokens and output tokens per minute (RPM, ITPM, OTPM) [9]. Headers like anthropic-ratelimit-input-tokens-remaining show how close you are [9]. Prompt caching helps: for most models, tokens read from cache do not count toward ITPM [9]. Put cache_control: { type: "ephemeral" } on the last tool to cache all tool definitions. Changing any tool definition clears the whole cache, so keep the tool list stable [13].

For errors left after retries, catch the SDK's typed errors:

try {
  await runAgent(userText);
} catch (error) {
  if (error instanceof Anthropic.RateLimitError) {
    // 429 after all retries: queue the job or show "please try again soon"
  } else if (error instanceof Anthropic.InternalServerError) {
    // 500+ (including 529 overloaded) after all retries
  } else if (error instanceof Anthropic.APIConnectionTimeoutError) {
    // the request took longer than your timeout
  } else if (error instanceof Anthropic.APIError) {
    console.error(error.status, error.name); // for example 400 BadRequestError
  } else {
    throw error;
  }
}

The SDK's default timeout is 10 minutes, and timed-out requests are also retried 2 times [10]. For a chat app that is too long, so set a lower timeout, like the 60 seconds above. For long outputs, use streaming [8].

Problem 8: What should I log for Claude tool calls?

Short answer: log every turn with the request ID, stop reason, tool calls, tool results and token usage.

When a user says "the bot gave a wrong refund", you need to replay what happened. Log for every turn:

  • Request ID: response._request_id, from the request-id header [8][10]. Anthropic support needs it.
  • Model and stop_reason.
  • Tool calls: id, name and input (remove personal data first).
  • Tool results: tool_use_id, is_error, duration and a short version of the content.
  • Usage: input_tokens, output_tokens, cache_read_input_tokens, cache_creation_input_tokens [9].
  • Step number and a trace ID for the whole conversation.
function logTurn(step: number, response: Anthropic.Message & { _request_id?: string | null }) {
  const calls = response.content.filter(
    (block): block is Anthropic.ToolUseBlock => block.type === "tool_use"
  );
  console.log(JSON.stringify({
    step,
    request_id: response._request_id,
    model: response.model,
    stop_reason: response.stop_reason,
    tool_calls: calls.map((c) => ({ id: c.id, name: c.name })),
    usage: response.usage
  }));
}

For deeper debugging, set ANTHROPIC_LOG=debug or logLevel: "debug". This logs full request and response bodies, which may include personal data, so use it in development only [10].

Also watch tokens. Tool definitions, tool_use and tool_result blocks all count, plus a tool use system prompt of 286 tokens on Opus 5.5 and Sonnet 5.5 with auto [1]. To check a request's size first, use client.messages.countTokens(). It accepts tools, it is free (with its own rate limit), and it returns an estimate [12].

Problem 9: How do I test tool calling before it reaches users?

Short answer: build a test set of real user messages with the tool you expect, and run it on every prompt, tool or model change.

Many tool bugs come from a changed description, a new tool with a similar name, or a model upgrade. A small evaluation set catches them. Anthropic advises mirroring your real task mix, including edge cases, and automating grading, with many simple automated cases over a few hand-graded ones [15]. Check: the right tool (or no tool), the right arguments, a follow-up question when data is missing, and a clear recovery after a tool error.

const cases = [
  { input: "Please refund ORD-123456, it arrived broken", expectTool: "issue_refund" },
  { input: "What is your refund policy?", expectTool: null },
  { input: "Refund my last order", expectTool: null } // should ask for the order ID
];

for (const testCase of cases) {
  const response = await client.messages.create({
    model: MODEL,
    max_tokens: 1024,
    tools,
    messages: [{ role: "user", content: testCase.input }]
  });
  const call = response.content.find((b) => b.type === "tool_use");
  const actual = call && call.type === "tool_use" ? call.name : null;
  console.log(actual === testCase.expectTool ? "PASS" : "FAIL", testCase.input, actual);
}

Start with cases from your real logs, and add one each time you fix a production bug. See our LLM evaluation guide for the full process.

Quick debugging checklist

  1. Log the request ID, stop_reason and full response content.
  2. On a 400 error, check tool_result order, tool_use_id values and unchanged thinking blocks.
  3. Wrong arguments: add strict: true, zod validation and better parameter descriptions.
  4. Tool not called: improve the description, then the system prompt. No any or tool on Opus 5.5 or Sonnet 5.5.
  5. Endless loop: check the step limit and your tool error messages.
  6. 429 or 529: check retry-after, maxRetries, prompt caching and traffic spikes.
  7. Add the failing case to your test set.

Need help with Claude in production?

Tool use is easy to start and hard to make reliable. At Eunix Tech, we build and stabilise production LLM systems: agent loops, tool design, validation, logging, evaluation and cost control. If your Claude integration works in a demo but breaks with real users, see how we approach LLM architecture.

Sources

  1. Tool use with Claude - Anthropic, 2026
  2. Define tools - Anthropic, 2026
  3. Handle tool calls - Anthropic, 2026
  4. Strict tool use - Anthropic, 2026
  5. Structured outputs (JSON Schema limitations) - Anthropic, 2026
  6. Parallel tool use - Anthropic, 2026
  7. Handling stop reasons - Anthropic, 2026
  8. Claude API errors - Anthropic, 2026
  9. Rate limits - Anthropic, 2026
  10. TypeScript SDK - Anthropic, 2026
  11. Tool runner (SDK) - Anthropic, 2026
  12. Token counting - Anthropic, 2026
  13. Tool use with prompt caching - Anthropic, 2026
  14. Models overview - Anthropic, 2026
  15. Define success criteria and build evaluations - Anthropic, 2026

Frequently Asked Questions

Does Claude have custom actions like GPT Actions?

No. Claude does not use "custom actions" or OpenAPI schemas. It uses tool use. You define tools with a name, description and input_schema in the API request, Claude returns tool_use blocks, and your app sends back tool_result blocks [1][3].

Why do I get a 400 error about tool_use ids without tool_result blocks?

Your next user message does not answer every tool call correctly. Each tool_use block needs a tool_result with the same ID, in the message right after it, and the results must come before any text [3].

Can I force Claude to call a tool?

On some models, yes, with tool_choice set to any or tool. Claude Opus 5.5, Claude Sonnet 5.5 and Claude Fable 5.1 do not support this and return a 400 error. On these models, use auto with strict tools and a clear system prompt [2][8].

Does strict tool use remove the need for validation?

No. Strict mode makes sure the input matches your JSON Schema, but the supported schema does not include limits like minimum, maxLength or pattern [4][5]. You still need code checks (for example with zod) for value ranges and business rules.

How many times does the Anthropic SDK retry failed requests?

The official SDKs retry 2 times by default with exponential backoff. They retry connection errors, 429 and 5xx errors, and honour the retry-after header. You can change this with maxRetries [8][10].

Which Claude model should I use for tool use?

Anthropic suggests starting with Claude Opus 5.5 (claude-opus-5-5) for most workloads. Claude Sonnet 5.5 (claude-sonnet-5-5) is faster, and Claude Haiku 4.5 (claude-haiku-4-5) is the fastest [14]. Test your own tool cases on each model before you choose.

How do I stop an agent from calling tools forever?

Put a hard step limit in your loop, check stop_reason on every turn, and write tool error messages that tell Claude what to do next. If you use the SDK tool runner, set maxIterations [11].

Rajesh Dhiman

Written by

Rajesh Dhiman

Founder & CTO, Eunix Tech

Rajesh leads Eunix Tech's engineering practice, building production-grade applications, AI systems, and platform modernizations for global clients. He writes about the practical side of shipping software: what works in production, what fails, and why.

Let's Get Your AI MVP Ready

Book a free 15-minute call and see how fast we can fix and launch your app.

Related Articles

Fix Common Replit AI Errors: Complete Troubleshooting Guide 2025

Struggling with Replit AI errors? Our comprehensive guide covers the most common issues and their proven solutions.

🚀 Need your AI MVP ready for launch? Book a free 15-minute call.