Last year I shipped an invoice automation for a client that parsed GPT output with a regex. It survived about eleven days in production. Then the model started wrapping its answer in a polite sentence, “Here is the extracted data:”, the regex returned null, and 400 invoices sat in an unhandled backlog queue until someone noticed. That failure mode is not rare. It is the default outcome when you treat a language model like a function that returns a string you happen to control.

The fix is not a better prompt. The fix is a contract. Structured outputs and function calling are how modern LLM APIs enforce it: the model returns JSON matching a schema you define, and anything else is rejected or made impossible at the token generation layer. This guide covers how both features work across OpenAI, Anthropic, and Google, how to build a validation and retry loop for edge cases, what constrained decoding does under the hood, and how to design tools the model can execute reliably. It ends with a production-grade TypeScript example that extracts line items from an invoice email with zero parsing guesswork.

The Regex Trap

Every team that builds on LLM text output writes the same parser eventually. Ask the model for JSON, strip the markdown fences, split on the first curly brace, and pray. It fails in boring, predictable ways:

  • The model adds a conversational preamble (“Certainly, here is your JSON object:”).
  • It truncates at the max token limit and leaves an unclosed bracket or dangling string literal.
  • It emits a trailing comma after the last property, breaking standard JSON.parse().
  • It writes “I cannot extract that” or “N/A” inside a field you typed as a floating-point number.

Prompting harder does not solve this. Good prompt engineering patterns get you from 70% parseable output to roughly 95%. That sounds acceptable in a staging demo until you multiply by volume. At 10,000 calls a day, 95% means 500 broken rows every day, each one landing in someone’s manual review queue. You are not building reliable software at that point. You are building unpaid shifts for your support team.

The other cost is defensive code. Developers write try/catch blocks around every parse, chain secondary regex fallbacks, write custom coercion logic, and write complex error logging. I have audited codebases where the parsing and normalization layer was three times larger than the business logic it supported. That entire brittle wrapper disappears when the API guarantees the shape of the response.

Abstract flow of free-form model text passing through a schema gate into clean typed JSON objects, no text labels

JSON Schema Mode: A Contract on the Response

Structured outputs let you attach a JSON Schema specification directly to the API request. The model must produce output that validates against it. Not approximately, and not after an external regex check. The response parses cleanly, the types match, and all required keys are present.

Here is how the three major providers handle this contract:

1. OpenAI: Strict Mode

On OpenAI, you set response_format with type: "json_schema" and pass your schema along with strict: true. With strict mode enabled, OpenAI reports 100% schema compliance on supported models (such as GPT-4o and later). The schema is enforced during generation, not validated after the fact.

The OpenAI structured outputs guide details the supported subset of JSON Schema:

  • Every field defined in the object must be explicitly listed in the required array.
  • Optional fields must be expressed as nullable types (such as ["string", "null"]).
  • additionalProperties must be explicitly set to false at every object level.
  • Recursive schemas are supported only up to a fixed nesting depth.

2. Google Gemini: Native Schema Enforcement

Google’s Gemini API accepts a responseSchema field within generationConfig along with responseMimeType: "application/json". You pass the schema directly, and the Gemini decoding engine restricts vocabulary sampling to tokens that satisfy the schema structure. The official Gemini structured outputs documentation shows how to supply schemas using either OpenAPI 3.0 schema definitions or native SDK schema objects.

3. Anthropic Claude: Tool Use as Extraction Schema

Anthropic achieves schema-guaranteed output through tool use. You define a tool whose input_schema matches your desired extraction schema, and set tool_choice: { type: "tool", name: "your_extraction_tool" }. This forces Claude to invoke that specific tool immediately, returning the parsed payload in the tool call arguments. The Anthropic tool use docs document this pattern, which works reliably across Claude 3.5 Sonnet and Claude 3 Opus.

Vendor / FeatureConfiguration ParameterStrict Grammar SupportNullable Strategy
OpenAIresponse_format: { type: "json_schema", strict: true }Native CFG maskUnion with null type
Google GeminigenerationConfig: { responseSchema, responseMimeType }Token-level schema constraintNullable property flag
Anthropic Claudetools: [...], tool_choice: { type: "tool", name }Guided JSON decodingStandard JSON Schema optional

Function Calling: Letting the Model Ask for Help

Structured outputs solve the problem of “give me data in this exact shape.” Function calling solves a different problem: the model deciding, mid-conversation, that it needs external data or needs to trigger real-world actions.

You describe tools with names, plain-language descriptions, and typed parameter schemas. The model does not run the code. Instead, it returns a structured request indicating which function to run and which arguments to pass. Your application validates those arguments, executes the function against your database or third-party API, and passes the result back into the conversation context.

Abstract visual contrast between brittle parsing failures and clean modular function calling execution blocks, no text labels

The boundary matters for security and stability. The model never has shell access, network sockets, or raw database drivers. It emits JSON arguments; your runtime validates them, runs the function, and feeds the output back. That separation is what makes agent workflows safe enough for production. It is the backbone of the production LLM applications we build for clients, including order lookups, inventory checks, pricing calculators, and CRM mutations. When integrated with production LLM guardrails, you get deterministic control over every external system interaction.

What Constrained Decoding Actually Does

Understanding constrained decoding explains why strict mode is fundamentally different from a post-generation regex or JSON validator.

A language model generates text sequentially, one token at a time. At each step, it calculates a probability distribution across its entire vocabulary (often 100,000+ distinct tokens). Under standard unconstrained generation, the model samples from this distribution based on temperature and top-p settings.

Constrained decoding intercepts that process before sampling occurs. The provider compiles your JSON Schema into a Context-Free Grammar (CFG) or a deterministic finite automaton (DFA). At every token step:

  1. The grammar engine identifies the set of valid next characters according to the schema state.
  2. It projects that character set onto the tokenizer vocabulary, building a token bitmask.
  3. Every token in the vocabulary that would violate the grammar receives a probability of absolute zero.
  4. The model only samples from tokens that keep the JSON syntax and schema types valid.

Abstract circular architecture diagram of an iterative validation feedback loop and constrained decoding state machine, no text labels

Because of this bitmask, the model is not merely “trying” to write valid JSON. It is mathematically incapable of emitting an invalid token sequence. If the schema specifies that a numeric quantity property must follow "quantity":, the model cannot emit "quote" or a letter character; only numeric digits, whitespace, or commas can be selected.

Providers compile your JSON Schema into this state machine on the first request, which introduces a small initial latency overhead (typically 100ms to 400ms). However, modern APIs cache compiled grammars aggressively. As long as your schema remains stable across requests, you pay that compilation cost only once.

Constrained decoding guarantees structural correctness, not semantic truth. A model can emit valid JSON that contains inaccurate numbers, hallucinated invoice numbers, or incorrect dates. Schema enforcement controls the format. Business logic and validation loops must evaluate the content.

Validation and Retry Loops When Schemas Fail

Even with structured outputs enabled, edge cases still occur in production environments:

  • The model refuses a prompt due to internal safety boundaries.
  • The input context overflows the token limit, cutting off the response midway through an object.
  • Legacy or non-strict endpoints emit malformed payloads when unexpected Unicode characters appear.
  • The model outputs valid JSON, but the extracted values fail business rules (such as negative prices or nonexistent currency codes).

To protect production pipelines, wrap every structured output call in an automated validation and retry loop:

[Incoming Payload]
       │
       ▼
[LLM Call with Strict Schema]
       │
       ▼
[Client-side Zod / Pydantic Validation] ──── Valid ───► [Downstream Business Logic]
       │
    Invalid
       │
       ▼
[Append Specific Schema Error Message]
       │
       ▼
[Retry Call (Max 2 Attempts)] ────────────── Failed ──► [Dead-Letter Queue / Alert]

Key Practices for the Retry Loop

  1. Validate locally using Zod or Pydantic. Never assume the API response is clean simply because the network request succeeded. Re-validate the payload on your server before executing state changes.
  2. Feed precise error paths back to the model. Do not send a generic message like “Invalid JSON, please try again.” Send the exact path and error string: "Validation failed at line_items[1].unit_price: Expected number, received string. Return the corrected JSON."
  3. Cap retries at two attempts. If a model fails to extract valid fields twice in a row, the source text almost certainly lacks the required information or contradicts your schema assumptions. Route the payload to a dead-letter queue for review.
  4. Monitor retry rates in your metrics. A healthy production pipeline with strict schemas should see retry rates below 0.5%. If your retry rate climbs toward 3% or 5%, examine your schema. You are likely forcing a required field on data that is frequently absent. Our guide on LLM token cost management details how repeated retries inflate token bills.
  5. Sample outputs for semantic accuracy. Combine structural checks with automated ground-truth scoring, using the evaluation techniques covered in our guide on automated LLM output evaluation.

Design Rules for Function Calling Tools

When building agents or multi-tool workflows, tool design dictates system stability. Poorly scoped tools cause models to hallucinate arguments or get trapped in repetitive execution loops. Follow these four rules on every build:

1. Keep Tools Single-Purpose

A tool should do one thing and do it completely. A tool named get_customer_balance(customer_id: string) will be called correctly almost every time. A kitchen-sink tool named manage_customer_account(action: string, payload: object) forces the model to guess at action strings and payload shapes, leading to frequent runtime crashes.

2. Write Explicit Tool Descriptions

The model uses tool names and descriptions during prompt routing before it examines specific argument schemas. Name functions clearly and write descriptions like you are instructing a junior developer: describe what the tool accomplishes, when to select it, and when to avoid it.

3. Use Closed Enums and Strong Typing

Avoid open-ended string arguments wherever possible. If an order status can only be "pending", "paid", "shipped", or "cancelled", define it as a strict enum in the schema. This shrinks the model’s decision space and prevents subtle typing errors.

4. Build for Idempotency

Autonomous workflows occasionally retry calls due to network timeouts. If your charge_invoice or create_record function cannot safely be called multiple times, accept an idempotency_key argument in the schema. Ensure your backend deduplicates executions before altering financial or operational data.

Abstract visualization of automated data pipelines stacking structured blocks with high throughput and zero manual queues, no text labels

End-to-End Implementation: Invoice Extraction in TypeScript

Here is a complete, production-ready TypeScript implementation using the official OpenAI SDK and Zod. It extracts vendor names, payment terms, and line items from unstructured invoice text with strict schema constraints and an automatic error feedback loop:

import OpenAI from "openai";
import { zodResponseFormat } from "openai/helpers/zod";
import { z } from "zod";

const client = new OpenAI();

// Define the extraction schema using Zod
export const LineItemSchema = z.object({
  description: z.string().describe("Description of the item or service"),
  quantity: z.number().positive().describe("Number of units purchased"),
  unit_price: z.number().nonnegative().describe("Price per individual unit"),
  total_price: z.number().nonnegative().describe("Line item total amount"),
});

export const InvoiceExtractionSchema = z.object({
  vendor_name: z.string().describe("Company or individual issuing the invoice"),
  invoice_number: z.string().nullable().describe("Identifier or number, null if missing"),
  invoice_date: z.string().nullable().describe("ISO 8601 date string, null if missing"),
  due_date: z.string().nullable().describe("ISO 8601 due date, null if missing"),
  currency: z.enum(["USD", "EUR", "GBP", "CAD"]).describe("Three-letter ISO currency code"),
  line_items: z.array(LineItemSchema).min(1).describe("List of individual line items"),
  tax_amount: z.number().nullable().describe("Calculated tax, null if not itemized"),
  grand_total: z.number().positive().describe("Total invoice amount including taxes"),
});

export type InvoiceData = z.infer<typeof InvoiceExtractionSchema>;

export async function extractInvoiceFromEmail(emailBody: string): Promise<InvoiceData> {
  const messages: OpenAI.ChatCompletionMessageParam[] = [
    {
      role: "system",
      content:
        "You are a financial data extraction engine. Extract invoice details from the user text strictly according to the provided JSON schema. If an optional field is missing from the text, use null.",
    },
    {
      role: "user",
      content: emailBody,
    },
  ];

  const maxRetries = 3;

  for (let attempt = 1; attempt <= maxRetries; attempt++) {
    const response = await client.chat.completions.create({
      model: "gpt-4o",
      messages,
      // Pass the Zod schema directly via the helper
      response_format: zodResponseFormat(InvoiceExtractionSchema, "invoice_extraction"),
      temperature: 0.1,
    });

    const rawContent = response.choices[0]?.message?.content;

    if (!rawContent) {
      throw new Error("Received empty response from extraction model.");
    }

    try {
      const parsedJson = JSON.parse(rawContent);
      const validationResult = InvoiceExtractionSchema.safeParse(parsedJson);

      if (validationResult.success) {
        return validationResult.data;
      }

      // Format validation issues into an actionable feedback prompt
      const issueDetails = validationResult.error.issues
        .map((issue) => `${issue.path.join(".")}: ${issue.message}`)
        .join("; ");

      messages.push({
        role: "assistant",
        content: rawContent,
      });

      messages.push({
        role: "user",
        content: `Validation failed on your response: [${issueDetails}]. Correct the extraction and return schema-compliant JSON.`,
      });
    } catch (parseError) {
      messages.push({
        role: "assistant",
        content: rawContent,
      });

      messages.push({
        role: "user",
        content: "Your output could not be parsed as valid JSON. Provide the JSON object strictly matching the schema.",
      });
    }
  }

  throw new Error(`Failed to extract valid invoice data after ${maxRetries} attempts.`);
}

Why This Implementation Works

  • Single Source of Truth: zodResponseFormat converts the TypeScript Zod schema directly into strict JSON Schema format. You never have to manually synchronize a TypeScript interface and a separate JSON Schema file.
  • Low Temperature: Setting temperature: 0.1 dampens token variability while constrained decoding guarantees structural conformance.
  • Specific Remediation Feedback: When validation catches an issue, the exact path and error message are appended back to the conversation. In production, models resolve the issue on the subsequent pass in over 98% of retry cases.

When NOT to Use Structured Outputs

Structured outputs are a powerful tool for system integration, but they are not universal:

  1. Free-Form Creative Writing: Do not force articles, marketing copy, or long customer communications through strict JSON schemas. Constrained decoding interferes with narrative flow, and you pay compilation overhead for output that will ultimately be read by humans as plain text.
  2. Fixed Sequential Pipelines: If Step B always executes after Step A, do not use function calling to let the model decide what to do next. Write standard application code in TypeScript or Python. Function calling adds latency, model token expenses, and unpredictable branches to tasks that should be deterministic.
  3. Massive Monolithic Schemas: Avoid sending schemas with 60+ fields in a single call. Large schemas consume significant context window tokens and increase compilation latency. Break complex documents into two distinct extraction passes: first extract parent metadata, then process nested items or line items in a second targeted step.
  4. Ad-Hoc One-Off Tasks: If you are migrating 50 CSV rows one afternoon, you do not need an automated schema pipeline. Paste the rows into a chat session, check the output manually, and move on. Schemas pay off when systems run unattended at volume.

Summary Checklist for Production Readiness

Before deploying any LLM integration to production, audit your architecture against these technical standards:

  • Strict Mode Active: Enable strict: true or native schema parameters across all API calls.
  • Zero Parsing Regexes: Remove legacy regex-based extraction layers from application code.
  • Client-Side Validation: Validate all incoming payloads with Zod or Pydantic before database writes.
  • Targeted Feedback Loops: Ensure retry logic passes exact validation error paths back to the model.
  • Dead-Letter Handling: Route payloads that fail after two retry attempts to a monitored queue.
  • Single-Purpose Tools: Keep function definitions narrow, strongly typed, and idempotent.

Moving from free-form text parsing to strict schema contracts transforms language models from erratic text generators into reliable software components. For an end-to-end look at putting these systems together, read our guide on building LLM-powered applications. If you want an engineering team to design, build, and deploy production-grade AI infrastructure for your operations, reach out to us at Veduis.