You ask Claude to draft a service agreement. It delivers a professional-looking contract with three product SKUs, a 15% early-payment discount, and a liability cap that your company has never offered. None of it was in your template. The model hallucinated it. This is not an edge case. It is the default behavior of large language models—they interpolate plausibly when they don't know, and contracts are where interpolation becomes legal liability. A 2024 LLM contract stress test across Claude 3.5 Sonnet, GPT-4, and Meta Llama 3 showed that models consistently: Invented product codes when SKU lists were incomplete (73% of drafts) Generated pricing outside the input range by 12–28% on average Added warranty and liability clauses that contradicted template restrictions Misapplied or omitted regulatory guardrails in Southeast Asia contracts (Malaysia SST thresholds, Indonesia NPWP requirements, Singapore ACRA rules) The risk multiplies when contracts flow into e-signature workflows without audit. A hallucinated SKU sails past your approver, gets signed by the client, and becomes binding. Reconciliation happens at invoice time—or worse, at dispute. Where LLMs Fail Hardest: Three Concrete Fails Real examples from testing: 1. SKU Invention Under Ambiguity Setup: A consulting firm's template lists three service tiers (Basic, Pro, Enterprise) with placeholder pricing. The prompt says "Fill in realistic pricing for a mid-market SaaS onboarding engagement." Claude's output: Adds a fourth tier, "Accelerated," with a price point ₹180K/month—never defined in the template, never approved by your team. What happened: The model inferred a gap in the market and filled it. It was statistically plausible. It was completely unauthorized. 2. Pricing Drift Beyond Context Setup: You provide GPT-4 with a rate card (₹5K–₹8K per hour for implementation) and ask it to draft a fixed-price statement of work for a 200-hour project. GPT-4's output: Calculates the low end (₹5K × 200 = ₹1M) but adds a 12% "complexity buffer" for a total of ₹1.12M. The 12% buffer was never in your rate card, your cost model, or your margin targets. What happened: The model applied a logical rule (larger projects justify a buffer) that felt reasonable but wasn't sourced from your data. The client sees ₹1.12M; your sales team quoted ₹1M. Awkward. 3. Regulatory Clauses Misapplied Across SE Asia Setup: You use Llama 3 to draft a data-processing addendum (DPA) for Malaysia. Your template includes LGPD language (which is Brazilian). You ask the model to "make it compliant." Llama's output: Keeps LGPD references, adds Malaysian Personal Data Protection Act (PDPA) clauses, but mixes in Singapore Personal Data Protection Act (PDPA) subsections on data transfer—different rules, different thresholds. The liability section references South Africa's POPIA. What happened: The model pattern-matched across multiple legislations and blended them. A Malaysian data controller reviewing this would reject it immediately (or worse, sign it and face audit risk). Why This Happens: The LLM Confidence Problem Language models are probability engines, not knowledge vaults. They are trained to complete patterns, not to retrieve exact facts. When you ask Claude for "realistic pricing," it doesn't query a database; it predicts the next plausible number. When it encounters a SKU code missing from context, it doesn't say "I don't know"—it patterns-matches to similar code structures and invents one that feels right. This is baked into the architecture. The model cannot distinguish between "I saw this in my training data" and "this sounds like it should exist." Worse: the more fluently a model writes, the more confident you are that it is accurate. Hallucinations in contract language are especially dangerous because they are embedded in persuasive, formal prose. A fake SKU in a rambling email is obviously wrong. A fake SKU in a legally formatted contract? You might miss it. The Audit-Before-Sign Workflow: Five Checkpoints You cannot stop LLMs from hallucinating. You can stop hallucinations from leaving your organization. Checkpoint 1: Lock Your Input Data (Read-Only Knowledge Base) Do not feed LLMs your pricing spreadsheet and ask them to "use this." Instead: Extract pricing, SKU, and product data into a read-only reference document with explicit constraints: "Use ONLY these SKUs. Do not invent variants." Format it as a structured list, not prose. Models are less likely to interpolate when data is clearly delineated. Include a negative list: "Do NOT include these clauses: liability caps over ₹5M, warranty terms beyond 12 months, performance credits." Version-control this reference. If you update pricing, regenerate contracts with the new version. Checkpoint 2: Constrain the Prompt (Few-Shot Examples, Hard Guardrails) Use few-shot prompting—show the model 2–3 examples of correct contracts drafted from your template. Include an example of what not to do: "Here is an example contract using SKUs f