AI Hallucinations and Saudi Tax Compliance

In September 2026, the New Mexico Supreme Court held criminal defense attorney Stephen Aarons in direct contempt after he submitted an appeal brief containing false testimony from four witnesses who do not exist [1]. ChatGPT had invented them. Aarons was fined $5,000, referred to a disciplinary board, and the court noted he "demonstrated a lack of remorse and a lack of concern for his client." [1] His explanation — "I didn't know that AI could hallucinate facts" — is the sentence that should be printed above every compliance workstation in Saudi Arabia.
The case is not a legal novelty. It is a precise illustration of what happens when a professional records an AI-generated output as established fact, without anchoring it to a verifiable source. The domain changes; the failure mode does not.
What the New Mexico Case Actually Shows About AI in High-Stakes Workflows
Aarons had practiced criminal defense for over 40 years [1]. His seniority did not protect him. What failed was a process: he admitted he did not verify the factual claims and legal authority in his AI-generated brief before signing and filing it [1]. He also did not inform his client. Two professional failures, both downstream of a single assumption — that the AI's output was traceable to something real.
The parallel for Saudi accounting offices is direct. A compliance officer who asks an AI tool to summarize a ZATCA notification, records the AI's figure as the penalty amount, and closes the ticket without attaching the original notice has made the same assumption Aarons made. If a ZATCA field auditor later requests the source of that figure, "the AI said so" is not a response that closes the audit.
The structural problem is not that AI tools are unreliable. It is that their outputs are self-contained — they do not carry a citation chain. When that output enters a compliance record, the record inherits the AI's confidence without inheriting any traceable authority.
Why Saudi Regulatory Compliance Has Zero Tolerance for Unverifiable Claims
الامتثال الضريبي في المملكة العربية السعودية يعمل ضمن بيئة تنظيمية تتسم بالتقاطع المستمر بين جهات رقابية متعددة: هيئة الزكاة والضريبة والجمارك، والمؤسسة العامة للتأمينات الاجتماعية (الغوسي)، ومنصة قوى، ووزارة التجارة. كل جهة تُصدر إشعارات بصياغات وتواريخ ورقم مرجعية محددة.
Three features of this environment make AI hallucination particularly dangerous:
- Obligation specificity. A GOSI contribution deadline is not an approximation — it is a calendar date tied to a payroll cycle. A hallucinated date recorded as fact does not cause a minor error; it causes a missed filing and a financial penalty.
- Cross-authority dependencies. A Qiwa Saudization ratio affects a company's Nitaqat tier, which in turn affects its ability to access certain government services. An AI that misreads or invents a ratio can trigger a cascade of incorrect downstream entries across multiple compliance domains. See also: Nitaqat Compliance for Accounting Offices.
- Audit recall depth. Saudi regulations require records to be retrievable and verifiable, not merely present. A compliance entry that cannot be traced to its originating notice does not satisfy the recordkeeping standard — it simply takes up space in a file. For a detailed treatment of what the law actually requires, see Audit Trails for Regulatory Notices: What Saudi Law Actually Requires.
The Specific Failure Modes: Where AI Hallucinations Hit ZATCA, GOSI, and Qiwa Filings
The failure modes are not theoretical. They follow predictable patterns:
ZATCA e-invoicing compliance. Phase Two onboarding requires specific technical integration records and clearance confirmations. An AI tool that summarizes a ZATCA notification without preserving the original XML clearance reference leaves a gap that cannot be closed retrospectively. The clearance number is the record; the summary is not. For a checklist-level treatment, see Phase Two e-Invoicing Audit Readiness.
GOSI monthly contributions. Each contribution cycle requires a matching confirmation from the GOSI portal. If an AI tool extracts a figure from a previous month's notification and applies it to the current period, the resulting entry looks correct in the compliance file but is wrong in the GOSI system. The discrepancy surfaces only at audit — at which point the absence of the source notice makes correction difficult. See GOSI Contribution Deadlines and the Recordkeeping Standard That Protects You.
Qiwa workforce compliance. Nitaqat calculations depend on headcount figures that change monthly. An AI summary of a Qiwa report that rounds, estimates, or misreads a workforce ratio produces a compliance record that does not match the authority's own data. See What a Qiwa–GOSI Compliance Dashboard Must Show an Accountant.
Zakat filings. Zakat calculations are fact-specific and entity-specific. Hallucinated ownership structures or asset classifications in a Zakat worksheet are not detectable from the AI output itself — they look like legitimate entries until a ZATCA assessor compares them against the registered entity data. See The Zakat Filing Lifecycle: Where Internal Compliance Teams Lose Control.
The FDA's experience with AI in regulated industries is instructive here. A 2026 warning letter cited a drug manufacturer for using AI to generate compliance records — procedures, specifications, production records — without adequate review by its quality unit, and for attributing its own lack of awareness of certain process requirements to the AI system's failure to flag them [4]. Regulators in high-stakes domains do not accept "the AI didn't tell me" as a compliance defense. There is no reason to believe ZATCA or GOSI would treat it differently.
What a Defensible Audit Trail Looks Like Under Saudi Recordkeeping Standards
A defensible compliance record has a structure that any auditor can traverse. It is not a summary. It is a chain.
The chain has four links:
- The source notice. The original document issued by the regulatory authority, preserved in its original format, with its timestamp, reference number, and issuing entity intact. This is the anchor. Everything else is derived from it.
- The extracted obligation. The specific action, amount, or deadline identified from the source notice, with a notation of who extracted it and when. If an AI tool performed the extraction, that is noted — and the extracted value is confirmed against the source notice by a human reviewer before the record is closed.
- The assigned owner. The named individual or team responsible for the action, with the date of assignment. Unowned obligations are the most common cause of missed deadlines in multi-client accounting offices.
- The completion evidence. The portal confirmation, payment receipt, or filed document that proves the obligation was discharged — again linked to the source notice by reference number.
Any compliance record that cannot produce all four links on demand is not audit-ready. It is a summary dressed as a record. See The Case for a Unified Compliance Register Across Client Portfolios for a structural approach to maintaining this chain across multiple clients simultaneously.
The risk of AI-generated compliance summaries is well recognized in the broader accounting and audit community: when AI outputs cannot be traced to source documents, they introduce a class of error that is difficult to detect precisely because the output looks authoritative [2]. Human oversight — specifically, a reviewer who checks the AI's extraction against the original notice — is not a redundancy. It is the control that makes the record defensible [3].
MAKYN's View: AI Should Surface Notices, Not Substitute for Them
The New Mexico case did not end Stephen Aarons' career because he used AI. It ended badly because he used AI as a replacement for source verification rather than as an aid to it [1]. That distinction is the entire compliance question.
MAKYN's position is that AI is genuinely useful in compliance workflows — for ingesting large volumes of regulatory notices, classifying them by urgency and authority, and routing them to the correct responsible party before the action window closes. Saudi accounting offices managing 180 to 400 client obligations per month cannot read every notice at the speed regulators issue them. AI triage is not optional at that scale; it is a structural necessity.
What AI cannot do is substitute for the source document. The compliance record — the entry that will be examined if ZATCA opens an assessment or GOSI flags a discrepancy — must link to the raw notice. The AI's extraction is a starting point for human review, not the record itself.
This is why MAKYN anchors every extracted obligation to its originating notice before it enters a client's compliance register. The AI reads and classifies; the notice travels with the obligation through every stage of the workflow. When an auditor asks for the basis of a recorded figure, the answer is a timestamped document from a named authority — not a language model's summary of what that document probably said.
If your compliance files currently contain AI-generated summaries that cannot be traced back to the original ZATCA, GOSI, or Qiwa notice, that gap exists right now. It will not surface until an audit opens it. اطلب عرضاً توضيحياً to see how a sourced, notice-anchored compliance register works in practice — before the auditor asks.
For a broader treatment of AI's role and limits in Saudi compliance automation, see Why AI Alone Is Not Enough in Tax Compliance and When AI Agents Go Rogue: The Audit-Trail Risk Saudi Finance Teams Can't Ignore.
Frequently asked
- What is an AI hallucination and why does it matter for Saudi tax compliance?
- An AI hallucination is a confident, plausible-sounding output that is factually wrong and untraceable to any real source. In Saudi tax compliance, a hallucinated ZATCA penalty figure or a fabricated GOSI deadline recorded as fact creates a compliance entry that cannot survive audit scrutiny. The original regulatory notice — not the AI summary — is the only legally defensible record.
- What happened with the New Mexico lawyer and ChatGPT?
- In September 2026, the New Mexico Supreme Court held attorney Stephen Aarons in direct contempt after he submitted a brief containing false testimony from four wholly fabricated witnesses, generated by ChatGPT without verification. He was fined $5,000 and referred to a disciplinary board. He admitted he did not verify the factual claims before filing — a cautionary example for any professional relying on AI in high-stakes recordkeeping.
- What does a defensible audit trail for ZATCA or GOSI obligations actually require?
- A defensible audit trail links each recorded obligation to its source: the raw regulatory notice (with its official timestamp and reference number), the obligation extracted from it, the team member assigned to act, and the evidence that action was completed. Any gap in this chain — particularly the absence of the source notice — transforms a compliance file into an unverifiable assertion that a ZATCA or GOSI auditor will reject.
- Can AI tools be used in Saudi compliance workflows at all?
- Yes, but within a defined scope. AI is effective for triage — surfacing which notices require action, flagging approaching deadlines, and drafting response templates. It is not a substitute for the source document itself. Every AI-generated output used in a compliance record must be anchored to the original authority notice, with a human reviewer confirming the match before the record is closed.
Sources
- 1. ChatGPT-using lawyer punished for citing fake testimony from made-up witnesses — rss:arstechnica-ai
- 2. AI Hallucinations in Accounting and Audit | Trullion — trullion.com
- 3. Guidelines for Handling Hallucinations in AI Tools | CXC — www.cxcglobal.com
- 4. FDA AI Compliance: Warning Letter Signals New Scrutiny of AI Use — www.morganlewis.com