- 20 August 2026
- No Comment
- 9
THE FOOTNOTE PROBLEM
By Nabeel Shaikh, FCA, MSc, FMVA, CME-1
Nabeel is a Fellow Chartered Accountant and financial expert specializing in corporate finance, advanced financial modeling, and capital market strategies. He brings deep industry expertise to navigating complex corporate governance, risk management, and the technological transformation of professional services.
What a year of reported citation failures says about AI governance in professional services
A note on sourcing: every factual claim below is attributed to a named publication and dated. Where a firm has responded publicly, its response is included alongside the finding. Where a firm disputes a characterisation, that dispute is stated. Statements of my own opinion are marked as such. Corrections are welcome and will be published.
Over the past fifteen months, research published by each of the four largest professional services networks has been reported to contain fabricated citations, unverifiable claims or case studies that the organisations named in them have disputed. Two reports were withdrawn by the publishing firms. One government contract was partially refunded. One provincial government requested a review.
The findings in each case were produced by GPTZero, an AI-detection research company, and in the more recent instances were verified independently by the Financial Times before publication.
My argument, stated plainly as opinion: the significant issue here is not that generative models produce unreliable citations. That property is well documented, and every firm involved has published client-facing material saying so. The issue is what the sequence suggests about the verification controls sitting between a draft and a published document, at firms whose commercial proposition in this market rests substantially on advising others how to govern exactly this technology.
1. What has been reported
Deloitte Australia, welfare compliance review. According to the Associated Press and the Australian Financial Review (October 2025), Deloitte Australia agreed to partially refund the A$440,000 (approximately US$290,000) paid by the Department of Employment and Workplace Relations for a 237-page independent assurance review of the IT framework used to automate welfare penalties. AP reported that the department stated Deloitte had confirmed some footnotes and references were incorrect, and had agreed to repay the final instalment under its contract. The errors reported included a quotation attributed to a Federal Court judgment and references to academic papers that could not be located. Chris Rudge, a researcher in health and welfare law at the University of Sydney, alerted media outlets to the references. AP reported that the department said the substance of the report was maintained and its recommendations were unchanged, and that the revised version disclosed the use of a generative AI language system, Azure OpenAI, in its preparation. Deloitte told AP the matter had been resolved directly with the client, and did not respond when asked whether the errors were AI-generated.
Deloitte Canada, Newfoundland and Labrador health workforce plan. The Independent, a Newfoundland news outlet, reported in November 2025 that a 526-page Health Human Resources Plan published by the province in May 2025 contained at least four citations to research papers that could not be found. CBC reported that the province contacted the firm and that Deloitte described the four citations as incorrect while standing by the report. Deloitte Canada told Fortune that it firmly stands behind the recommendations put forward, that it was revising the report to make a small number of citation corrections which do not affect its findings, and, importantly for accuracy here, that AI was not used to write the report but was “selectively used to support a small number of research citations.” Access-to-information records reported by The Independent put the contract value at just under C$1.6 million.
EY Canada, loyalty fraud study. The Financial Times reported in May 2026 that EY Canada removed a study titled Points of Attack: Uncovering Cyber Threats and Fraud in Loyalty Systems from its website after GPTZero published findings on apparent AI hallucinations and unverifiable footnotes. Per GPTZero’s published analysis, 16 of the study’s 27 references were fabricated, misattributed or led to unavailable pages, including a reference to a McKinsey report the researchers could not locate; the study also cited the same $200bn figure for both the size of the loyalty market and the value of unclaimed points. EY said it had removed the report and was examining how it was approved, that the study was not tied to any client assignment, and that EY Canada takes the accuracy of its published content seriously and maintains an organisation-wide commitment to responsible AI use.
KPMG, agentic AI study. TechCrunch and the Financial Times reported that on 13 June 2026 KPMG withdrew Total Experience: Redefining Excellence in the Age of Agentic AI, published in October 2025, after UBS, the UK’s National Health Service, Swiss Federal Railways and Transport for London told the FT that the report’s claims about their AI use were untrue or misleading. GPTZero’s published review states that five of the report’s 45 citations pointed accurately to their stated sources. A KPMG spokesperson said the firm removed the report from its websites while conducting its own investigation.
PwC Middle East, thought leadership reports. In late July 2026, the Financial Times reported on GPTZero research, which the FT verified, identifying fabricated citations, misattributed material and unverifiable information across four PwC Middle East publications issued between 2024 and 2026. GPTZero’s published analysis states that one 2025 report described a governance framework for which the researchers could find no evidence of the deployments claimed. A spokesperson for PwC Middle East said the firm takes the accuracy of its published research seriously and is “updating a limited number of supporting citations” in the reports identified, adding that, consistent with its approach to responsible AI, it maintains quality control processes for research and content development that it expects its people to follow.
Three of these documents concerned AI. Two concerned deploying it responsibly. That coincidence is a matter of public record, and readers can draw their own conclusions from it.
2. What the pattern suggests, and what it does not
Let me be precise about the limits of the inference, because precision is the point of the article.
These documents were not audit opinions. Four of the five were marketing or thought-leadership publications, subject to nothing resembling the quality architecture that governs a statutory audit or a signed assurance opinion. Nothing reported here bears on the quality of any firm’s audit work, and no reader should extend it there. Citation failures of the same kind have been reported over the same period in law firm filings, academic conference submissions and government policy documents; this is not a phenomenon confined to professional services. Vendors are not exempt either, Anthropic’s model produced an erroneous legal citation in a US court filing in 2025, a fact reported at the time and relevant to anyone treating AI-native providers as inherently safer.
What the reported sequence does support, in my view, is a narrower and more useful observation: verifying whether a footnote resolves to a real document is a mechanical check, it is cheap, it can itself be automated, and in these five instances the published findings indicate it did not happen before release. An organisation can hold a comprehensive AI usage policy and still have no operating control that tests output against source. Those are different things, and the distinction is the one worth taking to your own firm.
Three questions follow for any organisation publishing AI-assisted work:
Is there a verification step with a named owner? Not a principle in a policy document, a control, with an owner, an exception log and a rate.
Is disclosure default or retrospective? In the Australian case, per AP’s reporting, the use of a generative system appeared in the corrected version, not the original. Disclosure that arrives after discovery serves a different function from disclosure that arrives before it.
Does named authorship carry a verification obligation? Signing a document and having verified its evidence base are separable acts unless the process makes them the same one.
3. The tension in the market position
Big 4 leadership has consistently held in public that AI augments professionals rather than replacing them, and that agentic tools supplement human employees. I take that position as sincerely held. It also sits alongside a set of operating facts that clients can read independently.
On the product side: Deloitte’s Zora AI, built with Nvidia, targets finance and procurement workflows including invoice processing and variance reporting, with the firm stating it can release thousands of hours annually and reduce costs by up to 25%. Business Insider reported that EY has given roughly 80,000 tax professionals access to 150 agents, advanced around 1,000 agents into development or production during 2025 with a stated ambition of 100,000 by 2028, and invests over $1bn annually in AI platforms. PwC introduced its agent OS platform in March 2025 and has since deployed tens of thousands of agents across client operations. KPMG launched Workbench with Microsoft in June 2025 under a previously announced multi-billion-dollar alliance. Bloomberg Tax reported in March 2026 that agentic capability has changed the economics of managed services, allowing firms to run corporate service functions with fewer people and higher margins.
On the internal side, as reported by the Financial Times, City AM and Accountancy Age: UK graduate intake has been reduced across all four firms, with KPMG’s cohort falling from 1,399 in 2023 to 942 and reductions of 18%, 11% and 6% reported at Deloitte, EY and PwC respectively, against a 44% fall in UK accountancy graduate job adverts. PwC reduced US headcount by approximately 3,300 roles between September 2024 and May 2025 and, per City AM, saw global headcount fall by 5,600 in its last reported year. KPMG UK’s head of advisory, Lisa Fernihough, told the FT that she wants the organisation to still exist, offering that as her measure of how disruptive she believes AI will be.
I want to be careful here. Reducing graduate intake in a soft consulting market is an ordinary commercial decision, and all four firms have said they continue to invest in early-career talent. The firms have not claimed AI is headcount-neutral; several have said the opposite about their clients’ functions. The observation I would put to a board is simply this: a firm advising you that AI will not reduce your finance headcount, while its own published product literature quantifies the hours that same technology removes, is asking you to hold two propositions at once. Ask which one applies to your engagement, and get the answer in writing.
4. Where the competitive pressure is coming from
The model developers have moved into services. In May 2026, Anthropic announced an enterprise services venture with Blackstone, Hellman & Friedman and Goldman Sachs, reported by Fortune and the Wall Street Journal as backed by approximately $1.5bn in committed capital and structured to embed Anthropic engineers inside client operations; it has since been branded Ode with Anthropic. OpenAI has established a comparable deployment vehicle. The stated logic is the ratio of services spend to software spend, roughly six to one, which makes deployment, not licensing, the larger prize.
This is not a simple displacement story, and I would resist writing it as one. Deloitte announced in October 2025 that it would make Claude available to more than 470,000 people across 150 countries, described by Anthropic as its largest enterprise deployment, alongside a joint certification programme for 15,000 practitioners. The labs are supplier, distribution partner and prospective competitor at the same time. That is a more precarious position for the firms than outright rivalry would be, because the dependency runs in the direction of the party that can bypass it.
Former partners are building without the audit constraint. Unity Advisory launched with up to $300m from Warburg Pincus, chaired by former EY UK chair Steve Varley with former PwC COO Marissa Thomas as chief executive, deliberately excluding audit. Varley told the FT that clients are looking for a proposition that is “AI-led rather than based on legacy infrastructure,” and free of conflicts. Queen’s Tower Advisory, also founded by former Big 4 partners, follows a similar model; the Management Consultancy Association reports smaller firms growing at rates up to 50%.
The FT reported in February 2026 that PwC had sent letters to former partners who joined Unity, warning that post-retirement annuities funded from firm profits could be at risk, and had withdrawn healthcare scheme access for some. A person close to PwC told the FT that partners who move to competitors routinely forfeit such payments and that no contrary promises had been made. I include the firm’s position because it is a legitimate one: enforcing partnership terms is not misconduct, and readers should assess the dispute on both accounts.
The pricing model is under strain. The FT reported in 2026 that AI is forcing McKinsey and its peers to reconsider time-based pricing, with clients pushing toward outcome-linked fees; McKinsey now ties roughly a third of its work to performance-based arrangements. Where delivery cost falls and price does not, the difference is margin, and clients have begun asking about it.
5. What to require before you sign
The practical response is not to avoid the Big 4. Their audit licences, regulatory standing, global capacity and balance sheet depth remain genuinely difficult to replicate, and none of the reporting above touches their assurance work. The response is to procure differently, and to apply the same standard to every bidder:
- Disclosure of AI use in the engagement letter, which systems, which stages, which deliverables.
- Evidence that citation and claim verification exists as a control, with an owner and an exception rate you can see.
- Named human accountability, with the verification documented rather than assumed.
- A contractual remedy for fabricated content, agreed in advance rather than negotiated after discovery. The Australian case establishes that remedies are available; it also establishes that they get negotiated under pressure.
- Pricing tied to outcomes rather than effort, so that efficiency gains are visible in what you pay.
Apply all five to the AI-native challengers as well. A newer firm with no verification control carries the same exposure with less indemnity standing behind it.
Closing
The profession’s oldest product was never analysis. It was assurance, the willingness to put a name to a claim after checking it.
Nothing in the past fifteen months suggests these firms cannot rebuild that around AI; several have said publicly that they are doing exactly that. What the reporting does suggest is that the check itself has to be a designed control rather than an assumed professional instinct, and that the firms selling AI governance will be held to the standard they sell, which is as it should be.
That is a demanding standard. It is also the one this profession chose.
All figures and findings as reported at the time of publication and attributed to the sources named. Firms’ published responses are included where issued. This article expresses my opinion on the governance implications of publicly reported events; it makes no allegation of misconduct against any firm or individual.