top of page

The Mirage of Total Autonomy: Why Human Oversight is the Ultimate Risk-Mitigation Framework in the Age of AI

Aug 23
9 min read

This article is adapted from my paper, “The Imperative of Human Oversight in the Age of AI,” which serves as the third module in the three-part “Strategic Financial Leadership” series. The module explores why organizations must balance rapid technological advancement with meaningful human accountability.


My interest in Artificial Intelligence (AI) grew after I co-founded BURCH Business Services (BBS), which focuses on bookkeeping, fractional CFO support, and management training. AI supports each area: QuickBooks Online automates bookkeeping data entry; AI-assisted workflows help build client spreadsheets and consolidate dashboards for fractional CFO work; and, in training, AI helped connect three seemingly separate papers into one coherent course. After I completed the human oversight paper, I asked AI how it could relate to our earlier work on the Theory of Constraints (TOC) and Internal Controls over Financial Reporting (ICFR). It recommended combining the papers into a three-part “Strategic Financial Leadership” module series and even drafted the initial outline. It was a brilliant idea! With human review and edits, we'll release that series soon. For us, this ability to turn separate papers and ideas into an integrated learning experience captures AI's true value.


Now for the “but.” This is where AI gets tricky: it’s not always right, and if you are not paying close attention, problems can follow. I shared a recent example in a LinkedIn post. Before a Zoom conversation about a potential service collaboration, I asked AI to research the person and suggest ways to approach the discussion. It returned with some interesting background information on the person and recommended using some of it as an icebreaker. The problem was that some of the information was wrong, so the icebreaker did not land as planned. In my case, it was a little embarrassing—but also somewhat funny.


Some time ago, I had the opportunity to meet with prominent AI researcher Dr. Mark Chang, author of Foundation, Architecture, and Prototyping of Humanized AI: A New Constructivist Approach. His book focuses on Humanized AI (HAI)—artificial social beings designed to closely resemble human behavior and interaction. That possibility is both exciting and unsettling. During our conversation, I raised the “black box problem,” and his work on that issue eventually led me back to the governance framework I discuss later. In simple terms, the black box problem occurs when a deep neural network’s internal mechanics become too complex for humans to trace or understand.  (Rudin, 2019)


Let that sink in! 


DNNs are simply too complex for humans to trace or understand.

The irony is striking: humans design these systems, yet once data passes through millions or billions of interconnected parameters, the path to the final output can become impossible to trace. To reduce the risks of this black box problem, Dr. Chang proposes an evolving research approach called Third-Wave Contextual Adaptation. In simple terms, the system learns gradually from small, basic data blocks, much like a child. This approach makes its reasoning layers more organized, traceable, and aligned with human cognitive development—hence the term “Humanized AI.” Still, Dr. Chang’s HAI concept remains a work in progress, which means organizations must manage the black box problem today. Too many still fall for the “mirage of total autonomy”—the mistaken belief that AI can be treated as a “set-and-forget” asset. Now, to be honest, I have fallen into that trap myself at times, accepting AI outputs without enough validation. Even small organizations must protect their reputations; for large corporations, an AI mistake can be catastrophic, as the Air Canada and Zillow failures below show.

Air Canada and its Rogue Chatbot (2022)

What happened. In 2022, a customer named Jake Moffatt asked Air Canada's website chatbot about bereavement fares. The chatbot "hallucinated" a fake policy, telling him he could claim a refund retroactively within 90 days. When he applied for the refund, Air Canada refused, absurdly arguing in court that the chatbot was a "separate legal entity" responsible for its own actions. A Canadian tribunal ordered the airline to pay compensation. (Moffatt v. Air Canada, 2024)


Where the AI went wrong. This is a classic case of Large Language Model (LLM) unpredictability and poor source-binding governance. The AI drifted from the official knowledge base and generated false information without a safety system catching the contradiction before the user saw it. (Moffatt v. Air Canada, 2024)

Zillow’s Machine Learning & Predictive Model Failure (2021)

What happened. Zillow launched "Zillow Offers," an iBuying program that used predictive machine learning algorithms to automatically value, buy, renovate, and flip houses. The algorithm failed to accurately capture local nuances and panicked pandemic-era real estate shifts. It aggressively overvalued homes, leading Zillow to buy thousands of properties it couldn't sell for a profit. The program collapsed, resulting in losses exceeding $500 million, prompting a full shutdown of Zillow Offers and layoffs affecting roughly a quarter of the company's workforce.  (Zillow Group, Inc. 2021)


Where the AI went wrong. The algorithm failed to adapt to a changing environment because it was trained on stable, pre-pandemic historical data. Corporate leadership also made the mistake of relying on the predictive model as an autonomous purchasing engine rather than an advisory tool.  (Zillow Group, Inc. 2021)


In AI terms, the Air Canada chatbot case illustrates algorithmic hallucination: the chatbot gave a customer false information. Although the direct financial cost was modest—$812.02 CAD in compensation and legal fines—the greater damage was reputational. A small customer-service error became a textbook example of corporate tone-deafness and poor AI governance.  (Moffatt v. Air Canada, 2024)


The Zillow case shows what can happen when a model is trained for one environment but deployed in another. As market conditions shifted, the model’s assumptions no longer matched reality, creating a silent gap between past data and present risk.  (Zillow Group, Inc. 2021)

Balanced Scorecard as an AI Governance Tool, Not Just a Calculator

These two AI failures raise an important question: what governance tool can best reduce the risk of similar breakdowns? One answer is the Balanced Scorecard (BSC), a strategic performance management framework that first captured my attention early in my career as a corporate controller and later as a finance director. What appealed to me was the BSC’s move beyond purely financial metrics toward broader drivers of value, including customer service, retention, and organizational learning.


Developed back in 1992 by Robert S. Kaplan and David P. Norton, the concept was easy—“what gets measured is what gets done.” Kaplan and Norton argued that relying solely on financial metrics, like ROCE, profit margins, ROI, etc., is like “driving a car by only looking at the rearview mirror. It will tell you where you’ve been—not where you’re going.”  (Kaplan & Norton, 1992)


The traditional BSC model consists of four perspectives described in the table below:  (Kaplan & Norton, 1992)


Kaplan and Norton didn't intend these as four separate silos. They designed the model as a bottom-up hypothesis: investment in the foundational tier is supposed to flow upward and eventually show up in the financial results at the top. Kaplan & Norton’s “vertical vector,” shown below in Figure 1, is used to describe a chain of “cause-and-effect relationships" that links non-financial drivers to financial outcomes. The figure shows how the Learning & Growth perspective—skills, capabilities, and processes—supports the Internal Business Perspective, which measures quality, efficiency, and timeliness. These internal drivers strengthen the Customer Perspective, including relationships and loyalty, which ultimately supports the Financial Perspective through profitability measures such as ROI, net profit, and gross and operating margins.  (Kaplan & Norton, 1992)


Figure 1: Kaplan & Norton Vertical Vector

Image created with ChatGPT (August 2026)

When I used BSC years ago, our challenges centered on balancing physical automation with manual oversight. Today, as AI automates financial reporting and algorithmic decision-making at a much faster pace, what interests me is BSC's potential shift from a static, backward-looking tracking tool to something closer to a real-time predictive decision engine.


What the BSC offers isn't a better calculator; it's a governance structure that assigns each timescale to its own organizational perspective, so acute, chronic, and catastrophic AI risk are each owned and reported by name, instead of getting collapsed into one undifferentiated "AI risk" line that hides which failure mode is actually driving the number.


This is where Kaplan and Norton’s “cause-and-effect chain” proves its continued value rather than becoming a relic. Their model already treats Financial performance as a lagging outcome driven by upstream performance in Customer, Internal Process, and Learning & Growth.  (Kaplan & Norton, 1992)

Structural Governance to an AI-Augmented Scorecard

A key distinction between Air Canada and Zillow has to do with the timing of the failures. Air Canada’s breakdown was acute, occurring in a single customer interaction, while Zillow’s was chronic, unfolding gradually as its model failed to adapt to changing market conditions. To capture and quantify these multi-timescale risks, the BSC framework can be modernized into a predictive AI-Augmented Scorecard.


By tracing value and failure through a bottom-up cause-and-effect chain, we’re able to transition from backward-looking financial reporting to real-time risk mitigation, as shown in Figure 2 below.


Figure 2: AI-Augmented Scorecard

Image created with ChatGPT (August 2026)

Mapping the Three Oversight Archetypes into the Scorecard

To operationalize this structure, firms need to map the European Commission's three models of human oversight directly into the scorecard's leading performance indicators: Human-in-the-Loop (HITL), Human-on-the-Loop (HOTL), and Human-in-Command (HIC).  (European Commission, 2019)


  1. Human-in-the-Loop (HITL)—the Air Canada archetype. HITL means a human reviews or approves every individual AI output before it takes effect. It's the right mode for acute, transaction-level risk, which was exactly Air Canada's failure, where one ungated chatbot response became a binding commitment with no review step at all.  (European Commission, 2019)

    -- Internal AI Operations Metric: Share of customer-facing generative outputs passed through a pre-execution review gate.

    -- Learning & Growth Metric: Total headcount certified to review and override automated commitments.


  2. Human-on-the-Loop (HOTL)—the Zillow archetype. HOTL shifts the human role to monitoring the system once it is in use, intervening when the system's behavior drifts from expectation rather than reviewing every individual output. This is the right mode for chronic, cumulative risk. Zillow's failure was invisible at the level of any single prediction and only became visible in aggregate, which is exactly the pattern HOTL is built to catch.  (European Commission, 2019)

    -- Internal AI Operations Metric: Live dashboard tracking of forecast-versus-actual variance against automated escalation thresholds.

    -- Learning & Growth Metric: Cross-functional staff trained to interpret statistical data drift and authorized to pause live models.


  3. Human-in-Command (HIC)—a Bank Lending archetype. Because HIC does not yet have a well-known scandal to illustrate it, consider a regional bank deciding whether to let AI approve consumer loans on its own. The model can underwrite at exceptional speed and scale, scoring thousands of applications against income, credit history, and repayment risk far faster than any human team. That makes HIC the right model for catastrophic or existential risk: it governs whether autonomous authority is granted at all and under what conditions it can be revoked. Before the model approves even one loan without a human signature, bank leadership—such as the Chief Risk Officer (CRO) or the board’s risk committee—must authorize deployment, define clear limits on loan size, applicant risk tier, and product type, and retain the authority to stop the system at any time, for any reason. The model cannot expand its own mandate, and once a human says stop, it must stop.

    -- Internal AI Operations Metric: Percentage of autonomous-decision authority grants (by dollar volume or risk tier) with a documented, time-bound reauthorization date, rather than standing indefinitely.

    -- Learning & Growth Metric: Number of board or executive-committee members formally trained and certified to exercise stop authority over a live AI system.

Final Word

The true value of AI doesn’t lie in total autonomy, but in augmented accountability. As tools like ChatGPT, Claude, Gemini, and Copilot help us bridge separate ideas into cohesive strategies, they also introduce unprecedented dimensions of risk—from acute conversational hallucinations (like Air Canada) to chronic algorithmic drift (like Zillow). Treating AI as a "set-and-forget" asset is a high-stakes gamble that modern organizations simply cannot afford to lose.


By modernizing Kaplan and Norton’s Balanced Scorecard into an AI-Augmented Scorecard, we gain more than just a metric calculator. We establish a rigorous, bottom-up governance framework that maps human oversight directly to risk timescales. Whether utilizing Human-in-the-Loop (HITL) for daily transactions, Human-on-the-Loop (HOTL) for systemic market shifts, or Human-in-Command (HIC) for catastrophic guardrails, meaningful human intervention remains the ultimate risk-mitigation tool.


To build sustainable, resilient organizations in the age of AI, we must look beyond the rearview mirror of financial damage. True leadership requires actively managing the foundational layers of people, process, technology, and governance. Don’t let your organization fall for the mirage of total autonomy; ensure your human commanders always hold the keys to the engine.

 

About Carl Burch

Carl Burch holds an MBA, CMA, CIA, FCCA, and is a QuickBooks ProAdvisor. He is also the co-founder of BURCH Business Services (BBS) located in Boston, MA.


You can contact Carl at carl.burch@burchbusinesservices.com, or for more information on how BBS can help you improve your business operations, visit www.burchbusinesservices.com 


References

European Commission High-Level Expert Group on Artificial Intelligence. (2019). Ethics guidelines for trustworthy AI. https://digital-strategy.ec.europa.eu/en/library/ethics-guidelines-trustworthy-ai


Kaplan, R. S., & Norton, D. P. (1992). The balanced scorecard—Measures that drive performance. Harvard Business Review, 70(1), 71–79. https://hbr.org/1992/01/the-balanced-scorecard-measures-that-drive-performance-2


Zillow Group, Inc. (2021, September 28). Zillow Group Reports Third Quarter 2021 Financial Results—Shares Plan to Wind Down Zillow Offers Operations. Zillow Group Investor Relations. https://investors.zillowgroup.com/news-and-events/news/news-details/2021/Zillow-Group-Reports-Third-Quarter-2021-Financial-Results--Shares-Plan-to-Wind-Down-Zillow-Offers-Operations/default.aspx



Rudin, C. (2019). Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nature Machine Intelligence, 1(5), 206–215. https://doi.org/10.1038/s42256-019-0048-x


Zillow Group, Inc. (2022). 2021 annual report (Form 10-K). U.S. Securities and Exchange Commission. https://www.sec.gov/Archives/edgar/data/1617640/000161764022000013/z-20211231.htm

 
 
 

Comments


Contact Us

Brighton, MA 02135
Tel: 571-389-1037
Email: admin@burchbusinesservices.com

  • LinkedIn
  • Facebook
  • Instagram
  • X

© 2023 by BURCH Business Services. All rights reserved.

bottom of page