Executive Leadership

AI Literacy for Board Directors: Supervise One Agent for a Quarter

Most AI education offered to boards is a deck. Someone walks the directors through what a model is, runs a demo that works, and leaves forty minutes for questions. Everyone walks out feeling caught up, and almost nobody is.

October 11, 20266 min read

I have sat through a few of these sessions from the director’s chair at ACT | The App Association. I have also been on the other end of the question, as CEO of Visiting Media, where I run a fleet of specialized agents across my own operating work. The two seats give me different complaints about the same problem. From the board side, the sessions are pleasant and forgettable. From the operator side, I can see exactly what they leave out, because it is the part of running agents that cost me the most sleep.

What the literacy guides agree on

The pages that rank for board AI literacy line up well. Directors do not need to be technologists, they say, any more than they need to be accountants to read a balance sheet. One of them builds its case on an Australian court decision from 2011 that turned financial literacy into a baseline duty for directors, and argues AI is heading the same way. Most offer a short list of things a literate director understands: how the tools fail, where the company already uses them, which risks are specific to AI, and when to stop relying on the output. The training advice is similar everywhere. Get briefings from management, bring in outside educators, maybe add a director with AI depth, and put AI on the skills matrix.

I would sign almost all of it.

The trouble is that every recommendation is about asking. None of the guides suggest a director should ever supervise an agent, even a small one, even for a few weeks. That gap matters more than it looks, because the questions a director asks are only as good as their picture of what failure feels like from inside. You can memorize the phrase “how the tools fail” in a morning. Recognizing it in the wild takes longer.

Demos teach the wrong lesson

A demo is a success story by construction. The presenter picked the prompt, rehearsed it, and stopped before anything drifted. Directors leave believing that when an agent fails, it fails loudly, with an error or an obviously wrong answer someone will catch.

The failures that hurt me were quiet. An agent of mine published a fabricated claim at 2:14 one morning, and the claim read exactly like the true ones around it. No error. Nothing in the tone gave it away either. Another agent, one I had replaced in the spring, was still waking up on its old schedule in late summer with a token nobody had revoked. I built that system. I still missed it for months.

A director who has only seen demos will hear “we have human review” and relax. A director who has watched a confident wrong answer slide past their own eyes will ask how many outputs the reviewer actually reads per day, and what happens on the days the reviewer is out.

Supervise one agent for a quarter

My proposal is modest. Each director picks one small agent and supervises it for ninety days. Keep it away from board materials and anything confidential. A weekly brief on the company’s competitors, built on the company’s approved tools, is about right. The point is to own something that does real work on a schedule, so that its mistakes land on you.

The rules I would set are the ones I hold my own fleet to. Write down, before it runs, what the agent may and may not do. Read every output for the first two weeks. After that, pull a sample each week and grade it, the same weekly graded sample I run across my own agents, and keep a short log of what was wrong. At the end of the quarter, bring one page to the board: where it slipped, and how many minutes a week supervision actually took.

That last number is the one I care about. Every AI proposal a board approves assumes some amount of human oversight, and almost nobody on the board has a felt sense of what an hour of it buys.

What the quarter teaches

  • Drift is gradual and polite. Output gets slightly worse over weeks, never all at once, and you only notice because you graded week two and can compare. This is the lesson that changes how a director reads every quality metric management reports afterward.
  • Your instructions were worse than you thought. Most directors will rewrite their agent’s brief at least once, and learn that “the AI got it wrong” often means the person who set it up was vague.
  • Checking is tedious. Somewhere around week five, skipping the sample starts to feel reasonable.
  • Turning it off has steps. Retiring even a toy agent means finding where its credentials live, which is a small taste of retiring an agent in production.

The first item is worth the whole exercise. The third is the one that makes directors honest about how oversight really works inside a company with a deadline.

Embarrassment is the curriculum

A director who has never been embarrassed by an agent will believe the dashboard.

Ask for raw output once a year

Supervising your own agent is half of it. The other half is looking at the company’s. Once a year I would ask management for ten unedited outputs from a production agent, chosen by someone other than the agent’s owner, with the grader’s notes attached. Read them in the meeting packet as you would read a sample of customer contracts.

Directors who have spent a quarter grading their own agent read those ten outputs very differently. They spot the confident sentence with no source behind it, and the grader’s notes that are suspiciously short for a page that long. They also read the AI page in the board pack with more suspicion, which is healthy, since more of it is drafted by agents every quarter.

What I would not do

Don’t solve this by recruiting one AI expert to the board and letting everyone else off the hook. A technology board advisor or a director with deep AI experience is useful, and I have argued for that seat. A board where one person understands agents and seven defer to them has concentrated a risk it was supposed to spread. The same concentration problem shows up on the management side, which is why it sits on my list of what boards should ask about AI.

I would also skip certificates. A course completion badge tells the nominating committee that a director sat through something, which is exactly the deck problem again with a logo on it.

Where to start

At the next board evaluation, add one line to the skills matrix: has supervised an agent for at least one quarter, yes or no. Don’t grade it beyond that. The line makes the gap visible, and directors who see a “no” next to their own name tend to fix it before the next cycle.

The CEO should offer to help set the agents up, and then step back, because a director supervising an agent the CEO quietly maintains has learned nothing about supervision. Expect the first month to be clumsy. My own delegation matrix was wrong within six weeks of writing it, and I had been running agents for a long time by then. A director’s first brief will be wrong sooner, and that is the most useful thing that can happen to them all year. It is also the moment they start asking management better questions about the AI risk appetite statement they are about to vote on.

This article is part of the Executive Leadership cluster, focused on board governance and the operating discipline required to run AI systems responsibly at the executive level.

Related Reading