Friday, 14 August 2026
Your paper, your pace.
Technology

AI and Your Job: The Five Levels That Actually Matter

AI is already doing real slices of skilled work, from hospital paperwork to heart-attack screening. How much control it is given over a task, not the technology itself, decides what changes for you.

7 min 4 sources Confidence 94/100

In short

What happened. Researchers now grade AI systems by how much control they hold over a task — from “tool,” where a human does everything, up to “agent,” where the system acts alone — rather than by a single yes-or-no verdict on whether AI can “do the job.”

What it means. In real workplaces today, AI already sits at different levels for different tasks inside the same job: drafting a doctor’s notes is closer to “collaborator,” while reading a cardiac scan is closer to “expert,” and almost nothing in a skilled job has reached full “agent” autonomy yet.

Risks and impact. The risk is not a single dramatic layoff moment. It is tasks quietly sliding up the scale, one at a time, while the people doing them may not notice which parts moved.

What can be done. Break your own job into its component tasks and ask which level each one occupies today — that tells you more than any general prediction about “AI taking jobs.”

What to watch. Whether a task in your field moves from “consultant” (you still decide) to “expert” (its output gets acted on without a check) — that shift, not a headline, is the one that changes what you do all day.

Shown as a summary because of your reading settings.

What happened

A doctor finishing a twelve-hour shift used to spend the last hour typing up notes. Increasingly, an AI tool listens to the consultation, drafts the note, and the doctor edits and signs it. That is not science fiction — generative AI tools are already used this way to transcribe consultations and draft medical records, and early studies suggest they reduce the administrative burden that is a known driver of burnout among clinicians.

That single example holds the whole story of what AI is currently doing to jobs. It is not replacing the doctor. It is taking over one task inside the job and leaving the doctor to check the result.

In 2023, researchers at Google DeepMind proposed a framework for grading exactly this kind of change. Instead of asking whether an AI system is “intelligent” in general, they graded its autonomy over a specific task on a five-step scale: tool (fully under human control), consultant, collaborator, expert, and agent (fully autonomous). The same underlying software can sit at different levels for different tasks, and companies including OpenAI, Google, xAI and Meta are explicitly working toward systems with much higher scores across many more tasks — what researchers call artificial general intelligence, or AGI. A 2020 survey counted 72 active AGI research projects across 37 countries.

What the evidence supports

Some claims here are well supported by direct comparison. Wearables, smartphones and AI algorithms have shown they can flag patients at risk of coronary artery disease, and a 2019 study found AI predicting heart attacks with up to 90% accuracy; other research found AI performed as well as trained humans at interpreting cardiac echocardiograms, and in some emergency-room settings diagnosed heart attacks better than physicians, cutting both missed diagnoses and unnecessary tests. That is evidence of AI genuinely operating at “expert” level for a narrow, well-defined task.

Other claims are much thinner than the headlines around them suggest. A widely cited 2023 study found that people preferred ChatGPT’s answers to physician answers in 78.6% of 585 evaluations on the online forum r/AskDocs, rating them more thorough and more empathetic. But the questions were pulled from an open forum rather than an actual doctor-patient relationship, the answers were never checked for medical accuracy, and critics have pointed out that the people scoring the answers were the study’s own coauthors. That is a real finding about tone, not proof that AI gives better medical advice.

What we still do not know is how far this generalizes beyond narrow, data-rich tasks like cardiac imaging. A study from the Centerstone research institute found predictive models built from electronic health records reached only 70–72% accuracy forecasting how an individual patient would respond to treatment — useful, but far from a replacement for clinical judgment.

How the story is being framed

The alarmed view says AI is about to take over skilled work wholesale. History gives real reason for caution here: in 1965, AI pioneer Herbert Simon predicted machines would be capable of any work a human can do “within twenty years.” In 1967, Marvin Minsky said the problem of creating artificial intelligence would “substantially be solved” within a generation. Neither happened on schedule, and the field went through repeated cycles of hype followed by funding collapse, known as AI winters. A prediction about pace deserves the same scrutiny now as it did then.

The dismissive view says none of this really changes anything, that AI is just a tool like any other. That undersells what is already measurable: real clinical documentation is being drafted by AI today, real triage decisions are being informed by AI risk scores today, and by DeepMind’s own classification, systems like current large language models already qualify as “emerging” AGI — comparable to an unskilled adult across a wide range of tasks, not zero capability.

A third view, closer to the evidence, says the honest unit of change is not “the job” but the task. A radiologist’s job includes reading scans, writing reports, talking to anxious patients and coordinating care; each of those tasks can move independently, and at different speeds, along the tool-to-agent scale. This view is less dramatic than either of the others, but it is the one the data actually supports.

The background

The tool-to-agent scale is worth understanding in plain terms, because it explains why two people can look at the same AI system and disagree completely about how big a deal it is.

At the “tool” level, a human does the work and the software only assists mechanically — spellcheck is the oldest example. At “consultant,” the AI offers an opinion or a draft, but a human reviews it before anything happens; an AI flagging a patient as high-risk for coronary disease, which a doctor then double-checks, sits here. At “collaborator,” human and AI go back and forth, each shaping the output — a doctor editing an AI-drafted note is an example. At “expert,” the AI’s output is trusted enough to be acted on with only light or no human review, which is closer to where AI-read echocardiograms already sit. At “agent,” the system acts entirely on its own.

Very little skilled work has reached “agent” level yet, for a concrete reason: the newest generation of AI systems, called reasoning models, emerged only in 2024 and can produce convincing but false answers — a failure mode researchers call hallucination, which older rule-based systems did not share. A system that can be confidently wrong is a poor candidate for full autonomy over anything consequential, which is exactly why most real deployments today keep a human in the loop, at consultant or collaborator level, rather than agent level.

Where AI has been tested at something closer to full autonomy, it has mostly been in narrow physical or financial tasks designed as public benchmarks rather than real jobs: a 2025 University of Edinburgh system called ELLMER autonomously made coffee in a real kitchen, adapting to obstacles in real time, and entrepreneur Mustafa Suleyman has proposed a benchmark test of giving an AI $100,000 and asking it to turn that into $1 million with no human help. Nobody has publicly reported an AI passing that second test.

Who it touches

For a hospital documentation clerk, the arrival of AI-drafted notes is already a change in what the job involves, even without a layoff — the work shifts from writing to checking. For a cardiologist, an AI flag on a borderline scan is a second opinion that changes the odds of catching a real problem, not a replacement for the decision about what to do with it. For a patient posting a symptom question online at midnight, an empathetic-sounding AI answer can feel more attentive than a rushed clinical visit — which is exactly why the gap between “sounds better” and “is medically correct” matters most to the person who has no way to tell the difference.

The deeper story

The question people actually ask — “will AI take my job?” — is really asking something narrower and more useful: which of the things I do all day still needs a human to decide, and which ones have quietly stopped needing that?

That reframing matters because autonomy is not a light switch. It is granted, task by task, usually by an employer or a regulator deciding how much checking is still required. A hospital does not wake up one day and hand a diagnosis to software; it moves one task at a time from consultant to expert, as evidence accumulates that the software is reliable enough to trust with less supervision. The same will likely be true in law, writing, customer service and skilled trades — not a single verdict, but hundreds of small, mostly invisible decisions about how much checking a given task still deserves.

What tends to resist that shift longest is not intelligence in the abstract. It is accountability — the fact that when a decision goes wrong, someone specific has to own it, explain it, and live with the consequences of having made it. A model can draft an empathetic-sounding reply to a frightened patient; it cannot be the person who sits with that patient afterward if the answer turns out to be wrong.

Something to sit with

Which tasks in your own work would you trust an AI to do at “collaborator” level — and which ones would you never move past “consultant,” no matter how good the software got?

If a task you do moved from consultant to expert level tomorrow, what would you actually do with the extra time?

Sources

We report facts from the sources above in our own words and link to the originals. Interpretation is ours, not theirs.

QUICK UNDERSTANDING CHECK

According to the Google DeepMind autonomy framework, what separates a "consultant"-level AI from an "expert"-level one at the same task?

♻︎ Free to republish

Copy this HTML into your CMS. Credit line and licence are included. Republish our work — free

Every headline has a deeper story. This is ours.

What we are doing here

One good piece of thinking a day

The day's most worthwhile story, and the question underneath it. No spam, one click to leave.

One click to leave. We never sell or share your address.