netzstrategen
AI Operations

The end of AI magic: why companies need to know their human share

Published on 9/24/2026 · Sven Maier

Type a sentence, get a finished text. AI looks like magic. Anyone putting it into a business process cannot afford to believe in magic. Behind the surface sits no autonomous intelligence but organized human work, in model training and in daily operations. That share has had no name so far. We call it the human share.

Where do you stand?

Discuss your next step in a free diagnosis call. Book a slot →

Contents

The illusion of magic

The surface hides the construction. A prompt goes in, a usable result comes out. Everything in between stays invisible.

For private use that does not matter. For operations it does. Whoever puts a system into a business process is liable for its output. Liability without an understanding of the production chain is not a position. It is an open risk.

This holds twice over when AI takes on tasks people used to do. The work does not disappear. It moves somewhere else, usually somewhere nobody documents any more.

How AI actually gets built

Language models are trained in two phases. First, pretraining on large volumes of text. Then fine-tuning toward desired behavior. The second phase decides how a model behaves in everyday use, and it runs on human judgment.

The standard method is called reinforcement learning from human feedback, or RLHF. People are shown two model answers and pick the better one. Many such comparisons produce a reward model, which then retrains the language model.

How small that group can be is documented in the paper that made the method known. For InstructGPT, the direct predecessor of today’s chat models, OpenAI describes a team of about 40 contractors, hired on Upwork and through Scale AI (Source: Ouyang et al., OpenAI, 2022). The training sets held roughly 13,000 prompts for supervised fine-tuning and roughly 33,000 for the reward model (Source: Ouyang et al., OpenAI, 2022).

Three details from the same paper matter more for operations than the headline numbers.

The labelers often disagreed. Agreement among them ran at roughly 73 percent (Source: Ouyang et al., OpenAI, 2022). In a good quarter of cases there was no shared verdict on which answer was better.

Most comparisons were single-staffed. “Most comparisons are only labeled by 1 contractor for cost reasons,” the authors write themselves (Source: Ouyang et al., OpenAI, 2022). One judgment, one person, no four-eyes principle.

The yardstick did not come from the labelers. The researchers wrote the instructions: “we write the labeling instructions that labelers use as a guide,” states the section on whose preferences the model is actually aligned to (Source: Ouyang et al., OpenAI, 2022).

Four points where humans intervene From the training data to the released result Training data Model fine-tuning Configuration in operations Release of the result Curation and annotation Rating of answers Prompts, rules, limits Correction before it is sent Every grey box is part of the human share. Schematic illustration: netzstrategen · not based on survey data netzstrategen
The two boxes on the left sit with the provider. The two on the right sit in-house, and those are the ones that can be documented.
For presentations:

None of this is a criticism of the method. RLHF works, and it works efficiently: fine-tuning the 175-billion-parameter model took 60 petaflops/s-days against 3,640 for pretraining GPT-3 (Source: Ouyang et al., OpenAI, 2022). A small addition of compute, a large effect on behavior.

That is the point. A small, human-shaped step decides how the system behaves that now answers inside the company.

Why the magic narrative survives

The story of autonomous intelligence is useful. It attracts capital, it sells licenses, it spares awkward questions. The mechanical part appears in no keynote: data curation, labeling instructions, four-eyes principles.

Research takes a cooler view. A paper at the Conference on Fairness, Accountability and Transparency describes three distinct roles annotators can play in RLHF: they extend the judgment of the developers, they supply independent evidence, or they act as representatives of the wider public (Source: Coyne, FAccT ‘26, 2026). Which role applies is, in most pipelines, never stated.

The paper calls the consequence “responsibility laundering”: from the outside a pipeline sounds like broad societal participation, while inside it remains the decision of a few developers (Source: Coyne, FAccT ‘26, 2026). Responsibility is not shared. It just becomes harder to assign.

For companies that has an uncomfortable consequence. Plugging into a ready-made API does not just adopt a technology. It adopts decisions someone else made and did not disclose.

A company that does not know its human share does not know its system. It knows only the surface.

The counter-movement changes little here. Anthropic introduced Constitutional AI, a method that works without human labels identifying harmful outputs: “The only human oversight is provided through a list of rules or principles” (Source: Anthropic, Constitutional AI, 2022). The human share does not disappear in that setup. It moves forward, into the choice of principles.

Human share: a term we are introducing ourselves

One clarification belongs here: human share is not an established metric. No standard, study or norm defines it. We are introducing the term deliberately, because there is no name for something that occurs in every AI operation and is named nowhere.

Human share means the amount of human work needed to make an AI output usable. Not the training work at the provider, which can barely be quantified from outside. The work in your own operation: prompts someone sharpens. Results someone corrects. Cases someone filters out before they reach a customer.

Today that share counts as a flaw. As proof that automation is not finished yet. That reading points in the wrong direction. A high, known human share beats a low, unknown one. Known means controllable.

That human intervention does not vanish is also built into regulation. The EU AI Act requires data governance for high-risk systems, with training, validation and testing data examined for relevance, representativeness and freedom from errors (Source: Regulation (EU) 2024/1689, Art. 10). An examination that is not documented cannot be evidenced.

How closely data quality and AI output are linked is covered in the article on data quality. What applies to disclosure duties under Art. 50 is covered in AI content disclosure rules.

Three questions for tomorrow morning

Nobody can put an exact figure on the human share today. Neither can we. That is no reason to skip the question. It can be answered long before there is a metric.

  1. Where do people intervene in your AI output? Do you know, or do you assume? The difference between the two is the difference between operations and improvisation.
  2. Is the intervention routine or exception? Routine means the process is not finished. Exception means the process has a limit, and somebody knows it.
  3. Does the intervention still happen when that person leaves? If not, the system depends on a person. Not on a procedure.

Answering these three questions does not produce a metric. It produces control.

Which processes are affected in a specific case is something we sort out in a free diagnosis call.

What AI Operations does with this

AI Operations means running AI as a permanent business function, not as a project. For the human share that translates into four concrete commitments.

  • Keep provenance traceable. Which model, which version, which configuration produced a result? Without that record an error cannot be narrowed down.
  • Log the interventions. Every correction to an AI output is recorded as an event, not as silent handwork. The volume of those events produces the first defensible number.
  • Name the owner. One person per process, as for any machine in the plant. Without an addressee, responsibility spreads until nobody carries it.
  • Keep the know-how in-house. A company that understands its own processes can switch providers. One that does not, cannot.

These four points are not a software project. They are a question of organization, which is why we build systems with the team rather than for the team. The reasoning is in Built with the Team.

How decision rights and autonomy levels can be set in practice is covered in Governance, control, autonomy.

Conclusion: only documented work becomes measurable

AI is remarkable engineering. It is not magic. Between prompt and output sit decisions people made, partly at the provider, partly in-house.

The provider’s share stays largely hidden. The in-house share does not. It can be described, logged and owned. That is where AI Operations starts: handovers get documented, responsibility gets assigned, interventions stay traceable.

Scattered handwork turns into an operation that can be steered. Less spectacular than magic. It does, however, survive an audit.

Frequently asked questions about the human share

What is the human share?

The amount of human work needed to make an AI output usable, from data curation to the correction before it is sent. The term is not an established metric. We are introducing it because no name for this share existed.

How many people sit behind the training of a language model?

For most current models this is not public. It is documented for InstructGPT: OpenAI describes a team of about 40 contractors producing the demonstration and comparison data (Source: Ouyang et al., OpenAI, 2022). Comparable figures for newer models are not available.

Is a high human share a bad sign?

No. An unknown human share is the bad sign. A high, documented share shows where a process is not finished and where automation pays off next.

What does the EU AI Act say about this?

For high-risk systems, Art. 10 of Regulation (EU) 2024/1689 requires data governance, with training, validation and testing data examined for relevance, representativeness and freedom from errors. The duty applies regardless of whether personal data is involved.

How do we start measuring the human share?

With a single process. For one week, note every correction to an AI output: who, at which point, why. The result is usually surprising, and it is the basis for everything that follows. Where the starting point pays off most is something we sort out in a free diagnosis call.

Sources

How this article was produced

Human
  • Topic selection
  • Source selection
  • Fact-checking
  • Approval
AI
  • Research
  • Drafting
  • Diagrams
  • Publishing

This article was produced with AI support. Ideation, editorial planning, substantive review and approval rest with a human; copy-editing sits with the AI. Editorial responsibility is held by Sven Maier.

How we produce our content →

What's next

Self-Check

Assess AI potential

In 5 minutes: a concrete assessment of where the company stands with AI.

Start the Self-Check →
Newsletter

Digital Impact straight to your inbox

One sign-up, three newsletters: the AI Insights Newsletter every week with the latest insights articles, the Digital Impact Longread Newsletter and the Digital Impact Update once a month each. Double opt-in, unsubscribe anytime.

Podcast

AI Operations as a podcast

Experts including André Hellmann, Christina D'Ilio, Christian Sattel, Sarah Stock and regular guests from practice: all AI Operations topics as audio for on the go.