By clicking “Accept”, you agree to the storing of cookies on your device to enhance site navigation, analyze site usage, and assist in our marketing efforts. View our Privacy Policy for more information.
September 27, 2026

AI Regulation Around the World: Why Human Oversight Is Moving to the Centre of the Debate

September 27, 2026

AI Regulation Around the World: Why Human Oversight Is Moving to the Centre of the Debate

Artificial intelligence regulation is entering a new phase.

For several years, much of the discussion around AI governance focused on principles: fairness, transparency, safety, accountability and responsible innovation.

Today, those principles are increasingly being translated into laws, technical standards, testing requirements and operational controls.

But there is no single global regulatory model.

The European Union is implementing a horizontal, risk-based legal framework through the EU AI Act. The United States continues to debate how far federal regulation should go, while states are also developing their own rules. Other regions are adopting different combinations of sector regulation, technical standards and governance frameworks.

The underlying question, however, is becoming increasingly similar across jurisdictions:

How can organisations demonstrate that an AI system is not only powerful, but also reliable, controllable and appropriately supervised?

That question becomes even more important as AI moves from generating content to taking actions, supporting high-impact decisions and operating through increasingly autonomous agents.

A global debate between innovation and control

AI regulation has always involved a trade-off.

Regulate too early or too rigidly, and governments risk constraining innovation in a technology that is evolving extremely quickly.

Regulate too little, and companies, citizens and public authorities may be exposed to systems whose behaviour is difficult to understand, predict or control.

In the United States, this tension is particularly visible.

Part of the debate is focused on maintaining an innovation-first approach and avoiding excessive constraints on AI development. At the same time, there are growing discussions around stronger safeguards, independent evaluations, incident reporting and clearer responsibilities for developers of advanced AI systems.

The debate also has a strong federal-versus-state dimension, with a growing number of local initiatives and continued discussions around the role of a national framework.

Behind these discussions is another powerful factor: geopolitics.

AI leadership is now viewed as a strategic issue, and regulation is increasingly being discussed not only in terms of safety, but also competitiveness, sovereignty and technological leadership.

Europe has chosen a clearer regulatory path

Europe has taken the clearest legislative route with the Artificial Intelligence Act.

The EU AI Act is based on a risk-based approach, with different obligations depending on how AI systems are used and the level of risk they may create.

For companies, the important point is that the Act is not only about principles.

It creates operational expectations around areas such as:

  • risk management;
  • data governance;
  • traceability and record keeping;
  • transparency;
  • accuracy and robustness;
  • human oversight;
  • post-market monitoring.

For isahit, this creates a particularly relevant field around Human-in-the-Loop.

The working assumption behind our current thinking is simple:

the AI Act does not say “use a Human-in-the-Loop provider”, but it creates exactly the type of operational need that Human-in-the-Loop can help address.

Human oversight is more than having a person “in the loop”

One of the most relevant parts of the AI Act for our field is Article 14 on human oversight.

For high-risk AI systems, the regulation requires that they can be effectively supervised by natural persons.

That word — effectively — matters.

Human oversight should not simply mean putting an employee at the end of an automated process.

The people responsible for supervision should, depending on the context, be able to:

  • understand the system’s capabilities and limitations;
  • detect anomalies or unexpected performance;
  • remain aware of the risk of automation bias;
  • interpret AI outputs correctly;
  • decide not to follow an AI recommendation;
  • override or reverse a decision;
  • intervene or stop the system where necessary.

This changes the question.

Instead of asking:

“Is there a human in the loop?”

organisations may increasingly need to ask:

“Does the human loop actually work?”

For example:

Can an operator recognise an incorrect recommendation?

How quickly?

Does the operator understand when the AI may be unreliable?

Will the operator override an incorrect decision — or simply trust the machine?

Can the process be interrupted safely?

Does performance change depending on the operator’s training or expertise?

These are not abstract governance questions.

They are testable questions.

From compliance documentation to measurable evidence

This is where human evaluation becomes particularly relevant.

A written procedure explaining that an AI system is supervised is useful, but it does not necessarily demonstrate whether the supervision works in practice.

Likewise, saying that an AI system is “accurate” is less useful than being able to show how it behaves across hundreds or thousands of representative scenarios.

The document we have been working on internally identifies a broader operational layer around several AI Act requirements:

  • Article 9 — Risk management: edge-case and adversarial testing;
  • Article 10 — Data governance: human QA and dataset validation;
  • Article 12 — Record keeping: audit trails and evidence logs;
  • Article 13 — Transparency: human evaluation reports;
  • Article 14 — Human oversight: human oversight testing;
  • Article 15 — Accuracy and robustness: human benchmarking and QA;
  • Article 72 — Post-market monitoring: continuous human monitoring.

Taken together, these requirements point toward an important shift:

AI governance increasingly needs evidence.

Testing the whole system, not only the model

One of the most interesting ideas is to evaluate not only the AI model, but the complete operating system around it:

AI + interface + human operator + procedure.

For example, a controlled test could introduce incorrect or ambiguous AI outputs into a workflow.

The evaluation could then measure:

  • how many errors are detected by the operator;
  • how many are correctly corrected;
  • average detection time;
  • inappropriate acceptance of incorrect AI decisions;
  • override rate;
  • ability to interrupt the process;
  • performance differences depending on training or expertise.

This creates a form of measurable Human Oversight Effectiveness.

And this is fundamentally different from simply documenting that “a human can intervene”.

Human-in-the-Loop is evolving

Historically, Human-in-the-Loop services were mostly used upstream:

Human → Data → Train AI

Then, with RLHF and model evaluation:

Human → Fine-tune / Evaluate → Improve AI

The next step may increasingly be:

Human → Evaluate → Supervise → Produce evidence → Monitor AI

In other words, the human role is moving from being mostly before the model to increasingly being around and after the model.

This is particularly relevant for:

  • generative AI;
  • AI agents;
  • decision-support systems;
  • computer vision;
  • robotics and Physical AI;
  • regulated or sensitive enterprise use cases.

Where isahit is exploring a new role

At isahit, this is a natural extension of work we already perform across:

  • data annotation;
  • QA;
  • computer vision;
  • RLHF;
  • agent evaluation;
  • human review.

The key point is that we do not see isahit as an AI Act certification company.

The position we are exploring is different:

“We provide the human testing, evaluation and evidence layer supporting AI Act compliance.”

The customer, legal adviser, Responsible AI team or specialised partner remains responsible for regulatory qualification and overall compliance.

isahit can provide the operational layer:

tests → evaluators → results → metrics → logs → reports

Four possible building blocks

The model we are currently exploring can be summarised around four pillars:

Evaluate

Structured human evaluation of AI outputs, decisions and behaviours.

This could include:

  • relevance;
  • factual accuracy;
  • hallucination;
  • instruction following;
  • safety;
  • bias;
  • consistency;
  • quality of escalation to humans.

Red Team

Human-led testing of difficult, adversarial or unexpected scenarios.

Examples include:

  • edge cases;
  • ambiguous instructions;
  • jailbreak attempts;
  • conflicting instructions;
  • unexpected situations;
  • prohibited actions;
  • unsafe behaviour.

Oversee

Testing whether the human oversight process actually works.

The focus is not only the AI system.

It is the combination:

AI + human + interface + procedure.

Monitor

Continuous human review of a sample of AI interactions in production.

For example:

1% of live interactions → human review → metrics → alerts if a threshold is exceeded.

This is particularly important because AI assurance should not stop after deployment.

Expertise matters as much as scale

Another major implication is that not every AI system can be reviewed by the same population.

A general evaluator may be suitable for:

  • language quality;
  • instruction following;
  • basic hallucination detection;
  • A/B comparisons.

But higher-risk or more specialised use cases may require:

  • engineers;
  • legal experts;
  • finance professionals;
  • HR specialists;
  • cybersecurity experts;
  • automotive experts;
  • aerospace specialists;
  • domain-specific adjudicators.

This suggests that the future of Human-in-the-Loop may rely increasingly on qualified evaluator pools.

The internal concept we are exploring is an Evaluator Passport, combining:

  • identity;
  • languages;
  • country;
  • qualifications;
  • professional background;
  • sector expertise;
  • training;
  • calibration scores;
  • quality history.

This turns human evaluation into a more structured and auditable capability.

Working with partners will be essential

AI Act implementation is multidisciplinary.

It involves:

  • legal interpretation;
  • AI governance;
  • risk management;
  • cybersecurity;
  • technical architecture;
  • audit;
  • operational testing.

We therefore do not believe that isahit should become a general AI Act consulting firm.

A more relevant model is to work in synergy with specialised partners.

The division of roles can be simple.

A partner can manage:

  • regulatory qualification;
  • AI system mapping;
  • gap analysis;
  • governance;
  • definition of control requirements.

isahit can manage:

  • test protocol creation;
  • evaluator sourcing;
  • human evaluation;
  • human oversight testing;
  • red teaming;
  • bias testing;
  • double validation;
  • metric production;
  • audit trails;
  • recurring human monitoring.

This creates a complementary model.

The partner defines what needs to be controlled.

isahit helps execute how it is tested with humans and documented.

From concept to pilots

The objective now is to move from theory to real projects.

Our current roadmap is to begin with a small number of pilots in October 2026, ideally across different types of AI systems.

The aim is to test the methodology, evaluation criteria, evaluator qualification, reporting format and operational workflow in real conditions.

The initial pilots are expected to build on existing isahit capabilities:

  • isahit.lab;
  • Smart Forms;
  • QA workflows;
  • workforce management;
  • human evaluation;
  • structured reporting.

The idea is not to build a large AI Act platform first.

It is to start with practical use cases and then progressively industrialise what proves useful.

Toward a broader offer by the end of 2026

Based on these pilots, our objective is to progressively structure a broader offer by the end of 2026, in synergy with specialised Responsible AI, AI governance and regulatory partners.

The purpose will not be to add another abstract layer of AI governance.

It will be to translate governance principles into something operational:

tests, human judgements, measurable results, monitoring and documented evidence.

The core position can be summarised in one sentence:

isahit provides the human layer to test, validate and continuously monitor enterprise AI systems — turning human evaluation into measurable, auditable evidence.

As AI systems become more capable and more autonomous, organisations will need better ways to demonstrate that these systems can be challenged, supervised and controlled.

That may become one of the next major roles of Human-in-the-Loop.

Not only helping AI learn.

But helping organisations trust, govern and monitor AI in the real world.

Stay tuned — more details, pilots and news to follow in the coming weeks and months.

You might also like
this new related posts

Want to scale up your data labeling projects
and do it ethically? 

We have a wide range of solutions and tools that will help you train your algorithms. Click below to learn more!