Laava LogoLaava
Back to news
News & analysis

AI that excludes an employee does not automate the work

A field study with blind screen-reader users shows that computer-use agents can complete useful tasks while also taking control of the work away from the user. Accessibility should therefore be an acceptance criterion for the entire AI workflow: from instruction and progress to correction, confirmation and recovery.

Why this matters

News only becomes relevant when you can translate what it means for process, risk, investment, and decision-making in your own organization.

Accessibility is not a box to tick at the end of implementation

An AI agent can technically perform a task while excluding the employee who is expected to work with it. This happens when someone can issue an instruction but cannot independently follow what the agent is doing, correct an error, stop an important action or recover after a failure.

The work has not truly been automated. Operation and control have shifted to a colleague who can oversee the visual interface. The organisation has created a new dependency exactly where AI was supposed to deliver autonomy and scalability.

Accessibility therefore belongs beyond screen design. In operational AI, it is a property of the entire workflow: issuing an instruction, understanding context, monitoring progress, confirming choices, verifying output and recovering safely.

A field study shows where computer-use agents break down

A new study of computer-use agents for blind screen-reader users followed eight participants during three weeks of everyday computer use. Through an accessible interaction layer, they issued 1,258 commands across twelve Windows applications. The researchers also replayed the same commands with four other agent models and used human annotators to determine whether each task was fully completed, partially completed or failed.

The strongest model fully completed 52.5% of the attempted commands. Partial completion was common across every model studied. That matters operationally: a nearly finished task can still be useless or risky when a formula misses one range, a document is not finally saved or a setting is changed only halfway.

The failures were not limited to language understanding. Agents invented controls that were absent from the current interface, failed to discover hidden menu paths, dropped requirements across multiple steps and sometimes failed to recognise that a task was complete or stuck. Simple actions with visible controls worked better than tasks involving multiple constraints, applications or hard-to-find options.

These percentages are not a general score for office automation. The study covered eight adult screen-reader users, English-language commands, five models and Windows desktop applications; web applications, other operating systems and other accessibility needs were outside its scope. Participants also sometimes avoided sensitive tasks or task types that had repeatedly failed before. Those limits sharpen the useful lesson: test the actual work with the people, assistive technologies, language and applications for which the workflow is being built.

Users do not always want the agent to take over

A second finding is at least as relevant to businesses. Participants did not view the agent only as an autonomous executor. They also wanted explanations of what was on screen, situated help when they were stuck, guidance through an unfamiliar process and the ability to inspect, modify or stop a planned action.

That changes the design question. Do not ask how quickly the human can be removed from the loop. Ask which form of support gives this employee the greatest independence at this step. Sometimes that is execution. Sometimes it is a description, a suggested next step, a verifiable proposal or recovery support.

Full takeover may sound attractive, but it creates dependence when the agent provides no understandable status or can only be corrected visually. A well-designed agent expands a user's ability to act. It does not place a second inaccessible interface powered by AI in front of the first one.

Measure independent control, not only task completion

Organisations commonly measure pilots through accuracy, speed and cost per task. Those metrics miss an important part of the business outcome when not every employee can operate the workflow independently.

Add a concrete measure: the share of tasks the intended user can independently start, monitor, verify, correct and complete. Also track how many handoffs to a colleague are required solely because of the AI system's interface. That makes accessibility visible as operational capacity rather than an abstract quality ambition.

This aligns with W3C guidance on involving users with disabilities in evaluation. Automated checks and standards remain necessary, but real users identify problems that technical checks do not predict. That is even more important for AI workflows, where behaviour depends on context, task sequence, interface changes and recovery paths.

Five acceptance criteria for an accessible AI workflow

Make accessibility testable before the pilot with five questions:

  1. Instruction: can the user enter and review the task, conditions and exceptions through their own assistive technology?
  2. Progress: does the agent make clear, without relying on vision, what it is doing, which step is complete and where it is uncertain?
  3. Control: can the user inspect a proposal, correct assumptions and approve, change or stop an action before it becomes irreversible?
  4. Outcome: can the resulting state be verified in the source system, rather than relying only on a message that the agent is done?
  5. Recovery: after a failure, can the user independently go back, retry, switch to an accessible alternative or request focused help without losing context?

Do not test these points only on the happy path. Use real documents, multi-step tasks, unexpected dialogs, changed labels and work that moves between applications. Let employees using the relevant assistive technologies determine whether a handoff works. Their task completion is the criterion, not the impression of a sighted tester watching along.

Practical adoption starts with the people who must do the work

Laava builds operational AI around existing systems, human control and explicit failure paths. The study adds a necessary design boundary: that control is only real when every intended user can exercise it.

Start with one bounded workflow and map not only data, actions and exceptions, but also the different ways employees interact with the work. Treat status, confirmation and recovery as first-class interactions. Test with the people who will use the system in practice before making higher automation rates the goal.

An agent that completes more tasks while making part of the team dependent does not deliver scalable execution. Good operational AI removes manual work while allowing people to retain independent control over what happens on their behalf.

Translate this to your operation

Determine where this affects you first for real

The practical question is not whether this news is interesting, but where it directly changes your process, tooling, risk, or commercial approach.

Related Laava approach: Production AI agents

First serious step

From news to a concrete first route

Use market developments as context, but make decisions based on your own operation, systems, and risk trade-offs.

You leave with a clear view of the first workflow, the key dependencies, and the right next step.

Included in the first conversation

Assess operational impactSeparate relevant risks from noiseDefine the first route
Start with the bottleneck. Build from there.
AI that excludes an employee does not automate the work | Laava News