SecurityBrief Canada - Technology news for CISOs & cybersecurity decision-makers
Canada
AI agents fail to ace browsing tasks in Decodo test

AI agents fail to ace browsing tasks in Decodo test

Fri, 28th Aug 2026 (Today)
Joseph Gabriel Lagonsin
JOSEPH GABRIEL LAGONSIN News Editor

Decodo has published research ranking 45 AI agents across 10 browsing and task-based functions. No tool achieved a perfect score.

Claude for Chrome and Amazon Nova Act shared the top overall score with 18 points out of 20, followed by Fellou and ChatGPT Desktop on 17. The assessment compared vendor documentation with tests of how the tools handled practical tasks such as filling forms, completing purchases, switching tabs and carrying out multi-step work.

The results point to a gap between the range of functions AI agent vendors describe and how consistently those tools perform with limited supervision. The weakest area was transactional activity, where agents often managed parts of a purchase journey but fell short at the final step.

Transactional actions recorded an average score of 0.43 out of two, the lowest of any category in the study. Six tools received full marks for transactions: Amazon Nova Act, BrowserOS, ChatGPT Desktop, Minded, Sigma Browser and Skyvern.

Yet the ability to complete a payment flow did not always come with a documented stop point before an irreversible action. Only Amazon Nova Act and ChatGPT Desktop documented both successful purchase completion and a pause for user confirmation before a final commitment.

Safety divide

That created a split between agents that can carry out tasks and those that also define when human approval is required. Ten tools in the report documented a confirmation step before irreversible actions, including purchases.

Claude for Chrome was one example, with documented controls requiring user approval for purchases and financial actions. The study suggested this matters most when users allow agents to access payment details, accounts or other sensitive information.

Multi-step automation was another area of concern. Of the 35 tools that documented multi-step task completion, 51% did not document any confirmation step before irreversible actions.

The issue became more pronounced when background execution was added. Seventeen tools could run without an active user session, and seven of those combined that feature with no documented confirmation step, according to the findings. The group included Airtop, Axiom AI and BrowserOS.

Browser advantage

Cross-tab awareness, a measure of whether an agent can maintain context while moving between browser tabs, scored an average of 0.66 out of two. Browser-native products performed best in that category.

Claude for Chrome, ChatGPT Chrome Extension, Chrome Gemini, Sigma Browser, Opera Aria, Opera Neon and Brave Leo all achieved full marks for cross-tab work. Standalone and cloud-based agents generally scored lower, suggesting a structural advantage for tools built directly into the browser.

The methodology had two stages. First, each product was scored on how clearly the vendor documented each function in public material such as pricing pages, product documentation and changelogs. Researchers then tested selected claims through real workflows, logging problems including crashes, freezes, timeouts, lost sessions and unsuccessful task completion.

The 10 categories assessed were form filling, transactional actions, multi-step task completion, cross-tab awareness, session memory, persistent memory, cross-site execution, confirmation before irreversible actions, scheduled or background tasks and third-party integrations.

Business risk

The findings come as companies weigh whether AI agents are ready to take on more independent work in operations, procurement and research. While the strongest performers posted relatively high scores, the report argued that a strong overall mark does not necessarily make a tool suitable for every business process.

Research from Anthropic, OpenAI and Brave has already drawn attention to prompt injection risks, where hidden text in webpages or emails can influence browser agents. For tools able to buy goods, submit forms or access private data, weak supervision or confirmation can raise the consequences of those attacks.

That leaves businesses with a selection problem as well as a safety problem. Some may value form completion and reliability for internal administrative work, while others may need stronger controls around checkout, approvals or work spread across multiple sites and tabs.

Gabriele Vitke, Product Marketing Team Lead at Decodo, said: "Match the tool to the job, not to the longest feature list. Someone on your team will still have to own what happens when the agent gets it wrong. An agent buying the wrong inventory or emailing the wrong client won't announce itself, and a tool that fails and needs constant babysitting can ultimately cost more in staff time than a narrower tool that does one thing reliably."