Skip to content
Larnaca, Cyprus
BINA CYINNOVATION HUBLarnaca · est. 2026
AIAI23 August 20264 min read

Agents hit 100% on AGI-3; states step in on AI infrastructure

NVIDIA's AVO scores perfect on ARC-AGI-3, Pennsylvania gives communities a data-centre veto, and Varonis patches three Copilot flaws.

By BINA Editorial

Today's brief spans both ends of the capability spectrum — a landmark 100% benchmark score and a study showing frontier models still cannot think like scientists — alongside a US state ordering community consent before AI data centres can be built, a critical Copilot patch, and new labour-market data that puts numbers to a trend many workers already feel.

NVIDIA's AVO agent scores 100% on the ARC-AGI-3 benchmark

On 21 August, NVIDIA announced that its Agentic Variation Operators (AVO) system had completed all 183 levels across all 25 public environments of the ARC-AGI-3 benchmark, achieving a perfect 100.00 on the Relative Human Action Efficiency metric — the first agent to do so. AVO wraps an existing model, Claude Opus 5, inside a custom agent architecture; the underlying model alone scores roughly 30% on the same test. The system completed the benchmark in 6,624 actions, about 12% fewer than the nearest prior system, and was originally built to optimise CUDA kernels on NVIDIA hardware rather than to compete on benchmarks. ARC-AGI-3 gives agents no instructions, no rules, and no stated goals — the agent must infer what success looks like from the environment alone.

A new benchmark finds frontier LLMs still struggle with original scientific thinking

A paper published this month on arXiv introduces a benchmark called "Reconstruction," which tests whether language models can recover a research paper's central hypothesis from its bibliography alone — the kind of abductive reasoning a scientist exercises when reviewing a proposal with all citations intact but the core idea removed. The best-performing single models succeed only 3–15% of the time; a multi-agent Swiss-tournament pipeline raises the figure to 42%, still well below human expert level. The authors argue the deficit reveals a structural limit: current AI systems excel at pattern-matching across existing literature but struggle to generate genuinely novel hypotheses from incomplete evidence — a gap that matters as AI tools are increasingly marketed to support research.

Pennsylvania gives communities a veto over AI data centres

On 18 August, Pennsylvania Governor Josh Shapiro signed Executive Order 2026-05, requiring AI data centre developers to obtain legally binding community-approval agreements before the state will begin reviewing permit applications. Under the order's Responsible Infrastructure Development requirements, developers must fund all electricity infrastructure costs, source a material share of power from clean energy, and commit to local hiring and public transparency standards. The order removes AI data centre proposals from Pennsylvania's fast-track permitting process and bans non-disclosure agreements with local communities. "If the local community doesn't approve a project, the state won't approve it either," Shapiro said — a direct challenge to an industry that logged more than 100 data centre proposals in the state over the past year.

Researchers name and patch three vulnerabilities in Microsoft Copilot Personal

Varonis Threat Labs disclosed three vulnerabilities in Microsoft Copilot Personal — collectively named CoSnitch — patched on 18 August as part of Microsoft's monthly security update. The lead flaw, CVE-2026-24301, allows an attacker to craft a link that, when clicked, auto-runs a hidden Copilot prompt that silently pulls data from connected apps — calendar entries, emails, OneDrive files — and routes it to an attacker-controlled URL. No active exploitation was observed before the patch, and Microsoft's August 2026 Patch Tuesday addressed 415 vulnerabilities in total, including one exploited zero-day. CoSnitch illustrates a new class of prompt-injection risk: AI assistants that act on behalf of users become a lever for silent data exfiltration when their link-following and app-integration capabilities are abused.

AI's net employment effect turns negative at large firms, new data shows

An S&P Global analysis of private-sector employment data published this month finds that large companies — those with more than 250 employees — now report a net negative employment impact from AI investment at −13 percentage points, even as smaller firms still forecast positive net gains. Entry-level roles most exposed to automation have declined 13% since the rise of generative AI, according to Stanford data cited in the report. A companion report from the Center for Data Innovation cautions that aggregate statistics obscure wide variation by sector and task type, and calls for richer measurement frameworks before policymakers act on headline numbers. AI is simultaneously contracting some job categories and expanding others — but the balance for large employers has not yet reached neutrality.