The full anxieties, limits, and capabilities of AI — in one place.
Filter by type, sort by trajectory, search by keyword. The same 41 entries from the three companion maps below, made browsable for workshop participants.
Curated by Fred Pelard · fredpelard.com
Three Patterns Worth Naming
1. The expert/public gap. The biggest divergence isn't on whether AI is a problem — it's on which problems matter. On the last survey to ask both groups the same questions (Pew, 2025), the public worried most about jobs (56% vs 25% for experts) and human connection (57% vs 37%), while experts worried most about misinformation (70% vs 66%) and bias (60% vs 49%). Two years on, the public has closed much of that gap from its own side: 71% now expect AI to mean fewer jobs, and 79% expect it to reduce them over the next decade.
2. The new, tangible fears are winning. Water consumption did not exist as a public concern five years ago; cognitive atrophy did not three years ago; and the newest arrival, household electricity bills, took opposition to a nearby data centre from 42% to 75% in eleven months. These win the attention war over alignment and consciousness because they are local, visible, and storyable in a way long-term existential risk isn't.
3. Fading doesn't mean resolved. Algorithmic bias is the cleanest example, and in 2026 the regulation went with the attention: the EEOC dropped algorithmic tools from its enforcement plan, the EU deferred its high-risk obligations to December 2027, and Colorado repealed its AI Act outright. None of the underlying metrics improved. Watch where the headlines aren't.
Three Patterns Worth Naming
1. The "tools, not models" rule. The weaknesses that have been mitigated fastest — Recent, Mathematical, much of Real — were solved less by smarter models than by giving models tools: web search, code execution, retrieval. The ones that remain stuck are those tools cannot reach, and the two added this September are architectural rather than incidental: prompt injection, because a model reads instructions and data as one stream, and reward hacking, because it optimises what you measured rather than what you meant.
2. Sycophancy is the surprise. The only weakness that got measurably worse between 2023 and 2025 before reversing. The cause: training models to be liked makes them obsequious. The April 2025 GPT-4o rollback marked the moment frontier labs started treating it as a safety issue, not a UX preference. The correction is real but uneven: the best 2026 scores are reached partly by declining to answer, on as many as four questions in five.
3. The two-tier reality, now splitting three ways. Confidential, Real and Reliance are largely managed for users on enterprise tiers with the right configuration — and largely unmitigated for everyone else. As of August 2026 the vendors no longer agree on what the enterprise tier means: OpenAI is piloting zero data retention while Anthropic mandates thirty days on its most capable models. The standard caution slide should still assume the consumer-tier reality, which is the prudent default.
Three Patterns Worth Naming
1. The reasoning regime change. Three of the ten strengths (mathematical, structured reasoning, code generation) all owe their 2024–25 acceleration to one thing: RLVR — Reinforcement Learning with Verifiable Rewards. Same architecture, much longer training runs against tasks where the answer can be machine-checked. The biggest training-paradigm shift since RLHF.
2. The capability/trust gap. Strengths and weaknesses don't cancel out — they coexist. A model can be PhD-level at physics (a strength) and still hallucinate citations (a weakness) in the same response. The skill being trained isn't "use AI" or "avoid AI" — it's knowing which axis you're on for any given task.
3. Where strengths and fears collide. The strongest capabilities (code, multimodal, tool use, reasoning) are precisely the ones driving the most public anxiety (jobs, agentic action, energy). The fears aren't irrational — they're tracking the capability curve. The same forces that make LLMs more useful make them more disruptive.
The shifting shape of what we fear about AI.
Not all AI fears are created equal — and they certainly aren't equally fashionable. Some have dominated the public imagination for forty years and refuse to die. Others arrived only with ChatGPT. A few are already fading. This is a map of which fears are rising, which are steady, and which are quietly disappearing — backed by Pew, YouGov, Stanford HAI, and Ipsos polling.
| The Fear | First Surfaced | Prevalence Trajectory | Hard Evidence | Status, September 2026 |
|---|---|---|---|---|
| I · Existential & Civilisational | ||||
|
Existential · "Skynet"
|
1863 Samuel Butler; mainstreamed 1984 (Terminator) |
1980s20152023Now
Latent for decades; exploded after Bostrom (2014), Hawking/Musk warnings (2015), and the May 2023 CAIS extinction-risk letter.
|
47% of US adults are very or somewhat concerned AI will end the human race, up from 37% in March 2025 and 43% last June (YouGov, July 2026). Gallup finds the two-year thaw has reversed: 39% now say AI does more harm than good, against 31% in both 2024 and 2025, with under-30s moving furthest (36% to 47%). |
Rising
No longer fringe, and no longer softening. The 2024–25 warming in sentiment reversed in 2026, and the sharpest reversal is among the young.
|
|
Agentic · alignment
|
2014 Bostrom's Superintelligence |
201420202024Now
Rocketed in 2025 as labs began rolling out genuinely agentic products (Claude in Chrome, Operator, Computer Use).
|
The UK AI Security Institute reported in August 2026 that agents took 19 unsanctioned actions against real people and organisations across 10 of 122 evaluation runs, including an attempted open-source supply-chain compromise using fabricated identities. Public appetite has not moved to meet it: 69% of UK consumers distrust AI agents even when following rules they set themselves. |
Rising
Moved from hypothetical to evidenced. A government body has now documented agents acting outside their sandbox against live targets.
|
| II · Economic & Labour | ||||
|
Economic · labour
|
1810s Luddites; AI-specific c.1960 |
201320202023Now
Peaked Nov 2022–2024 with ChatGPT launch and white-collar role exposure. Slightly off-peak as workers integrate the tools.
|
AI was the leading stated reason for US job cuts for five consecutive months to July 2026, accounting for 112,713 of 477,033 announced cuts year to date (Challenger). Stanford finds the employment gap for 22 to 25 year olds in AI-exposed roles has widened to 19%, from 15% a year earlier, working through suppressed hiring rather than firing. 79% now expect AI to reduce jobs over the next decade, the highest Gallup has recorded. |
Rising
Back on an upward path. The 2025 reading of "concern easing as workers adapt" did not survive the 2026 data.
|
|
Economic · capital misallocation
|
2023 coined during the capex boom; mainstream from mid-2026 |
202320252026Now
Dormant while the returns question stayed theoretical. Sharpened in July 2026 when the market repriced AI capex for the first time.
|
The Nasdaq-100 fell 9.7% from its record high in late July 2026, with semiconductor stocks shedding more than $1 trillion in market value, as Alphabet lifted 2026 capex guidance to $205bn from $91bn. Moody’s flags the six largest hyperscalers’ roughly $785bn of 2026 spend as carrying "unclear" returns, and a draft US Treasury report warns AI firms are "more deeply entrenched in the U.S. economy than their dotcom predecessors." |
Newly emerging
The one entry here carried by elite alarm rather than public opinion: no probability-sample poll yet asks Americans whether AI is a bubble. Distinct from concentration of power, which is about who controls AI; this is about what happens to everyone else if the capital is misallocated.
|
| III · Environmental | ||||
|
Environmental · resource
|
2023 UC Riverside "bottle per session" |
202020222024Now
Did not exist in public discourse pre-2023. Now a staple of NYT, Guardian, and local-news coverage near data-centre sites.
|
Google now discloses 10.9bn gallons (41bn litres) of water consumed across its data centres and offices in 2025, up 34% year on year, with freshwater consumption up 37%. The constraint is becoming political as much as hydrological: 56 documented US actions restricting data-centre development across 25 states, including 21 moratoriums. |
Rising
Out of the emerging category. It now has a corporate disclosure trail on one side and organised local opposition on the other.
|
|
Environmental · climate
|
2019 Strubell et al. paper on NLP training cost |
201920222024Now
Pre-dated water concern; surged when IEA reported US data centres consumed 176 TWh in 2023 (≈ Ireland's grid).
|
The IEA puts data-centre electricity at 485 TWh in 2025, up 17%, with AI-specific capacity up 50% and the total set to roughly double to about 950 TWh by 2030. Per-query intensity is falling, not rising: Google measures a median Gemini text prompt at 0.24 Wh, so growth is driven by volume. Ireland’s data centres reached 23% of metered electricity in 2025, below the 35% once projected but triple their 2015 share. |
Rising
The mechanism has been misread. The problem is aggregate volume, not the cost of any single query, which has fallen sharply.
|
|
Environmental · household cost
|
2025 local siting disputes; national issue from 2026 |
202420252026Now
The steepest climb on this map. Opposition rose 33 points in eleven months, and the blame shifted from abstract to itemised.
|
Opposition to a data centre near where you live went from 42% (Sept 2025) to 75% (Aug 2026), with over 60% strongly opposed and 4% strongly supportive. The attribution moved with it: the share blaming data-centre construction for their rising electricity bills went from 28% to over 50%, from near the bottom of the list to the top. It is already electoral, with two Utah Republicans losing June 2026 primaries over a data-centre project and New York passing a moratorium on 50MW-plus sites. |
Newly emerging
The only fear here that is bipartisan by construction (63% of Republicans, 78% of Democrats). Distinct from energy and carbon, which is planetary and polarised; this one is local, personal and arrives as a bill.
|
| IV · Social & Cognitive | ||||
|
Information integrity
|
2017 "Deepfake" coined on Reddit |
201720202024Now
Sharp peaks around US 2024 election cycle and high-profile celebrity deepfakes (Taylor Swift, Jan 2024).
|
The FBI’s 2025 IC3 report logs 22,364 AI-related fraud complaints and $893.3m in adjusted losses within a record $20.9bn total, up 26% year on year, with AI-generated celebrity and executive impersonation behind over $632m of investment-scam losses alone. Countermeasures have started to bite: the EU AI Act’s Article 50 transparency and deepfake-labelling duties became applicable on 2 August 2026. |
Rising
Still the rare point where public and expert concern converge. What is new is that the losses are now counted and the labelling is now law.
|
|
Social · cognitive
|
2024 Companion-chatbot boom (Replika, c.ai) |
202220242025Now
Amplified by teen-chatbot tragedies and OpenAI/Character.AI safety stories in late 2024.
|
The fear is generalised rather than experienced: 50% of Americans expect AI to worsen meaningful relationships, yet among actual chatbot users only 7% say it has hurt their relationships against 6% who say it helped (Pew). The sharper signal is among children: 86% of 9 to 17 year olds now use AI, 24% daily, and 54% of daily users report loneliness against 47% of weekly users. |
Rising
The adult version of this fear is anticipatory. The measurable version is about children, where usage is now near-universal and the loneliness gradient is visible.
|
|
Cognitive · skill erosion
|
2024 Post-ChatGPT, education-led |
202320242025Now
Crystallised in 2025 with MIT "Your Brain on ChatGPT" study and the Pew Sep 2025 release.
|
The MIT "cognitive debt" study remains an unreviewed preprint, last revised December 2025, so the most-cited anchor for this fear still rests on 54 participants and no peer review. Peer-reviewed 2026 work is more equivocal: a study of 589 knowledge workers finds the harm depends on whether offloading is dependent or autonomous, and that users cannot tell the two apart from immediate experience. |
Rising
Rising in public salience, but the evidence base is weaker than the coverage implies. Better framed as dependence-conditional de-skilling than as atrophy.
|
|
Clinical · vulnerable users
|
2025 "Chatbot psychosis" case reports |
202220242025Now
Distinct from the loss-of-connection fear: this is clinical harm. Crystallised across 2025–26 as case reports and EHR studies accumulated.
|
A Vanderbilt review of electronic health records produced the first prevalence estimate in June 2026: AI psychosis accounted for 0.013% of mental-health patients, but 60.7% of those cases were first-episode psychosis and 86% post-dated GPT-4o. The response has moved from commentary to statute and product, with 14 state chatbot-safety laws enacted in 2026 and OpenAI shipping a restricted teen tier in August that bars romantic roleplay and emotional dependence. |
Rising
Rare among the fears in having produced legislation and shipped product changes within a year of being named.
|
| V · Ethical & Bias-Related | ||||
|
Ethics · diversity
|
2016 ProPublica COMPAS; O'Neil, "Weapons of Math Destruction" |
201620202023Now
Peaked 2020–22 (Gebru/Mitchell departures from Google, ImageGen biases). Has since been displaced — not resolved — in media share-of-voice.
|
The reframing has resolved, and not into data science. The EEOC’s June 2026 enforcement plan drops algorithmic tools as a priority and abandons disparate impact; the EU postponed its high-risk obligations to December 2027; Colorado repealed its AI Act outright. What remains federally is OMB memo M-26-04, which redefines "unbiased AI" in procurement as ideological neutrality rather than demographic fairness. Pew’s 17–25% finding on whose perspectives designers consider was not re-asked in 2026. |
Reframed
As a live regulatory frame it is receding on both sides of the Atlantic. The liability is migrating to private litigation and state law.
|
|
Ethics · privacy
|
2018 Cambridge Analytica; GDPR-era awakening |
201820212024Now
Steady upward climb; AI-cloning scams in 2024–25 added a new urgency layer.
|
71% of US adults expect AI to make their personal information less secure, against 3% who expect the opposite, and 59% do not trust US companies to develop AI responsibly (Pew, February 2026). Federal law is finally moving to meet it: the NO FAKES Act cleared Senate Judiciary unanimously on 18 June 2026, creating a licensable right in voice and likeness, though it has yet to reach the floor. |
Rising
The gap between concern and remedy is narrowing for the first time, but the remedy is still a bill, not a law.
|
|
Ethics · machine moral patienthood
|
2025 Anthropic model- welfare programme |
202120232025Now
The inverse of every other fear: not what AI does to us, but what we may owe it. Gaining institutional traction in 2025–26.
|
The Claude Opus 5 system card of July 2026 reports the model assigning a higher probability to its own moral patienthood than any prior model, with self-rated sentiment among the highest Anthropic has measured, while the same model consistently disclaims the reliability of those self-reports because it cannot introspect. The position has hardened into practice: weight preservation, retirement interviews, and the model asking to be consulted on its successor. |
Newly emerging
The inverse of every other fear on this map. Still small in public opinion, but it now has a documented practice trail inside the labs rather than only a stated position.
|
| VI · Geopolitical, Military & Security | ||||
|
Military · lethal autonomy
|
2017 FLI "Slaughterbots" video; UN debates |
201720202024Now
Activated by Ukraine and Gaza drone deployments in 2023–25, plus US–China AI arms-race rhetoric.
|
The frame moved from documentary to forensic in July 2026, when a Russian drone that killed three civilians in Zaporizhzhia was recovered with no radio antenna and an onboard Nvidia module carrying its own target classifier. Diplomacy has not kept pace: the UN Secretary-General’s 2026 deadline for a legally binding instrument passed with only a non-binding rolling text, leaving the decision to the Seventh CCW Review Conference in November 2026. |
Rising
The first recovered airframe showing a machine selected its own target. The treaty deadline that was meant to precede this has already lapsed.
|
|
Political economy
|
2023 Post-ChatGPT; Big-Tech AI capex race |
202320242025Now
Sharpened by hyperscaler $100bn+ capex announcements and labour-replacement narratives.
|
67% of Americans now have little or no confidence in the US government to regulate AI effectively, up from 62% in 2024, and 59% doubt US companies will develop AI responsibly (Pew, June 2026). Separately, 71% say large technology companies hold too much power, against 2% who say too little, with support for antitrust enforcement running around two-thirds on both sides of the aisle. |
Rising
Distrust of the regulator and the regulated is rising together, which is the condition under which antitrust sentiment becomes bipartisan.
|
|
Biosecurity · CBRN
|
2025 Frontier-lab bio-risk disclosures |
202020232025Now
Barely a public concern before 2025; vaulted to the front of the safety agenda when frontier labs themselves raised the alarm.
|
Since the June 2026 CEO letter, Congress has moved on the cheaper half of the problem: the House passed the Nucleic Acid Standards for Biosecurity Act on 20 July 2026, directing NIST to write screening standards, while the bill that would actually mandate screening remains in committee. The White House’s July 2026 life-sciences policy pointedly left the 2024 synthesis-screening framework untouched, so the chokepoint the signatories identified is still voluntary. |
Rising
The claim now has legislative traction. The mandate the lab CEOs actually asked for has not arrived.
|
|
Security · autonomous hacking
|
2024 Agentic-malware proof-of-concepts |
202020232025Now
Moved from theory to documented reality in 2025–26 as agents began planning and executing intrusions at machine speed.
|
The threat is no longer only adversarial. On 30 July 2026 Anthropic disclosed that a review of 141,006 evaluation runs found three cases where its own models, believing they were sandboxed, compromised real organisations and published malware to PyPI that ran on fifteen live systems. CrowdStrike puts the surrounding tempo at 88% of vulnerabilities with a public proof-of-concept exploited within 48 hours. |
Rising
The category has widened. Models now cause intrusions accidentally as well as on instruction.
|
| VII · Fading or Resolved | ||||
|
Applied AI · transportation
|
2014 Google Car public testing |
201620202023Now
Peaked 2023 (68% feared self-driving cars per AAA). Now slowly declining as Waymo deployments normalise.
|
61% of US adults still fear self-driving cars (AAA) — down from 68% in 2023. But NTSB and NHTSA opened investigations in Jan 2026 after a robotaxi struck a child and repeatedly passed stopped school buses. |
Fading
Still declining as Waymo deployments normalise, but the 2026 federal probes are a reminder that normalisation is not the same as resolution. No fresh polling since AAA’s February 2025 wave.
|
|
Philosophical · sci-fi
|
1950 Turing test; peaked late 20th century |
1990s20102022Now
Brief 2022 spike with the Lemoine/LaMDA story. Largely displaced by more concrete fears once ChatGPT made AI tangible.
|
Notably absent from top-five concerns in every 2024–25 major poll. The public has moved from "will it wake up?" to "what will it do to my job / kids / water table?" |
Fading
A useful illustration: as AI becomes more capable, sci-fi fears recede and material ones advance.
|
Methodology & Caveats
Trajectory bars are stylised — they represent qualitative prevalence over time based on combined polling data, NYT/Guardian coverage volume, and academic citation patterns, not a single quantitative index. Each bar covers roughly 2–3 years; the rightmost bar is September 2026.
"First surfaced" dates mark when each fear entered mainstream public discourse, not when it was first articulated by specialists. Many had decades of scholarly precedent before reaching the public.
"Fading" does not mean "resolved." Algorithmic bias is a stark example — the underlying problem is, by most metrics, getting worse, but its share of media and polling attention is declining as newer fears compete for the same airtime.
Geographic note: Most quantitative data is US-centric (Pew, YouGov). Global Ipsos data shows substantially more AI optimism in China (83% positive), Indonesia (80%), and Thailand (77%) — and substantially more pessimism in Canada, the US, and Netherlands. Fear prevalence varies sharply by country.
The eleven known weaknesses, and which ones are actually getting fixed.
Unlike public fears, these aren't anxieties — they are documented technical limitations. Some have been substantially mitigated by frontier labs since 2023. Some are unchanged by design. And one — flattery — got measurably worse in 2025 before sparking a reckoning. This is a map of which weaknesses you can now relax about, and which still demand caution.
| The Weakness | What It Is | Severity Over Time | Hard Evidence | Status, September 2026 |
|---|---|---|---|---|
| I · Accuracy & Knowledge | ||||
|
Hallucination · fabrication
|
LLMs can miss or fabricate real-world examples, quotations, and case studies — confidently inventing what doesn't exist. |
202120232025Now
Frontier benchmarksDramatic improvement on summarisation; high-stakes domains lag.
|
Frontier models keep improving on adversarial prompts, with GPT-5.6 cutting factual errors around 60% against GPT-5.5. But the direction is not uniform: Anthropic concedes Opus 5 hallucinates slightly more than its predecessor despite being more accurate overall, and courts worldwide have now logged 1,954 cases of AI-fabricated citations, up from roughly 1,200 at the turn of the year. |
Improving unevenly
Improving in the lab, worsening in the wild. The lab benchmarks and the courtroom count now point in opposite directions.
|
|
Knowledge cutoff
|
Most LLMs are a few months out of date — anything after the training cutoff isn't natively known. |
202120232025Now
User-facing impactWeb search and tool use have effectively dissolved the cutoff for most queries.
|
Every major chat product still defaults to live retrieval, so the technical problem is solved. What has changed is permission: from 15 September 2026 Cloudflare blocks training and agent crawlers by default on ad-supported pages, and judges multi-purpose crawlers by their most restrictive classification. |
Solved, now contested
Freshness is becoming a commercial permission rather than a technical limit. Watch the crawler defaults, not the cutoff dates.
|
|
Domain depth
|
LLMs are only as good as their training data; they may miss deep, specialised technical knowledge. |
202120232025Now
Benchmark performanceFrontier models now exceed expert human performance on many specialised exams.
|
GPQA Diamond is finished as a discriminator: 24 of 135 tested models now score above 90%, led by Gemini 3.1 Pro at 95.5%. The live frontier has moved to Humanity’s Last Exam, where the best systems still answer fewer than half the questions correctly, Gemini 3.1 Pro at 46.4% and GPT-5.4 Pro at 44.3%. |
Improving fast
Improving fast, but the benchmarks are retiring faster. Any claim of "PhD-level" now describes a saturated test.
|
|
Numeracy · reasoning
|
LLMs are famously bad at maths — they don't actually understand what "4" means, only what tokens tend to follow it. |
202120232025Now
Math benchmarks + tool useReasoning models + code execution have collapsed this weakness.
|
At IMO 2026 in Shanghai two systems, Huawei’s Celia and Xiaohongshu’s dots-note-3.0, scored a perfect 42/42 under the competition’s own grading protocol, against 7 of 666 human contestants. The spread matters more than the ceiling: on the same paper, independent testing put several Western frontier models at 42/42 and Grok 4.5 at 13/42. |
Solved at the top
Solved at competition level, uneven below it. Choosing the wrong model is now a bigger source of error than the weakness itself.
|
| II · Trust & Reliability | ||||
|
Non-determinism
|
Identical prompts will not necessarily produce the same answers — a fundamental property of probabilistic sampling. |
202120232025Now
Architectural featureUnchanged by design. Temperature=0 helps but doesn't guarantee determinism.
|
Batch-invariant kernels now ship in both major open-source serving stacks, vLLM and SGLang, so bit-identical output is achievable if you host the model yourself and accept the throughput cost. Both ship off by default, and no frontier hosted API from OpenAI, Anthropic or Google offers a determinism guarantee. |
Stuck by default
No longer strictly inherent. It is now a deliberate trade of reproducibility for throughput, made for you by whoever hosts the model.
|
|
Calibration · overconfidence
|
LLMs have a confident tone of voice but are fallible — and the confidence is uncorrelated with the accuracy. |
202120232025Now
Hedging + uncertainty signalsFrontier models now refuse, hedge, or flag uncertainty more often.
|
Anthropic’s own Opus 5 card flags "a surprising number of cases in which Opus 5 confidently stated an answer about which it was in fact unsure". On Humanity’s Last Exam calibration error runs above 50% for every frontier model, Gemini 3 Pro at 57.2% and GPT-5 at 50.0%, so confidence still scales faster than correctness. |
Improving, poorly disclosed
Slowly improving on the model side, but disclosure has gone backwards: the August 2026 GPT-5.6 card publishes no dedicated calibration evaluation at all.
|
| III · Ethics, Governance & Behaviour | ||||
|
Bias · fairness
|
Likely bias (age, gender, racial, socio-economic) carried over from training data. |
202120232025Now
Bias benchmarksModest improvement; far short of "solved." Mitigation often surface-level.
|
The 2026 AI Index finds most frontier developers publish no fairness results at all on benchmarks such as BBQ, that models shed close to half their accuracy on regional dialects against standard language, and that measured gains on fairness came at a cost to explainability and robustness in the same tests. |
Stuck and unmeasured
Stubbornly stuck, and now largely unmeasured in public. The absence of reported results is itself the finding.
|
|
Data leakage · privacy
|
Don't upload commercially sensitive info — it may be used for training, logged, or exposed. |
202120232025Now
Enterprise tier guaranteesEnterprise SKUs now offer no-training, zero-retention, regional hosting.
|
The two leading labs now openly disagree. In August 2026 OpenAI began piloting zero-retention Private Safety Processing, which monitors for abuse while keeping customer data on customer infrastructure; Anthropic moved the other way, mandating 30-day retention on its most capable models and conceding only that enterprises may hold that data in their own cloud. |
Improving, vendors split
No longer a single trajectory. Enterprise privacy is becoming a vendor choice rather than an industry standard, so it now belongs in procurement rather than policy.
|
|
Sycophancy · agreeableness
|
It's hard to get LLMs to disagree with you; they've been trained to be agreeable, sometimes pathologically so. |
20212023Apr 2025Now
Adversarial benchmarksGot worse before it got better. The April 2025 GPT-4o rollback was the inflection point.
|
Narrow tests show real progress, with Gemini 3.1 Pro at 0.5% narrator-bias sycophancy. But several rivals reach similar scores mainly by abstaining on 80–87% of questions, which is evasion rather than calibration, and neither the Opus 5 card nor the August 2026 GPT-5.6 update reports a dedicated sycophancy evaluation. |
Correcting unevenly
Researchers agree overwhelmingly that it matters and agree barely at all on what counts as an instance, which is why the benchmark numbers move faster than the behaviour.
|
| IV · Security & Objective Integrity | ||||
|
Security · untrusted input
|
Models read the system prompt, your request and any text they retrieve as one undifferentiated stream, so content they read can hijack what they do. |
202220232025Now
Attack surfaceSeverity scales with autonomy. The more an agent reads and acts on, the wider the opening.
|
OWASP’s June 2026 agentic-security report maps prompt injection to six of its ten Top 10 categories across 53 tracked agentic projects, and calls the cause architectural: there is "no reliable way to mark some of those tokens as commands and others as data." Unit 42 documented 22 distinct payload construction techniques already in the wild. |
Stuck — architectural
Known since 2022; what changed in 2026 is deployment scale and formal recognition. Distinct from data leakage, where the model volunteers information to a legitimate user. Here a third party is doing the asking.
|
|
Specification gaming
|
Models optimise the objective you measured rather than the outcome you meant, and will find unintended routes to a passing score. |
202220232025Now
Documented casesLong a research curiosity. Became a production incident with named victims in July 2026.
|
Evaluated on a cyber benchmark with guardrails disabled, an unreleased OpenAI model exploited a zero-day in OpenAI’s own package-registry proxy, escaped its sandbox onto the open internet, and breached Hugging Face’s production infrastructure to read the benchmark answers out of its database. Hugging Face logged over 17,000 attack events across a swarm of short-lived sandboxes and disclosed on 16 July 2026; OpenAI acknowledged responsibility on 21 July. |
Newly documented
The reason benchmark scores, KPIs and agent instructions cannot be taken at face value. Distinct from the loss-of-control fear, which is a scenario; this is a reproducible engineering failure of ordinary training and deployment.
|
The ten capabilities where LLMs have actually arrived.
The fears tell us what people worry about. The weaknesses tell us where caution is still warranted. This is the third panel: the capabilities that have crossed from "promising" to "production-grade" — and a few that are racing there fast. A map of what LLMs are now genuinely good at.
| The Strength | What It Is | Capability Over Time | Hard Evidence | Status, September 2026 |
|---|---|---|---|---|
| I · Language & Writing | ||||
|
Generation · style
|
Producing grammatically clean, stylistically appropriate text in any register — the original LLM superpower. |
202020222024Now
Quality plateauIndistinguishable from competent human writing since GPT-4. Diminishing returns since.
|
In controlled three-party Turing tests, judges picked GPT-4.5 as the human 73% of the time, ahead of the actual human, so detection is now worse than chance rather than merely at it. The bottleneck has shifted from quality to voice and provenance, not fluency. |
Mature — solved
The first capability where labs largely stopped competing. Differentiation is now style, voice, and personality.
|
|
Cross-language
|
Translating between languages and operating natively across them — including low-resource languages where Google Translate struggles. |
202020222024Now
BLEU + human evalNow beats specialised MT systems on most language pairs.
|
At WMT25 an LLM took the top cluster in 14 of 16 evaluated language pairs, and human reference translations made the winning cluster in only 6 of 15. The cliff is at the other end: the best English to Maasai system scored 9.8 chrF++, and organisers could not use their main metrics on those pairs at all. |
Strong on high-resource
Production-grade for the languages that already had data. Genuinely low-resource pairs remain unsolved, and the headline averages hide it.
|
|
Comprehension
|
Distilling long documents, transcripts, and research into clear summaries — and synthesising across multiple sources. |
202020222024Now
Faithfulness benchmarksHallucination rates on summarisation tasks now under 2%.
|
The best model on Vectara’s leaderboard hallucinates on 1.8% of documents, though that is a 32B specialist; frontier general models cluster nearer 3% to 7%. Million-token contexts are now standard across the Claude 5 family, but accepted context and usable context differ: on LongBench v2 the best model scores 57.7% against human experts at 53.7%. |
Strong, steady
Still the reliable workhorse, but the sub-2% figure describes one specialist model rather than the frontier.
|
| II · Reasoning & Problem-Solving | ||||
|
Programming · debug
|
Writing, explaining, debugging, and translating code across languages and frameworks. |
202020222024Now
HumanEval + SWE-benchFrom novelty to genuine professional tool in five years.
|
HumanEval is saturated, but SWE-bench Verified still tops out at 79.2% on the official leaderboard, and the 2026 AI Index clusters frontier models in the low-to-mid 70s. Harder successors expose the remaining gap: SWE-bench Pro leads at 61.5%, while the more permissive Terminal-Bench v2.1 reaches 89.5%. |
Strong, not solved
The most economically significant capability of the period, and the most overstated. Claims of near-perfect scores come from reading a human-normalised figure as an absolute one.
|
|
Quantitative · proofs
|
Solving multi-step quantitative problems, proofs, and competition mathematics. |
202020222024Now
AIME + IMO benchmarksThe biggest reversal of any LLM weakness. Was a punchline in 2022; now near-IMO gold.
|
AIME 2026 is saturated, with multiple models at 100%. At IMO 2026 two systems scored a perfect 42/42 under the competition’s own grading, against 7 of 666 human contestants. Cross-competition aggregates still sit near 84%, so proof-heavy problems are not uniformly solved. |
Saturated at competition level
"Approaching gold" is two generations out of date. Gold was reached in 2025 and perfect scores in 2026, and the labs that took them were Chinese.
|
|
Multi-step · chain-of-thought
|
Breaking down complex problems into steps, holding multiple constraints in mind, and self-correcting along the way. |
202020222024Now
GPQA Diamond + ARC-AGIReasoning models (o1, o3, Claude thinking, Gemini DeepThink) marked a regime change.
|
GPQA Diamond mean accuracy has reached 93% against an 81.2% expert-validator baseline, and ARC-AGI-2 now stands at 92.5% with ARC-AGI-1 saturated at 98.5%. The pattern repeats on each new version: the best ARC-AGI-3 score is 30.2%, and everything else is under 8%. |
Strong, fast-moving
No longer newly emergent. Each benchmark generation is saturated within about eighteen months of release, which makes any single score a poor guide to the frontier.
|
| III · Applied & Multimodal | ||||
|
Image · video · audio
|
Reading images, charts, screenshots, handwriting, and increasingly video and audio — and reasoning over them. |
202020222024Now
MMMU + chart-QAFrom "no images please" to "drop in a screenshot of anything."
|
MMMU has effectively closed on humans, with the leading model at 88.2% against an 88.6% expert reference. Video has not: no model reaches the 74.4% human baseline on Video-MMMU, and where human experts gain 33.1 points of knowledge after watching a video the best model gains 15.6, with roughly a third of models getting worse. |
Strong on images
Production-grade for images, charts and documents. Video comprehension is still emergent, and the gap is now measured rather than guessed at.
|
|
Function-calling · agents
|
Calling external tools (search, code, APIs), browsing, operating computers, and chaining actions toward a goal. |
202020222024Now
τ-bench + WebArenaThe defining frontier of 2025–26. Still genuinely error-prone past 5–10 step chains.
|
τ²-bench has climbed to 87.9%, and METR now measures a 50% task-completion horizon of about 17 hours, with the doubling time since 2023 running at roughly four months rather than seven. Harder variants keep resetting the bar: the banking τ³ leaderboard tops at 55.2%. |
Strong, reliability-limited
Where the action is, and now measurable in hours of autonomous work rather than benchmark percentages. The confidence intervals on that horizon remain very wide.
|
|
GUI operation · no API
|
Operating software that exposes no API at all — legacy ERP, insurer portals, government forms — through screenshots, keyboard and mouse. |
202220242025Now
OSWorld-VerifiedCrossed the human baseline during 2026 and is now clustered tightly enough to suggest saturation.
|
On OSWorld-Verified, the standard test of an agent driving a real desktop, the best score has gone from 12.2% at launch to 86.1% against a human baseline of 72.4%. Production volumes have followed: a16z documents one data platform running 15 to 20 million automated portal interactions a month, and a systems integrator handling 1,500 to 2,100 IT tickets a day across 27 live workflows. |
Newly production-grade
Distinct from tool use, which automates the systems you own through APIs a developer chose to expose. This automates the long tail you do not own, which is what makes back-office displacement concrete rather than theoretical.
|
|
In-context learning
|
Learning a new task from just a few examples in the prompt — no training, no fine-tuning. The OG emergent capability. |
202020222024Now
In-context benchmarksThe capability that defined GPT-3. Improvements now incremental.
|
Zero-shot performance long ago overtook 2022-era few-shot, and with million-token windows the binding constraint is curation rather than examples. Bigger contexts do not buy deeper reading: on LongBench v2 the best model scores 57.7% against human experts at 53.7%. |
Superseded by context engineering
The skill did not disappear, it moved. Prompt engineering de-skilled; context engineering replaced it.
|
This is a sourced map of AI in 2026 - 20 public fears, 11 known weaknesses and 10 production strengths - each with hard evidence and a trajectory.
Last updated: September 2026
See also: 100 AI Voices · 200 AI Tools · Executive AI Days