{"uuid": "cb4c415c-b32c-45b6-a681-c889a6d5fd32", "vulnerability_lookup_origin": "1a89b78e-f703-45f3-bb86-59eb712668bd", "author": "9f56dd64-161d-43a6-b9c3-555944290a09", "vulnerability": "CVE-2026-25253", "type": "seen", "source": "https://gist.github.com/iamalvisng/cec5751e8a7d57e4c0838f280dab962e", "content": "# The claim ledger \u2014 companion to \"The Best Lines Were the Wrong Ones\"\n\nA claim-by-claim fact-check of a one-hour Cantonese-language podcast interview about AI (the speaker stays unnamed; the point was never him). Every checkable assertion was pulled from a cleaned English transcript, verified against the best available source, and graded. High-stakes claims (specific numbers, named-entity actions, direct quotes) got a second, adversarial pass whose only job was to break them.\n\n**Source-quality tiers:** t1 = primary/official (the org itself, papers, regulators, official statistics) \u00b7 t2 = reputable secondary (established press, named analysts) \u00b7 t3 = single blog / aggregator, directional only.\n\n**Verdicts:** SUPPORTED \u00b7 MISSTATED (a source exists but says something materially different) \u00b7 UNSUPPORTED (no credible source found) \u00b7 NEEDS-SOURCE (plausible but uncited).\n\n&gt; **Correction (2026-08-04), left visible on purpose:** this ledger's original summary line said \"52 claims checked, 30 supported, 12 misstated, 7 unsupported, 3 needs-source.\" The table below has **64 numbered rows**. By row count: **12 MISSTATED, 8 UNSUPPORTED, 1 clean NEEDS-SOURCE (+2 hybrid), ~43 SUPPORTED with varying hedges.** The summary line was a compression that failed its own file \u2014 which became the second half of the essay this ledger accompanies. Cite the table, not summaries. Including this one.\n\n---\n\n## AI history\n\n| # | Claim | Verdict | Best source (tier) | Evidence / correction |\n|---|---|---|---|---|\n| 1 | GPT-3 released 2020 | SUPPORTED | arXiv 2005.14165 (t1) | Preprint 28 May 2020; API beta 11 June 2020 |\n| 2 | GPT-3 &gt;100x larger than GPT-2 | SUPPORTED | OpenAI paper (t1/t2) | \"over two orders of magnitude\" \u2014 175B vs 1.5B \u2248 117x |\n| 3 | ChatGPT launched Nov 2022, not Nov 2023 | SUPPORTED | openai.com/index/chatgpt (t1) | 30 Nov 2022. The transcript's \"Nov 2023\" was a speech-recognition/speaker error |\n| 4 | ~40 human contractors for early RLHF | SUPPORTED | Ouyang et al., InstructGPT (t1) | \"a team of about 40 contractors\" via Upwork and ScaleAI |\n| 5 | OpenAI never disclosed the GPT-3.5 labeler count | MISSTATED | InstructGPT paper (t1) | It *was* disclosed \u2014 ChatGPT is OpenAI's \"sibling model\" to InstructGPT, same method. Correct framing: not separately re-disclosed for ChatGPT |\n| 6 | Transformer came out of machine translation | SUPPORTED | arXiv 1706.03762 (t1) | Paper frames the Transformer against RNN translation quality/parallelization limits |\n| 7 | GPT-1/2 had almost no impact; GPT-3 first real influence | SUPPORTED (overstated) | TechTarget (t2) | GPT-2 did get real attention via the \"too dangerous to release\" controversy. GPT-3 was the commercial inflection |\n| 8 | Distillation + AI eval is the *main driver* of the narrowing gap | UNSUPPORTED as stated | Epoch AI (t1) | Real techniques, but no source names them the main driver. Compute efficiency, post-training recipes and competitive pressure cited more often |\n| 9 | Closed/open gap now ~6 months | MISSTATED | Epoch AI open-closed ECI gap (t1) | \"an average of four months\" (Jan\u2013May 2026). Six months is stale \u2014 the 2025 figure |\n| 10 | DeepSeek compressed the release cadence | SUPPORTED directionally; causation NEEDS-SOURCE | industry retrospective (t3) | Cadence did compress after DeepSeek R1 (Jan 2025) \u2014 6\u20139 month cycles to a 4\u20136 week competitive floor \u2014 but sources describe multi-lab dynamics, not isolated DeepSeek causation |\n\n## Money\n\n| # | Claim | Verdict | Best source (tier) | Evidence / correction |\n|---|---|---|---|---|\n| 11 | Anthropic ~$14B by Feb 2026 | SUPPORTED | Anthropic official (t1) | \"Our run-rate revenue is $14 billion, and has grown over 10x in each of the past 3 years\" |\n| 12 | Anthropic grew ~12x | MISSTATED | Anthropic official (t1) | Anthropic's own wording is \"over 10x\" per year. No official 12x figure |\n| 13 | OpenAI $2B \u2192 $20B by Feb 2026 | MISSTATED | OpenAI CFO disclosure via Epoch AI (t2) | Actual: $2B (2023) \u2192 $6B (2024) \u2192 $20B+ (2025) \u2192 ~$25B by Feb 2026. The \"20 months earlier\" framing is wrong |\n| 14 | 2026 AI capex \u2248 one year of US military spending | SUPPORTED, order-of-magnitude only | Futurum (t2/t3) + CBO (t1) | ~$1.04T both sides \u2014 but only on the broadest AI definition (hyperscalers + Oracle + neoclouds + sovereign + China). Big-Five-only is $660\u2013725B, ~two-thirds |\n| 15 | Most lab revenue is B2B, not consumer | SUPPORTED with caveat | Forbes (t2) | True for Anthropic (~80% API). OpenAI is closer to 50/50 by mid-2026 |\n| 16 | Jensen Huang's \"five-layer cake\" | SUPPORTED | NVIDIA blog (t1), Forbes (t2) | Davos 2026. Layers: energy, chips, infrastructure, models, applications |\n| 17 | Google trains on its own TPUs, not Nvidia | SUPPORTED | The Conversation (t2) | \"All phases of Gemini 3 training ran on Google-made TPU v5e and v6e pods without fallback to Nvidia GPUs\" |\n\n## Rankings and usage\n\n*(This whole block leans on tier-3 aggregators. Treat as directional.)*\n\n| # | Claim | Verdict | Best source (tier) | Evidence / correction |\n|---|---|---|---|---|\n| 18 | ChatGPT still #1 on mobile and web | SUPPORTED with nuance | TechCrunch (t2) | 53.9% of worldwide web visits vs Gemini 27.9%, Claude 9.2% \u2014 but share fell below 50% for the first time in June 2026; Gemini +450% YoY, Claude +855% YoY |\n| 19 | Big four = OpenAI, Anthropic, Google, xAI; Meta \"constrained\" | MISSTATED | capex trackers (t3), SemiAnalysis (t2) | Depends entirely on the metric. By capex Meta is ~$125B in 2026, third among hyperscalers \u2014 not capital-constrained. Its constraint is talent/model quality |\n| 20 | Doubao is China's strongest model | UNSUPPORTED | OpenCompass rankings via aggregators (t3) | No single Chinese model leads mid-2026. Qwen3-Max leads LiveBench/Arena-Hard, DeepSeek-V4-Pro tops aggregate coding, GLM-5.2 leads long-horizon agent tasks, Ernie 5.1 topped LMArena preference. Doubao is top-4-ish |\n| 21 | Doubao's top model is closed-weight | SUPPORTED (first half) | Doubao model docs (t3) | Closed-weight is correct. \"Stronger than the open models\" is not \u2014 see #20 |\n| 22 | Ranking: DeepSeek &gt; GLM &gt; Qwen &gt; Ernie &gt; MiniMax/Moonshot | UNSUPPORTED | multiple comparisons (t3) | No source supports this fixed hierarchy. Benchmark-dependent and contested |\n| 23 | GLM/Zhipu came out of Tsinghua | SUPPORTED | (t3, consistent across sources) | Spun out of Tsinghua's Knowledge Engineering Group (THUDM), 2019 |\n| 24 | DeepSeek, GLM, Qwen ship open weights | SUPPORTED | (t3) | All three, as of mid-2026 |\n| 25 | DeepSeek got famous by open-sourcing, investment followed | SUPPORTED | CSET (t1-adjacent), IISS (t2) | \"turning DeepSeek from a little-known research team into China's best-known AI company almost overnight\" |\n| 26 | Western labs rarely open their strongest weights | SUPPORTED with exceptions | OpenAI help center (t1), The Register (t2) | Real exceptions: gpt-oss-120b/20b (Aug 2025), Gemma 4 (Apr 2026). But GPT-5.x and Gemini 3 stay closed \u2014 holds for *strongest* |\n\n## Power and infrastructure\n\n| # | Claim | Verdict | Best source (tier) | Evidence / correction |\n|---|---|---|---|---|\n| 27 | China's power ~3x the US by 2027 | **MISSTATED** | Ember via Visual Capitalist (t2 from t1 data) | \"China generated 10,087 TWh of electricity in 2024, while the U.S. generated 4,635 TWh\" \u2014 ~2.2:1, trending to ~2.4\u20132.5:1 by 2027. Correct phrasing: \"more than double, and widening\" |\n| 28 | US has essentially nothing under construction | **UNSUPPORTED** | Rabobank/Wood Mackenzie (t2) | US builds ~40 GW/yr, needs ~80 GW/yr. Accurate version: \"building at about half the needed rate, far below China's scale\" |\n| 29 | Musk/Google proposed space data centers, power the driver | SUPPORTED | DataCenterDynamics, MIT Tech Review (t2) | Musk: \"you're power constrained on Earth\". SpaceX filed with the FCC; Google's Project Suncatcher targets two prototype satellites by early 2027. (The Musk quote could not be fetched directly \u2014 403 \u2014 treat as reported) |\n| 30 | US demand flat partly due to offshoring | SUPPORTED with nuance | EIA (t1) | \"U.S. electricity consumption was essentially flat for nearly two decades.\" EIA attributes this to efficiency + the manufacturing-to-services shift; offshoring is a widely-cited secondary attribution, not EIA's own wording |\n| 31 | Orbital data centers have an unsolved heat problem | SUPPORTED | IEEE Spectrum, 1-ACT (t2/t1) | Radiative rejection only: ~1,200 m\u00b2 of radiator to shed 1 MW; a single 700 W chip needs ~1.4 m\u00b2 |\n| 32 | Power is the binding constraint; demand will exceed supply | SUPPORTED | Gartner, IEA, Uptime (t2 from t1) | Gartner: 40% of AI data centers power-constrained by 2027. Global DC consumption +26% in 2026 to 565 TWh |\n| 33 | Musk timelines slip 3\u20135x | SUPPORTED directionally, figure UNSOURCED | prediction trackers (t3) | Of 20 tracked predictions: 4 delivered, 5 partial, 9 clearly missed. The \"3\u20135x\" multiplier is not a measured statistic anywhere \u2014 the speaker's rule of thumb |\n| 34 | GPU supply + power are the bottleneck, not architecture | MISSTATED (half) | industry analyses (t3) | By mid-2026, GPU supply is no longer primary \u2014 CoWoS capacity doubled in 2025 and again through 2026. Power is now the sole primary constraint |\n\n## Agents and OpenClaw\n\n| # | Claim | Verdict | Best source (tier) | Evidence / correction |\n|---|---|---|---|---|\n| 35 | OpenClaw is real, open-source, viral | SUPPORTED | openclaw.ai (t1), DigitalOcean (t2) | Peter Steinberger, launched Nov 2025 as \"Clawdbot\", most-starred repo on GitHub (346k+ stars) in under five months. Viral peak was Q1 2026, not mid-2026 |\n| 36 | Huang called it the most important software ever shipped | SUPPORTED | CNBC, HPCwire (t2) | Morgan Stanley TMT Conference, March 2026: \"probably the single most important release of software... probably ever\" |\n| 37 | One person, AI-assisted, no monetization | SUPPORTED | Euronews (t2) | Steinberger: \"It's a free, open source hobby project.\" MIT license. (He has since joined OpenAI, Feb 2026 \u2014 not mentioned by the speaker) |\n| 38 | Major publicized security holes from broad access | SUPPORTED | Dark Reading, Barracuda (t2) | \"ClawJacked\" website-driven hijack; CVE-2026-25253, one-click RCE; tens of thousands of internet-facing instances exposed |\n| 39 | Connects WhatsApp/Telegram, messages you proactively | SUPPORTED with nuance | OpenClaw docs (t1) | `post_message` tool. Proactive posting is gated to scheduled cron/heartbeat, not arbitrary spontaneity |\n| 40 | Persistent \"soul\" file that accumulates history | SUPPORTED | OpenClaw docs (t1) | `SOUL.md`, read at startup and injected as a system-level prompt |\n| 41 | Tencent's Shenzhen install event was overrun | SUPPORTED, understated | SCMP (t2) | \"Nearly 1,000 people lined up outside... Tencent Holdings' Shenzhen headquarters\" \u2014 20 staff installing free, 6 March 2026 |\n| 42 | Major AI companies issued a joint statement | **UNSUPPORTED** | \u2014 | No joint statement found. Individual industry responses and OWASP's Top 10 for Agentic Applications exist |\n| 43 | Chinese giants incl. Xiaomi shipped their own | SUPPORTED | Yicai Global, CNBC (t2) | Xiaomi's MiClaw, Honor's Yoyo, Huawei's Xiaoyi; Zhipu, ByteDance, Tencent + 10 other giants |\n| 44 | GPT-4o sycophancy update rolled back after safety objections | SUPPORTED | openai.com/index/sycophancy-in-gpt-4o (t1) | Update 25 April 2025; rollback began 28 April 2025 |\n| 45 | GPT-5 trained to decline over fabricating; less personable | SUPPORTED | (t3) | Real and widely reported personality controversy; direction is right |\n| 46 | Claude's first agent shipped with 26 tools | **NEEDS-SOURCE** | \u2014 | No source found for 26 at any first agent release. Computer Use shipped ~3 (computer, text_editor, bash); Claude Code has ~15+ built-ins |\n| 47 | Tool-use benchmarks exist | SUPPORTED | (t3) | BFCL (Berkeley Function-Calling Leaderboard, v4 as of Apr 2026), \u03c4-bench, ToolBench |\n| 48 | Official 2025 guidance said don't let agents book flights | **UNSUPPORTED** | \u2014 | Not found. Likely conflated with DOT statements on airline AI *pricing* \u2014 a different story |\n| 49 | Agents still fail at flight booking | SUPPORTED, imprecise | tau-bench (t1), Tech Times (t2) | ~56% task success on tau-bench Airline \u2014 coin-flip, not \"never\". Agents reached airline sites directly only ~5% of the time |\n\n## Safety, history, economics\n\n| # | Claim | Verdict | Best source (tier) | Evidence / correction |\n|---|---|---|---|---|\n| 50 | German unemployment 5% (1929) \u2192 29% (1933) fueled the Nazis | SUPPORTED as rounding | Nuffield paper (t1) | Registered unemployment ~4.5% (1929) \u2192 peak ~30% in 1932, ~6M in early 1933. Peak year is 1932, not 1933 |\n| 51 | Hinton, Sutskever, Musk all warned of extinction risk | SUPPORTED with nuance | CAIS statement (t1) | Hinton and Sutskever signed the May 2023 CAIS statement. Musk signed the separate FLI pause letter, not CAIS |\n| 52 | Musk funded OpenAI; it began as a nonprofit | SUPPORTED | CNBC (t2) | Launched 11 Dec 2015 as a nonprofit research lab |\n| 53 | Musk broke with Larry Page over AI safety, motivating OpenAI | SUPPORTED as *Musk's account* | court testimony reporting (t2/t3) | Sworn self-report, not corroborated by Page. Attribute, don't assert |\n| 54 | Anthropic's founders left OpenAI over safety | SUPPORTED | (t2/t3) | Standard account; follows the 2019 Microsoft deal |\n| 55 | A safety lead quit a major lab in early 2026 over safety | SUPPORTED | AP wire (t2) | Mrinank Sharma, Anthropic, effective 9 Feb 2026: \"a confession of futility... the pressures arrayed against safety were winning\" |\n| 56 | Reid Hoffman: market models lag internal by 12\u201318 months | **UNSUPPORTED** | \u2014 | No transcript or article found. Appears to be a floating industry talking point mis-attributed |\n| 57 | LeCun says LLMs alone won't get there; nobody has rebutted him | MISSTATED (second half) | Brown University (t1) | First half accurate. Second half false \u2014 Sutskever, Amodei, Altman and Meta's own leadership publicly disagree; the scaling-focused reorg reportedly contributed to LeCun leaving Meta in Nov 2025 |\n| 58 | Harari's \"useless class\" | SUPPORTED | *Homo Deus* (t1) | Correct attribution |\n| 59 | YouTube took ~10 years to produce full-time creators | MISSTATED | (t2/t3) | Partner Program launched Dec 2007; some creators over $100k/yr within a year; full-time creators as a category by ~2010\u20132013. Closer to 5\u20137 years |\n| 60 | AI coding crossed a step change in Q3 2025 | MISSTATED | (t3) | No source pins Q3 2025. Commentary points to late-2024 and late-2025 milestones and gradual growth with humans in the loop |\n| 61 | Japanese perfume founder's AI poster went viral | **UNSUPPORTED** | \u2014 | Not found. Nearest real story is an AI-generated Iwate Prefecture tourism poster controversy |\n| 62 | California has advanced AI safety legislation | SUPPORTED | CA Legislature (t1), FPF (t2) | SB 53, Transparency in Frontier Artificial Intelligence Act \u2014 signed 29 Sept 2025, in force 1 Jan 2026, applies above 10^26 FLOPs. Already law, not merely \"advanced\" |\n| 63 | US corporations are legally obliged to maximize shareholder returns | **MISSTATED** | Stanford Law (t1), Harvard Corp Gov Forum (t2) | The classic legal myth. *Dodge v. Ford* (1919, Michigan, closely-held) is \"mere dicta\"; Delaware courts have never cited it on corporate purpose. Directors have broad discretion under the business judgment rule |\n| 64 | Ukraine is largely AI vs AI, humans nearly absent | **MISSTATED** | CSIS (t2), IEEE Spectrum (t2) | \"most autonomous systems remain 'human-in-the-loop' rather than fully autonomous\". AI does terminal guidance, GPS-denied navigation, target recognition. Humans plan, launch and usually confirm strikes |\n\n---\n\n## Not checked, on purpose\n\nThe speaker's first-person business claims \u2014 his monthly AI spend, his former company's headcount, project timelines before and after AI, two hiring anecdotes. Only he can confirm these. They are plausible and internally consistent, but they are **testimony, not evidence** \u2014 and, as the essay notes, they are also the one section where nothing could fail.\n", "creation_timestamp": "2026-08-04T15:51:21.179913Z"}