Does publishing the number that cuts against you actually outperform claiming the one that flatters you?

The mechanism is not complicated. A developer audience checks claims, in public, at speed. So an adjective costs nothing to write and earns nothing, while a number someone can go and verify does work — and a number that cuts against you does the most work of all, because nobody publishes those by accident.

Three distinct mechanics are on the record, and it is worth keeping them distinct rather than filing them all under “be honest”.

Dogfooding that names its own cost. Cloudflare migrated cdnjs onto its own developer platform and measured the exercise by the internal limits it had to raise — the interesting output was not “it worked” but the list of places it did not, yet. The same company repeated the mechanic in late August, moving its own blog onto EmDash, its own CMS, and publishing the load figures — 75 requests per second normal, a 7,000 RPS burst load test, 99.5% of static-file requests served from cache, and a 28,000 RPS DDoS absorbed mid-migration. The migration is the campaign; the numbers are what make it citable. GitHub’s September 2026 port of the Copilot runtime to Rust is the largest specimen so far — about 430,000 lines of TypeScript to about 832,000 of Rust in fourteen and a half weeks, primarily one developer using Copilot — and it earns its “wasn’t affordable before agents” by the same mechanic: 158 unsafe blocks located by file, 135 releases counted, and the regressions listed by category rather than smoothed over.

A benchmark that reports where you lose. The honest-benchmark pattern: publish the comparison including the runs where the competitor wins, and the whole table becomes citable instead of dismissible. Neon’s September entry is the cleanest recent case — 42 models on its own gateway run through a fixed support-ticket task with a stated seven-check rubric, publishing the full ~1,446x cost spread including the finding that open-weight models undercut the proprietary ones it resells. PlanetScale’s Neki post (September 11) applies the rule to a headline number of its own: 118.5 million queries a second, sustained sixteen minutes across 512 shards, stated in the same breath as what the run did not do — single-row reads only, primary-only with no replicas, no failover attempted, 67 errors a second. The caveats a skeptic would have supplied are printed first, which is what lets the number travel.

A workflow that prints its counter-number. Sentry shipped its Seer review numbers with the close-without-merge rate up 12.5% — and argued that the increase was healthy rather than hiding it. That is the hardest version of the move, because it requires having a thesis about why your bad-looking number is good.

The deprecation window is the same argument

This is where the thread earned its guide edits. When a product goes away, the window is the headline and the mechanics decide whether anyone gets burned. The ledger the site has built up is stark once it is laid side by side: Spark at roughly 27 days, GitHub Models 29 days to dead, Cerebras about a month, MCP at twelve months, HCP Vagrant Registry at about five (announced August 3; new boxes stop October 1, support ends November 2, operations end December 31, 2026 — corrected 2026-09-14 from a ten-month figure the archive had carried without a HashiCorp page behind it).

A dated window a developer can plan against is a proof point in exactly the way a benchmark is. A generous but vague one is an adjective. The ledger gained another short entry in September: GitHub gave four Copilot models roughly four weeks from announcement to their October 2 cutoff.

The buyer side arrived, and proof became a campaign format

Until late August the thread ran entirely on seller-side evidence — vendors publishing checkable things and appearing to benefit. The week of August 17 added the missing half: LeadDev cofounder Dave Fletcher, citing LeadDev’s own buyer research (reportedly 150 interviews plus audience surveys since January — self-reported, dated, and stated on the podcast rather than published with underlying data), says AI-first velocity claims measurably push skeptical engineering buyers away, with only a little over half its audience positive on AI at all. That is the first buyer-side data point saying the adjective doesn’t just earn nothing — it costs something.

The same week, the sell side turned proof into a campaign format rather than a content genre. Vercel put a $1M bounty with a dated window and a $50K per-report price on breaking its own agent sandbox; Replit announced pen tests by naming the bugs each method caught; Docker argued one security thesis across four dated posts in five days and named a framework — the Agent Baseline — before any competitor named theirs; and GitHub’s outage postmortem printed demand doubling, its own Copilot client amplifying the failure, and countable fixes. The distinct move is packaging: the proof now ships with a date, a price, or a name, which is what makes it a launch asset instead of a compliance page.

Tension

Founder credibility buys attention, not a verdict — and the reverse of this thread is that proof does not always win the room either. Block shipped Buzz in late July to 304 points on Hacker News and a broadly “LLM slop” reception: the reach was real and the judgment went against it anyway. Publishing something checkable is necessary, and this thread has not shown it is sufficient.

The other side of the ledger got its cleanest specimen on September 13: PostHog’s homepage now says “97% of users pay us $0” under the signup button, with no method and no denominator. It is a number, and it is flattering, and nobody can go and check it — which is the case the question above asks about. Whether it outperforms the agent-tools block it replaced is PostHog’s to know; what the thread can say is that a number without a denominator is an adjective with digits, and the skeptic supplies the denominator.

Open loops

What would move this thread. Each is meant to be checkable by a date — if one goes quiet, that is itself an answer.

  • A second limits-raised case study — a vendor that dogfooded, hit its own published ceiling, and said so with the number.

  • The first vendor to market a migration plan as a feature rather than an apology. Deprecation is currently damage control everywhere; treating it as a selling point is the untested move.

  • Does LeadDev publish its buyer research with the underlying numbers, or does a second buyer-side dataset corroborate the AI-messaging backlash? One research org's self-reported panel is the only buyer-side evidence the thread has.

    by next survey wave, expected by late 2026

  • Vercel's own report on the $1M sandbox challenge. The tally reached the trade press on 2026-09-15 (SecurityWeek — 1,285 reports, 91 validated findings, one critical and seven high, the two most serious in the Linux kernel's networking stack rather than Vercel's code), so the result is public; what is still missing is the attack-techniques write-up the program page promised, from Vercel itself. Narrowed to that, once.

    by mid-October 2026

On the record

20 dated entries filed onto this thread, newest first. Each carries its own primary source.

September 2026

  • 09-17
    GitHub GitHub ports the Copilot runtime to Rust, mostly solo

    GitHub rewrote its Copilot agent runtime from ~430,000 lines of TypeScript to ~832,000 lines of Rust in 14.5 weeks, shipping 135 releases along the way and publishing dozens of regressions it hit and fixed. The post says the port 'wasn't affordable before agents': work once needing a team a year or two took primarily one developer a few months, using Copilot itself.

  • 09-14
    PlanetScale PlanetScale publishes its Neki benchmark, caveats and all

    PlanetScale ran its sharded Postgres engine Neki to 118,538,803 single-row point-select queries per second, sustained for 16 minutes across 512 shards and 1.22 PiB of data, and published the number on 11 September alongside its limits: no writes, joins or cross-shard queries, shards primary-only with no replicas, no failover attempted during the measured window, plus an exact error rate (67/second, about one query in 1.8 million).

  • 09-13
    PostHog PostHog swaps an agent-tools block for "97% of users pay us $0"

    PostHog rewrote its homepage this week: against the 9 September snapshot below, a "Built-in tools for your agents" block is gone and the free-signup call to action now carries a self-reported stat — "97% of users pay us $0", no credit card required. The hero still sells the agent story ("Make your product self-driving"); the pricing-transparency line sits under the sign-up button rather than at the top, and PostHog does not say how the 97% is measured.

  • 09-08
    HashiCorp HashiCorp winds down its hosted Vagrant box registry

    HCP Vagrant Registry, HashiCorp's free hosted service for storing and distributing Vagrant boxes, stops accepting new boxes October 1, loses support November 2, and shuts down completely December 31. HashiCorp is pointing users at self-hosting boxes on S3, Azure Blob or similar, using the still-maintained community Vagrant CLI.

  • 09-04
    GitHub GitHub sets a four-week runway for four Copilot model deprecations

    GitHub will drop four models on October 2 — Gemini 3.5 Flash, Gemini 3.6 Flash, Kimi K2.7 Code and Claude Opus 4.7 — across Copilot Chat, inline edits, ask and agent modes, and code completions, pointing users to Gemini 3.8 Flash, Kimi K3 and Claude Opus 5 in their place. The notice gives roughly four weeks from announcement to cutoff.

August 2026

  • 08-25
    Cloudflare Rebuilds its blog on its own stack, then publishes the numbers

    The Cloudflare Blog itself now runs on EmDash, Cloudflare's own CMS built for Astro on its stack, with production traffic ramped from 1% to 100% in a day. The write-up publishes the load numbers: about 75 requests per second normally, a 7,000 RPS burst load test, 99.5% of static files served from cache, and an unrelated 28,000 RPS DDoS absorbed mid-migration. The migration itself, not a new product, is the proof.

  • 08-24
    The week — The trust surface is the campaign now

    LeadDev's buyer research put numbers on developer skepticism the same week Docker ran a four-post trust campaign, Vercel paid to be hacked in public, and GitHub printed its own outage numbers. The vendors who reach skeptical buyers are marketing proof, not promises, and the human gates are coming down behind them.

  • 08-22
    Docker Docker's fourth agent-trust move in five days

    Docker has published four security and trust posts in five days, each with its own primary source and a distinct shipped or announced capability. Hardened Images now cover system packages in the same SLSA Build Level 3 pipeline, a "17,600 Actions" post answers the OpenAI/Hugging Face agent intrusion with the six-outcome Agent Baseline framework Docker co-authored, Verified Publisher applications went self-serve with pull analytics that name the companies behind anonymous traffic, and Docker Sandboxes became a supported runtime in GitHub Agentic Workflows.

  • 08-21
    GitHub Publishes the August 17 postmortem, and it names its own limit

    A Central US data center component failed to scale under load and cascaded into authentication failures across github.com, Actions, the APIs, pull requests, issues and Copilot for seven hours and 47 minutes. A retry loop in the Copilot client then amplified traffic during recovery. GitHub blames platform demand nearly doubling, with monthly commits up from 1.4 billion to 2.9 billion, and promises more capacity, architecture that scales linearly, and consistent retry limits.

  • 08-14
    Netlify Runs one prompt through 11 models to sell OpenRouter support

    Netlify ran the same website-build prompt through 11 models, three times each — Claude Opus and Sonnet 5, GPT-5.6, Gemini 3.x, Kimi K3 and K2.7, GLM 5.2 and DeepSeek V4 among them. Credit cost varied from 2.4 to 519 for comparable output. The piece doubles as a launch vehicle for its expanded OpenRouter-backed Agent Runners model support.

  • 08-12
    Val Town Rebuilt its docs on its own platform, and named the DX cost

    Val Town moved its docs off Astro and Cloudflare onto Val Town itself — server-rendered, no build step, edits live in about 100ms — while keeping `llms.txt` and "Copy as markdown" for agent readers. The post, older than this sweep, is open about the trade: a snappier cached UX given up for immediate-feedback DX, priorities ranked UX over AX over DX, slow spots named.

  • 08-06
    Sentry Publishes the numbers from routing its AI bug-fix PRs to Slack

    Sentry published a build-in-public post on its internal bug-fix workflow: its Seer agent opens pull requests for detected issues, then Claude routines pick the most relevant engineer and notify them in Slack rather than leaving the PR to be found. It reports roughly 21% more action on those PRs, 13% more 48-hour responses, and 12.5% more close-without-merge. Sentry argues that last number is healthy, since engineers often chose a broader fix than Seer's narrower one.

  • 08-05
    Postman Ships a TypeScript SDK that regenerates itself from the spec

    Postman released @postman/api-sdk, a typed TypeScript client for its full API: workspaces, collections, environments, monitors and mocks. The client regenerates itself from the API's own spec, so an API change files an automated pull request against the SDK repository instead of waiting on a maintainer. It's built with Postman's own SDK Generator, which customers can point at their own APIs.

  • 08-05
    GitHub GitHub Spark shuts down on 27 days' notice

    GitHub Spark stops accepting new users and new apps immediately, and existing users get 27 days to export their work. Already-deployed apps keep running, but any app using the `llm()` function needs a new inference provider because GitHub Models, the service behind it, has already retired. That is the shortest window in this site's running deprecation comparison: MCP got 12 months, GitHub Models six weeks, HCP Vagrant Registry ten months.

  • 08-04
    HashiCorp A hosted registry retires on a ten-month clock

    HashiCorp is retiring the HCP Vagrant Registry in three phases spread over ten months: new box creation stops first, support ends next, then all operations cease. It documents S3 export guidance and keeps the open-source community edition alive as the fallback. That window is short of MCP's 12-month offramp but far past GitHub Models' six weeks.

  • 08-04
    GitHub GitHub markets Copilot CLI by showing its lawyers using it

    GitHub's own legal team, not engineers, show two Copilot CLI workflows they built: one drafts contracts from a plain-language style guide and a library of pre-approved agreements, the other handles DMCA notices, NDA triage and compliance checks. Both run on structured Markdown instructions rather than code. One attorney reports cutting review and drafting time roughly in half, a self-reported figure.

  • 08-03
    The week — Twelve months or six weeks, the deprecation window is the positioning now

    MCP's biggest spec revision started a twelve-month migration clock under every MCP server the same week GitHub retired Models on a six-week runway with no like-for-like replacement. How you end things is becoming as much of a marketing surface as how you launch them.

July 2026

  • 07-31
    GitHub GitHub Models goes dark six weeks after closing to new customers

    GitHub Models is fully retired for every customer, including those with active usage, six weeks after it stopped accepting new customers. The two-year-old service bundled an AI playground, a model catalog, an inference API and bring-your-own-key access. Departing users are pointed at Microsoft Foundry or Copilot's own model access, adjacent products rather than like-for-like replacements.

  • 07-30
    Cloudflare Moves cdnjs onto its own platform and makes it the campaign

    Cloudflare rebuilt cdnjs entirely on its own R2, Workers, Workflows, Queues and Durable Objects stack, replacing a fragmented Google Cloud setup, and published the write-up. The numbers: 108,000 requests per second, 9 billion a day, a 98.6% cache hit rate. The migration also raised a public platform ceiling: subrequests on paid plans now go up to 10 million.

  • 07-27
    GitHub A researcher shows 10,000 deleted malware repos grew back

    GitHub deleted about 10,000 malware-distributing repositories after a researcher's post hit Hacker News. The researcher then ran the same search script again and found the repos back within hours, following the same recognizable template that had persisted for roughly two years. The charge is that GitHub answered publicity with a one-time sweep, not a lasting filter.

Where this lands in the guide