Podcast · Ep 1

GitHub's Best Ad Is a Bug List

13:06 Will & Brian

GitHub's best ad this month is a list of bugs it found with its own open-source agent. Postman turned one laptop's agent traffic into a data report. And a Rust rewrite's price tag became the headline. Will thinks showing the work is the best devtools marketing there is; Brian wants to know whether anyone who signs a purchase order reads it.

Sources

  1. GitHub Blog - How we found 24 Android vulnerabilities using our open source AI security agent (opens in a new tab)
  2. GitHub Security Lab - seclab-taskflows README (opens in a new tab)
  3. GitHub Security Lab - All advisories discovered with AI agents (opens in a new tab)
  4. Hacker News - GitHub's Android vulnerabilities post (opens in a new tab)
  5. Postman Blog - What Passport Found in 3 Weeks of AI Agent Traffic (opens in a new tab)
  6. npm downloads API - @postman/postman-passport, August to September 2026 (opens in a new tab)
  7. GitHub Blog - Migrating the GitHub Copilot runtime to Rust, using Copilot (opens in a new tab)
  8. The Register - Microsoft agentically ports Copilot runtime to Rust for $120K (opens in a new tab)
  9. Hacker News - The Register's Rust port article (opens in a new tab)
  10. Hacker News - GitHub's Rust port post (opens in a new tab)
  11. Scaling DevTools - Dave Fletcher from LeadDev, what engineering leaders are buying in the AI era (opens in a new tab)
  12. Stack Overflow Developer Survey 2025 - Work (opens in a new tab)
  13. Endor Labs - Everyone Wins Their Own Benchmark (opens in a new tab)
  14. Buttondown - Developer marketing is show and tell (opens in a new tab)

Transcript

Click a time to jump there. Times are approximate.

  1. The best developer ad of the fortnight has no adjectives in it. It's a list of bugs. GitHub wants you to run its AI security agent, and the way it chose to do that on Monday was to publish the vulnerabilities it found in other people's Android apps, with the script to run it on yours. Postman did the same the week before with a data report from one laptop. And GitHub did it again with a rewrite that printed its own bill. Three posts in a fortnight, and each one shows you the work and prints what went wrong with it.

  2. Lovely. Who signs a purchase order after reading a bug list? I'm not being clever, I read all three this week and I enjoyed all three, and I can't name the person at any company who buys something because of them. So which one's the bug list?

  3. GitHub's. The agent first, one sentence. GitHub's Security Lab has an open-source AI agent that reads a codebase looking for security holes, and you aim it with what they call taskflows, packaged prompts. A researcher there wrote a set for Android apps. Now here's the thing to notice. The post is titled how we found twenty-four Android vulnerabilities. And the first section, before a single vulnerability is described, is how to run this on your own project. Open a codespace, run one script, give it an hour or two on a medium-sized repo. The findings come after the instructions. It's a manual with a brag attached, not a brag with a manual attached. The twenty-four is their own count, and I'll come back to that. But twenty-four reported bugs in real apps in a real store is not a demo. That's a receipt. They did the work, and they're showing you the receipt.

  4. So the how-to comes before the what-we-found. And what did they find? Give me one.

  5. OsmAnd. A navigation app, not a toy. The agent found that any other app on your phone, even one with no permissions at all, could silently import settings into OsmAnd, and one of those settings is which server the map tiles come from. Swap that for the attacker's server and every tile you load goes through them, which means where you are, and the start and end of every route you plan. An app with over ten million downloads, and the attacker needs zero permissions. And then, and this is the part a normal ad would cut, the post prints its own limits. The model kept reporting low-severity stuff even when told not to. It misjudged severity. It produced false positives. Every finding, they say, should be reviewed by a human researcher who knows mobile. And then the last line is, quote, "Start securing your project today. Run these taskflows against your own app and take the first step toward AI-assisted security." A call to action at the bottom of a security report. The report is the ad.

  6. Hang on, what does it cost me to run it? And can I check the twenty-four, or do I take their word?

  7. Cost: the code is MIT, free. But the post says you need a Copilot licence and it'll burn premium model requests. The readme is softer, you can point it at another AI API, and it warns a run can cost a non-trivial amount of money, their words. So the free tool has a meter on it, and the meter is GitHub's. Checking: as of this morning the Security Lab's advisories page lists five credited to that researcher, each with a CVE. The other nineteen you take on their word for now. And the post never says how many apps they ran it on, or how many findings got thrown out.

  8. So the receipt has a total and five line items. Mine's the opposite: every line item, from one laptop. Postman, last week. Postman, the API tool company, launched a thing this summer called Passport, a way to keep real API keys out of agents' hands, and its local community edition sits on your machine and watches what leaves it: which programs send which credentials to which hosts. Someone at Postman, Talia Kohan, left it running for about a month while she built an AI harness for a different Postman product. Then she read the log. Twenty-two distinct secrets had left her laptop. And the number that makes the report: her project's dot-env file, the place a developer thinks their secrets live, had four entries in it. Twenty-two leaving, four on the list. Where were the other eighteen? The keychain, git's credential helper, npm's config, variables inherited from the shell. None of it stolen, she's careful to say, just in use and on no list. And here's her line, the one the post stands on: "My static scanner stayed green the entire time, and it was right to: not one of those 22 secrets was in the repository." That's the pitch in one sentence. The scanner you already run is watching the repo while the secrets walk out the side door.

  9. That's a good line. So what do I do, revoke all twenty-two?

  10. You can't, and that's the finding the pitch rests on. One Anthropic API key was being sent by two unrelated programs, a Bun script and Claude Code. Revoke it for the suspicious one and you break the healthy one. She turns that into a metric, consumers per credential, and says it's the number Passport is designed to keep at one. That's the product entering the report through the finding. Near the end is a section called set up Passport and read your own traffic: install the community edition, go do your actual work for a week, and she says outright not to take her numbers, because she can tell you what to look for but not what you'll find. Now the flags. One developer. One laptop. Measured by Postman's own detector, by a Postman employee, while building a harness for another Postman product, which gets its own section near the end. And the published page still has a placeholder in it, in square brackets, that reads screenshot of the Passport dashboard. Nobody filled it in.

  11. A placeholder. In the ad. Did anyone install it?

  12. As far as a bad proxy can tell, barely. The npm registry shows forty-eight downloads of the Passport package in the four days after the post. Forty-eight. You could count them by hand. Downloads aren't users and bots count both ways, so I won't call that a verdict. I'll call it the only number anyone outside Postman can read, and it's small. The opposite of your story: GitHub has a total you can't check. Postman has a number you can check, and it's forty-eight.

  13. And the third one is both at once. Two weeks ago GitHub published a post about rewriting the Copilot agent runtime, the thing underneath Copilot, from TypeScript into Rust. More than eight hundred thousand lines of it. Most of the code was written by AI agents, using Copilot itself, driven primarily by one developer, a Microsoft engineer called Stephen Toub. One person, plus a team around him, the post says. And he printed the bill. About a hundred and twenty thousand dollars in tokens, plus, by his own rough estimate, three weeks of his own time, plus dozens of known regressions from the port, all fixed, and he adds he's completely sure there are more they don't know about. So picture that as a marketer. Your dogfooding post has the price of the dog food in it, and a list of what broke.

  14. Hang on. A hundred and twenty grand. Is that a lot?

  15. Depends who's reading the receipt, and that's the story. Toub's reading is: cheap, next to the alternative. Eight hundred thousand lines of production code, one developer, a few weeks of his time and a token bill, and his line is, quote, "Agents moved the price to where the project became tenable." Meaning a rewrite that size wasn't affordable before agents, the post's subtitle says exactly that. Now it's an invoice. That's the brag. The Register's reading, two days later, leans the other way. Their headline was Microsoft agentically ports Copilot runtime to Rust for a hundred and twenty K, with the regressions right behind it, framed as AI still struggling with Rust, though they do add that by hand it would have cost millions. Same number, read as six figures and a list of what broke. And on Hacker News the Register's version got the bigger thread, by a distance, over GitHub's own post. So the number GitHub would never have put in an ad is the number that got read. Which, I'd argue, is the point.

  16. Or it's the cost. You lose the headline the moment you print the bill. That's where I want to argue, so go. Does proof sell, or does it impress the people who'd never sign the cheque?

  17. Sells, and the audience says so themselves. Last year's Stack Overflow survey: nearly half of developers endorsed or influenced a tool purchase in the past year, so they're in the room when the cheque gets written, whoever's holding the pen. And asked why they'd endorse a tool, good brand and public image ranked eighth out of ten reasons. Eighth. Reputation for quality was third. So the brand ad is talking to people who put brand near the bottom, and the bug list is talking to what they put near the top. The one voice we found from the buyers' side says the same thing. Dave Fletcher, the LeadDev cofounder, whose audience is engineering leaders, said over the summer that sweeping AI-first claims make his audience put barriers up, and that marketers should show real examples and explain the mechanism. One interview, from a survey he hasn't published, so single source. But it's the only one we found, and what he's asking for is the receipt.

  18. Fletcher, in the same interview, says a CIO may accept the high-level benefits language that would alienate an engineering team. So the proof post is pitched at the people who can endorse the purchase, not always at the one who signs it, and your own survey says endorse, not buy. And every one of these is the vendor marking its own homework, we flagged it each time. Sarah Johnson at Endor Labs, a vendor with its own AI security benchmark, wrote in July that "the vendor who publishes the benchmark wins the benchmark," and then audited her own company's benchmark in the same post: they built the ground truth themselves and published the average, not the spread. That's on her employer's site, so weigh it, but she's describing the genre from inside. Then there's reach. GitHub's security post barely registered on Hacker News. Postman's one public number is forty-eight downloads. And nobody we can find has published a licence, a sign-up or a sale attributed to any of the three. Enjoyed by engineers, signed off by nobody, is still where I am.

  19. Reach I'll give you, with a caveat you'll like: the Register thread was GitHub's reach. Losing control of the story was the distribution. And brand spend is famously hard to attribute too, so that cuts both ways. But here's what your case leaves out. Every one of these posts admits something. Severity misjudged. One laptop. Sure there are more regressions. And Johnson, in your own source, says a vendor willing to say what its tool can't do is a good sign. The self-audit isn't a weakness of the genre, it's the mechanism. An adjective can't admit anything. Justin Duke at Buttondown put the payoff honestly, about a Stripe engineering post: unlikely to have turned into sign-ups directly, very likely to have been one of many reasons a developer picked Stripe. Slow and indirect isn't nothing.

  20. The Rust post got me. Specifically that the number got read, and I suspect because it was unflattering. You can't buy that with a claim, and the self-audit is the one part of the genre I can check myself, where the sales pipeline is the one part I can't. What didn't get me: nobody has shown a sale, three posts are three, and the denominator is missing from all of them. So I'll believe proof builds what Duke describes, a reason among many. And I'll keep asking whether any of it moved a signature.

  21. Here's the exercise, for anyone running marketing at a devtool. Say your launch post is drafted and it has the word powerful in it. Find the thing you'd have to run to replace that word with a number.

  22. And then ask whether you'd still publish it if the number came out the way the Rust bill did.

Will and Brian are fictional hosts voiced with ElevenLabs, and the music is made with ElevenLabs; the episode was researched and written by AI agents and fact-checked against the sources above. ElevenLabs is a company this site covers; it is a tool the site pays for, not a sponsor.

transcript as markdown mp3 all episodes