Martin Källström
knowledge / stories

Clone the repo before you believe the demo

Martin's rule for AI announcements is short: a demo is a claim, and the only trustworthy evidence is running the thing yourself. "Like experiencing things yourself. I think that's, that's the, that's the only like benchmark that is trustworthy." ▶ 30:54 He has paid for it twice: at Narrative, where the hype was his own, and in 2024, when he believed a competitor's voice demo and abandoned a project two months from beta. The cases below come mostly from the Co-creating with AI podcast with Rasmus Adler Wahlberg. One episode is already summarised in AI adoption & society; this page goes wider, and separates the demos Martin forgives from the ones he does not.

Five takes and a repo

The rule predates the cases. In an autumn 2023 episode on voice AI, Martin warned: "It's very easy to produce really nice demos of stuff because then you can, you can, you can do 5 takes and just publish the best one." A demo "is a very different thing" from "a solid, robust solution that is ready for the real world". ▶ 23:43 The older wound is from Narrative, told on a Breakit podcast about bankruptcy: with that much positive press, "vi insåg inte hur mycket omvärlden skulle påverka vårt företag" (we didn't realise how much the outside world would affect our company), and "det gjorde att vi inte höll tillräckligt bra koll på vad som hände på resten av marknaden" (it meant we didn't keep good enough track of what was happening in the rest of the market). ▶ 18:43 ▶ 19:09 That is the growth-strategy version of the rule: hype, including your own, is a substitute for looking.

The method in the title is literal. In summer 2024 a paper promised eight per cent more robust code generation from a few-shot technique, exactly what Martin was building. He cloned the repository and ran the tests; they passed. Then he read the tests and found two layers, a solid one and one underneath, "which was sort of cleverly hidden", where the model was handed the expected results: "The few-shot learnings was actually just one shot, and it was their solution." ▶ 6:50 ▶ 7:41 He looked "back and forth, up and down, and and sideways, and there was no other way to interpret it than pure deception", from what seemed a serious Singapore–Harvard collaboration. ▶ 8:32 Nothing else happened: "I discovered the deception and it's never been publicized anywhere because it's so minor". ▶ 6:23

Two kinds of fake

Martin draws a line most hype discourse does not. On one side is Devin, marketed in early 2024 as the first reliable AI software engineer: "it turns out that the demo videos were built only on cherry-picked examples", and the flagship example, an Upwork job supposedly solved for a real client, was worse. "that person that in the demo video got their problem solved came forwards like a couple of weeks ago and said, I actually didn't get a solution to it". His verdict was two words, that's just lazy; Rasmus called it bad demo faking. ▶ 22:20 ▶ 22:22 Rasmus later supplied the number: about 14% on SWE-Bench. ▶ 20:41

On the other side is Google's Gemini 1.0 launch video, which "was faked, faked by some marketing agency like that. That was just supposing this is how AI should work. And they made a too good video." Martin's reading is generous: "it's fake, but it also, as you can see it as science fiction and inspiration for something that is to come in the near future". ▶ 24:04 He had said as much when the video appeared, betting "there's at least 20 teams trying to implement, like, or not trying, but actually working on implementing what was seen in that video right now". ▶ 25:49 Five months later OpenAI shipped roughly what Google had staged, the day before Google I/O; Google's launch was "where everybody was disappointed that they overhyped it". ▶ 16:28 The distinction: fiction that announces itself as a target is forgivable; a demo that claims a result it did not get is deception.

Figure's humanoid robot sat between the two. Rasmus was impressed that "it's not manipulated. Like, of course they practiced it beforehand, but it's not cut." ▶ 1:46 Martin granted that the language-to-motion layer was real, "research that has only been going on for the past year" ▶ 5:30, but fixed on the robot casually dropping an apple for the human to catch: "that's not a robust plan for the robot to come up with. I'm just going to drop it in the air and hoping the human catches it." ▶ 6:46 A month later Rasmus filed it with Gemini: "probably is pre-programmed and fake", but likely to happen for real as competence rises. ▶ 24:26

Demos with a shelf life

Some demos are not faked so much as immortal. Rasmus noted that the canonical agent demo, "Google Assistant, you know, calling up a hairdresser and booking a time for you", is "even pre-LLM", and that ChatGPT's first plugin demo was the same trip-planning idea. ▶ 17:47 When GPT-4 gained audio and function calling he was back at the salon, "the Google Assistant kind of demo from a few years back". ▶ 26:15 A capability that has been coming soon across two technology generations is hard, not merely unshipped.

The voice demos got the same treatment. When GPT-4o's voice was announced, Rasmus reported that Bill Gurley, the Benchmark founder, had used it and said "it wasn't as good as in the demos", then added the line that could headline this page: "Well-rehearsed demos. That's the hallmark of AI launches." ▶ 3:31

And Martin applied the rule to a tool he loves. In September 2024 he admitted Cursor Composer was something "I was really honestly hyped about when we did the last episode recording". ▶ 17:17 Two weeks of real use later he "started bumping into these really, like, immature behaviors": it "sometimes takes a whole file, deletes the existing working code, and replaces it with a placeholder saying, here, here's where your existing code should go", and across four or five files it introduced tiny flaws that made review more tedious, not less. ▶ 18:09 ▶ 19:00 ▶ 20:22 His summary of every YouTube demo: "from minute 3 and onwards, you're going to have problems, but the first 3 minutes is awesome". ▶ 20:47 Rasmus sorted real from hyped, "What is real is the UX innovation. What is real is getting started quickly." ▶ 21:27 That is the rule working: not cynicism, a two-week test.

The price of believing

The rule has a scar. Through the winter of 2023–24 Martin built, alone, "a good unique take on conversational AI", and took a demo to San Francisco in April for a voice hackathon. ▶ 11:19 Well-funded teams were describing their projects in his words, and "when I came home, I felt that some of the energy and excitement in my own work had been drained by that trip to San Francisco". ▶ 12:37 Then OpenAI showed its advanced voice demo, "the proper way of doing conversational AI", audio tokens straight into and out of the model. ▶ 13:27 In the moment it felt like liberation, "it's such a huge relief for me to get this", OpenAI providing what he had been assembling from open-source parts. ▶ 9:40 ▶ 10:05 He shut the project down.

Four months later he called it "the worst decision I made this year based on pure hype". ▶ 11:08 "since then, to my deep regret and frustration, none of those companies, including OpenAI, has delivered on that hype", and "I was probably something like 2 months away from a beta when I made a pivot". ▶ 14:19 The conclusion is the rule in the first person: "I should not have gone there. I should just stay home and finish my own code." ▶ 14:45 There was no repo to clone; a demo with no code to run is exactly the kind you cannot verify, and therefore should not pivot on.

Why the demos keep coming

Martin does not treat this as a moral failing of particular companies. "every startup that is VC-funded needs to build the vision before they build the product"; the VC industry demands a hype cycle. ▶ 28:33 Public companies have their version: when Jensen Huang promised inference a million times faster, Martin guessed Groq was beating NVIDIA on speed and "he just needs to keep his shareholders happy in some way". ▶ 6:19 Credit gets hyped as well as capability: "I believe Groq basically took just Flux and then launched it as their own", and got the applause for Black Forest Labs' model. ▶ 2:50 Rasmus added the Reflection fine-tune: "They just wrapped Claude in an API and told it not to be Claude." ▶ 4:36 His definition is the one Martin operates by: "hype by definition is, um, not reality"; Flux is not hype, Flux generating movies would be. ▶ 15:12

The rule follows from the incentives. If everyone must over-promise to be funded, seen or valued, the demo carries no information about the product. The repository, the two-week trial, the pineapple question, the client who did or did not get a solution: those do. See Failure & resilience for the other pivot he regrets, and Old fears, new machines for the historical instinct that makes him generous about science fiction and hard on fraud.

Worth remembering