Martin Källström
knowledge / projects

Twingly Blog Search

Twingly Blog Search is the core blog search engine indexing the global blogosphere — the product that gave Twingly its name. It was born from a pivot: in spring 2006 Martin Källström's team scrapped a nearly-finished Blocket-style classifieds clone and, watching blogging explode the way it had around the 2004 US election, rebuilt the company around indexing blogs worldwide instead. Over the next several years it grew into a spam-free, dozens-of-languages search engine covering tens of millions of blogs, a widget layer connecting bloggers to newspapers like Dagens Nyheter and Svenska Dagbladet, and eventually a mostly API-driven data business serving over 25 million searches a month.

The pivot that created Twingly. Martin and the Primelabs team had a Blocket-like classifieds service nearly finished when, during spring 2006, blogging started dominating the conversation — echoing what they'd seen in US election coverage. They shelved the almost-complete project and built a new one from scratch instead, carrying over technical groundwork (language handling, geographic ambition) from the abandoned attempt. The team grew from four people to nine over the following two years.

"Vi hade utvecklat en Blockit-site nästan klar... Vi tog ett jättesvårt beslut och helt enkelt la hela det stora projektet på hyllan och började istället utveckla en ny lösning" ("We'd built a Blocket-like site, almost finished... We made a very difficult decision and simply shelved the whole project, starting instead on a new solution.") ▶ 5:25

Product & technology

Spam-free by whitelist, not blacklist. Rather than index every blog and filter spam afterward, Twingly built a curated whitelist: starting from a hand-vetted seed of roughly 6,000 quality blogs from around the world, then spidering outward along outbound links on the assumption that legitimate blogs don't link to spam blogs. Each newly approved blog recorded which existing blog had vouched for it, forming an approval chain — so when a spam cluster was later discovered, cutting that chain could remove thousands of spam blogs in a single action. By June 2008 the whitelist held "hundreds of thousands" of blogs.

"Vi började med ett seed på ungefär 6000 bloggar från hela världen som vi stoppar in i en whitelist... Så kan vi klippa av dem med en enda knapptryckning. Och plötsligt har vi en algoritm som är ett verktyg som gör att vi kan vinna kampen mot spam." ("We started with a seed of about 6,000 blogs worldwide put into a whitelist... we can cut them off with a single button press. Suddenly we have an algorithm that lets us win the fight against spam.") ▶ 84:01 · TechCrunch ↗ · TechCrunch ↗

Ranking: TwinglyRank. Results were ranked by TwinglyRank, described as "a combination of keyword relevancy, number of inlinks, number of user recommendations, publishing date and time and some secret sauce." TechCrunch ↗

Language coverage. The engine detected and supported roughly 55–60 languages by mid-2007; by its April 2008 public launch it offered accurate language-specific search in 29 languages while tracking another 31 not yet reliable enough to deploy. ▶ 0:57 · TechCrunch ↗

Scale, by April 2008. The index held roughly 110 million blog posts and 170 million interlinked relationships on about half a terabyte of data. It reached 30 million indexed blogs, concentrated in Europe, by launch that June. ▶ 9:46 · TechCrunch ↗ · gHacks ↗

Architecture. Twingly ran on Sphinx, chosen over alternatives like Lucene specifically for its built-in scalability, on top of a horizontally-scaled MySQL layer: four database servers, each holding a quarter of 128 tables, kept under roughly 10 million rows per table for manageability. A queue-based task system let a pool of servers pull indexing work independently, and the index itself was split by time — a matrix of monthly indices plus a rolling 24-hour "fresh" index — with old data reindexed weekly and the last day's data reindexed every three minutes. Martin compared that cadence to Google's old quarterly "Google Dance" full-index rebuild, calling Twingly's three-minute refresh a much harder engineering problem on a smaller dataset. ▶ 71:45 · ▶ 47:09 · ▶ 76:26 · ▶ 31:01

Getting blogs into the index. Beyond the whitelist crawl, Twingly ingested content through manual pings (bloggers self-registering, including non-detected foreign-language blogs), automatic pings from blog-hosting platforms like blog.se, Expressen and Aftonbladet, a discovery crawler, and an internal "Roundhouse Ping" that continuously cycled through important blogs to refetch fresh RSS. Users could also report indexing errors directly by email. ▶ 1:23 · ▶ 62:14 · ▶ 3:31

Media partnerships & business model

Bridging bloggers and newspapers. Twingly's core mechanic let bloggers work alongside traditional media: a blogger writes about and links to a newspaper article, pings Twingly, and is then linked back from the newspaper's site — giving independent writers visibility on established mastheads and giving newspapers a feed of relevant blog reaction. Dagens Nyheter and Svenska Dagbladet were early adopters, later joined by Dagen and TV4's Idol site, and at least one South African newspaper; by 2007 Twingly was scanning roughly 250,000 blog posts a day to surface content relevant to newspaper feeds. ▶ 0:01 · ▶ 1:02 · ▶ 2:45 · Björn Rutgersson ↗

Nyteknik summarized the resulting business model plainly: Twingly earns its revenue as a bridge builder, connecting the blog world with traditional companies and organizations — letting a company link from its own site to blogs discussing its products, services, or articles. Källström also described plans to strike revenue-sharing deals with large content/news sites for showing related blog posts, a model TechCrunch noted Sphere and Technorati had already proven out. Ny Teknik ↗ · TechCrunch ↗

Becoming an API business. By 2011–12, the majority of Twingly's 25-million-plus monthly searches were happening through its API rather than the consumer website — pushing Twingly to acquire Bloggportalen from Aftonbladet to add more consumer-facing features alongside the API business. Bloggportalen was redesigned onto Twingly's own servers but kept as a distinct product from the main Blog Search. Arctic Startup ↗

Bloggbävningen case study (Sept 2008). Martin presented a Twingly-powered sentiment report at a Stockholm seminar analyzing how blog opinion of Swedish politicians shifted around "Bloggbävningen" (the parliamentary revolt over FRA surveillance legislation) — an early public demonstration of Twingly's real-time sentiment analysis. The data showed defense minister Sten Tolgfors's reputation swinging sharply in the blogosphere while PM Fredrik Reinfeldt was comparatively untouched, and Annie Johansson and Fredrik Fälldin both taking notable reputation hits. ▶ 0:51 · ▶ 1:17

Features at public launch (June 2008)

Twingly opened to the public on June 12, 2008, after a private beta of roughly 3,000–3,500 testers who could propose and vote on features before launch. TechCrunch ↗ · The Next Web ↗ · Teknikfreak ↗

The launch package, per contemporary coverage:

Reviewers judged the result favorably against the incumbents: gHacks called Twingly's Web 2.0-style interface "defiantly the superior" of Google Blog Search and Technorati, and The Next Web wrote that it "compares favourably to its better known rivals." gHacks ↗ · The Next Web ↗

Worth remembering