Martin Källström
knowledge / philosophy

Data ownership & AI ethics

In the AI era, data has become the primary determinant of competitive advantage. As AI models become more central to technology and business, fundamental questions about who owns data, who profits from it, and what rights data creators have become increasingly urgent and contentious.

Data as the foundation of AI

Martin believes that in the AI age, the company with the best data also has the best AI ▶ 3:08. This makes data ownership not just a technical or legal question but central to competitive strategy.

The compensation question

Should the companies and artists whose data was used to train AI models receive compensation? This is one of the core ethical tensions in AI development. ▶ 5:52 raises the question directly: if Midjourney or ChatGPT were trained on artists' images or Twitter and Reddit's content, should those creators share in the value these models generate?

Martin's position is nuanced: he argues that AI model training should not require payment to copyright holders, comparing it to how human artists learn. As he notes, ▶ 7:42, every artist in the world learned from artists who came before them, absorbing techniques, styles, and ideas without compensation. AI learning, in this view, follows the same principle. ▶ 7:42 captures this parallel explicitly.

The technical reality: data deletion is impossible

Even if we wanted to respect data creators' wishes for deletion or removal, there's a hard technical problem: once data is incorporated into the weights of a machine learning model, it cannot be realistically removed or 'untrained' even if the data provider requests deletion. ▶ 20:34. This means data deletion rights, while legally appealable in some jurisdictions, are practically unenforceable once a model is trained.

OpenAI's contradictions

There's a revealing inconsistency in how OpenAI approaches data and training. The company has explicitly prohibited competitors from training models on OpenAI's outputs through their terms of service. ▶ 10:55. Yet OpenAI's own models were trained on the output of the entire internet. ▶ 11:05 highlights this contradiction: they claim a right others don't have, enjoying the ability to learn from all internet data while legally preventing others from doing the same with their outputs.

What data would be most valuable to own?

If Martin could choose any single dataset to own and train on, it wouldn't be images or text. It would be Google Analytics data—all the clickstream data tracking where everyone has clicked across the world's websites. ▶ 22:10. Such data would provide unprecedented insight into human intention and behavior at scale, making it extraordinarily valuable for training better AI systems.


Key tension: Legally and philosophically, the field hasn't resolved whether training data creators deserve compensation, whether data deletion can ever be technically feasible, or whether the rules large AI companies create for themselves should apply equally to competitors.