TechCrunch Minute: Over 100K YouTube videos have been scraped to train AI for Apple, Nvidia

Mobile

What do MrBeast, John Oliver and the Wall Street Journal have in common? The transcripts of their YouTube videos have been scraped to train the AI used by companies like Anthropic, Nvidia, Apple and Salesforce.

An investigation from Wired and Proof News found that this dataset, which is called YouTube Subtitles, contains transcripts from over 173,000 YouTube videos on more than 48,000 different channels.

This AI scraping is a problem all across the tech industry. Artist and founder of the app Cara, Jingna Zhang, has tried to protect artists by building a social platform that won’t sell them out. And the University of Chicago is working on Nightshade, which can “poison” an image to limit what an AI can glean from it. 

But is there really any way for creators to protect themselves from being next? More on the TechCrunch Minute.

Products You May Like

Articles You May Like

Apply to Speak at TechCrunch Sessions: AI before the deadline
Fetii’s group rideshare app for young people attracts funding from Mark Cuban, YC
Thinking Machines Lab is ex-OpenAI CTO Mira Murati’s new startup
Amazon, Microsoft, and Exxon want to make scandal-plagued carbon markets more trustworthy
Lingo.dev is an app localization engine for developers

Leave a Reply

Your email address will not be published. Required fields are marked *