GAICC AI Conference & Awards 2026 "Governing the Future – Building Responsible, Safe and Human-centric AI"

Nvidia Video Scraping Case

Nvidia’s Video Scraping Case: What It Teaches About AI Data Provenance

In 2024, one of the most valuable companies in the world was reportedly pulling down the equivalent of 80 years of video from the internet every single day, without asking anyone’s permission first. When employees inside the company raised questions about whether that was legal, they were told it was simply an executive call.

That is the core of the Nvidia Cosmos scraping story, and it is one of the clearest examples in recent memory of what happens when an organization has the right instincts inside its own workforce but no governance structure to act on them. Employees asked the right questions. The system around them failed to answer.

This isn’t just a story about one chipmaker’s video pipeline. It is a preview of the kind of legal and reputational exposure that any organization building or buying AI systems now has to plan for, and a useful case study for what a functioning data governance program actually looks like in practice.

Watch the full case breakdown below, or keep reading for the regulatory context and the practical controls this incident points to.

What Nvidia’s Cosmos Project Reportedly Did

According to internal Slack messages, emails, and documents obtained by 404 Media from a former employee, Nvidia was building an internal video foundation model, codenamed Cosmos, meant to power its Omniverse 3D world generator, its self-driving car systems, and its “digital human” products. Ming-Yu Liu, an Nvidia vice president of research and a Cosmos project leader, described the goal internally as building a pipeline that could produce “a human lifetime visual experience worth of training data per day.”

To hit that target, the reporting describes a setup that ran an open-source video downloader across twenty to thirty virtual machines on Amazon Web Services, rotating IP addresses so YouTube’s blocking systems could not catch up. The targets were not limited to YouTube. Employees also discussed pulling content from Netflix and other platforms. Netflix later told 404 Media it had no content agreement with Nvidia and that scraping violates its terms of service. YouTube has separately stated that unauthorized scraping of its platform is against the rules.

What makes this case instructive rather than just another scraping story is what happened when employees pushed back internally. Workers reportedly asked whether it was appropriate to use academic datasets licensed only for non-commercial research inside a commercial product, and whether legal had signed off on the broader scraping effort. The answer, according to the leaked messages, was blunt: “This is an executive decision.” No documented legal risk assessment. No escalation to a governance body. Just a green light from leadership.

It’s worth noting Nvidia has taken a different approach elsewhere. The company has licensing arrangements with Shutterstock and Getty Images for other generative projects, which shows the organization understands licensed data pipelines exist and are workable. That makes the Cosmos decision look less like a resourcing gap and more like a deliberate choice to skip the harder, slower path.

Nvidia has maintained that its practices are compliant with copyright law, telling 404 Media it respects the rights of content creators and that its research is in full compliance with the letter and spirit of copyright law. The company argues copyright protects specific expression, not the underlying facts or data a model might learn from it.

Where the Legal Fight Over Training Data Actually Stands

The video is right not to predict how courts will rule on this. But by mid-2026, the picture has gotten considerably clearer, and it cuts against the idea that “we didn’t get sued yet” is a viable governance strategy.

The biggest signal came out of Bartz v. Anthropic. In June 2025, the court found that training an AI model on legally purchased books was fair use, even calling it “transformative, spectacularly so.” But the same ruling drew a hard line: keeping pirated copies of books downloaded from shadow libraries, regardless of what they were eventually used for, was not fair use. That distinction turned into a $1.5 billion settlement, with final approval in mid-2026 and payouts running to roughly $3,000 per covered work.

A parallel case, Thomson Reuters v. Ross Intelligence, went the other way on different facts. A federal court found that training an AI legal research tool on Thomson Reuters’ proprietary Westlaw headnotes was not fair use, since the underlying material was original, protected content rather than a public dataset. That case is now before the Third Circuit, and its outcome is expected to shape how appellate courts treat AI training claims more broadly.

The pattern across these rulings is consistent: courts are not asking “is AI training inherently legal or illegal.” They are asking “how was this specific data acquired, and does the model’s output compete with the market for the original work.” Provenance, in other words, is doing most of the legal work, not the training itself.

For a company like Nvidia running an unlicensed, high-volume scraping operation against platforms whose terms of service explicitly prohibit it, that emerging standard is not favorable ground.

Why “An Executive Decision” Is Not a Governance Control

The failure inside the Cosmos project was not a knowledge failure. Employees identified the exact risk: licensing terms, platform rules, legal exposure. The failure was structural. There was no defined path for that concern to reach a legal or governance function, get assessed against real criteria, and produce a documented decision.

This is precisely the gap that frameworks like the NIST AI Risk Management Framework are built to close. The RMF’s Map function calls for organizations to document training data sources, evaluation outcomes, and the accountability structure behind those decisions, and to be able to produce that documentation on request rather than reconstruct it after the fact. NIST’s generative AI profile specifically flags copyright and data-provenance uncertainty as a distinct risk category that needs to be evaluated at the design stage, not discovered later in litigation.

The distinction matters because “an executive decision” and “a documented risk assessment” can produce the exact same green light and still represent completely different levels of organizational exposure. One creates an audit trail that shows due diligence. The other creates a Slack message that becomes evidence in a lawsuit or leaks to a journalist.

Building a Data Provenance Program That Actually Holds Up

A workable data provenance program does not need to be complicated, but it does need to exist before a high-volume data pipeline goes into production, not after a reporter starts asking questions. Three controls do most of the work.

A data provenance record for every training source. Before any dataset enters a training pipeline, document where it came from, what license or terms of use apply, and the specific legal basis for using it commercially. If a dataset is licensed for non-commercial research, it does not get folded into a revenue-generating product, no matter how useful it is. The “Datasheets for Datasets” approach, now widely referenced in AI governance literature, is a practical template for this: origin, collection method, licensing terms, and known limitations, captured at the point the dataset is created or acquired.

A real escalation path, not a rubber stamp. When an employee raises a legal or ethical concern about a data source, it needs to route to a legal or governance function that performs an actual assessment and documents the outcome, approved or rejected, with reasoning. A verbal “we have clearance” from a project lead is not an assessment. It is the absence of one.

Sign-off before scale, not after. A pipeline capable of downloading eighty years of video a day is not a small internal experiment; it is a high-risk data operation. Those need a documented risk review and formal approval before they scale up, proportional to the volume and sensitivity of what is being collected.

Use this as a quick gut check for your own organization:

  • Can you produce, for any dataset currently in a production model, its source, license, and legal basis for use, on short notice?
  • Is there a named function, not an individual executive, responsible for approving high-risk data acquisition before it scales?
  • Are legal or ethical objections raised by staff logged and resolved with a documented decision, rather than a verbal assurance?
  • Does your risk assessment process distinguish between a small pilot dataset and a production-scale pipeline?

If any of those get a “no,” that’s the gap the next audit or the next leaked Slack thread will find first.

The Real Lesson Isn’t About Nvidia

The companies that will hold up under the coming wave of AI copyright litigation are not the ones that never touch a legal gray area. Nearly every foundation model has been trained on some data whose provenance is contested. The companies that hold up will be the ones that can show their work: a documented source, a real assessment, a decision that was made deliberately rather than assumed.

That is what AI governance actually is. It is the difference between “we’re confident this is fine” and being able to prove, with a paper trail, exactly why.

If you want to build that capability inside your own organization, GAICC’s Certified AI Law and Compliance Professional Certification covers exactly this kind of data provenance, risk assessment and escalation framework in depth, and it’s one of the fastest-growing credential tracks in AI governance right now for good reason.

The same governance gap can appear when employees use AI tools without clear controls over what data can be submitted, as the Samsung ChatGPT data leak demonstrates.

Share it :
About the Author

Dr Faiz Rasool

Director at the Global AI Certification Council (GAICC) and PM Training School

A globally certified instructor in ISO/IEC, PMI®, TOGAF®, SAFe®, and Scrum.org disciplines. With over three years’ hands-on experience in ISO/IEC 42001 AI governance, he delivers training and consulting across New Zealand, Australia, Malaysia, the Philippines, and the UAE, combining high-end credentials with practical, real-world expertise and global reach.

About the Author

Latha Karthigaa

Head of AI Governance at the Global AI Certification Council (GAICC)

A PhD-qualified AI governance leader in Software Engineering from the University of Auckland, she brings hands-on experience founding and exiting AI companies, and leading real-world AI solutions for finance and legal firms across the USA, UK, Australia, and New Zealand, combining governance, risk, compliance, and commercial expertise.

Start Your ISO/IEC 42001 Lead Implementer Training Today

4.8 / 5.0 Rating

Recent Post