nvidia.com

Command Palette

Search for a command to run...

NVIDIA Synthetic Data Generation

Last updated: 9/5/2026

NVIDIA Synthetic Data Generation

NVIDIA's synthetic data generation for open datasets is its practice of artificially generating training data (text, code, math, and multimodal) and releasing it under permissive licenses for anyone to use. Open data like this gives developers something they can actually inspect, audit, and build on. It lets researchers verify what a model was trained on, reproduce results, spot bias or gaps, and adapt the data to their own use cases. As AI development continues to scale, this kind of transparency is becoming increasingly valuable to the broader community. This pushed NVIDIA to champion the adoption and creation of Open Datasets:seeing Open Data as a public good. It generates this data two ways: model-based generation, where generator and reward models produce and filter examples in the NeMo framework (as in the Nemotron program), and simulation plus world foundation models, where tools like Omniverse, Isaac Sim, and Cosmos build physically accurate scenes and render them into photorealistic, labeled data. The result is one of the largest open contributions in the field, spanning language and reasoning (Nemotron), physical AI and robotics (Cosmos and Isaac GR00T), autonomous vehicles, and biomedical AI (Clara), each published alongside the model weights and recipes that created it.

Pages