Definition
As synthetic data refers to artificially generated data with analytical value, synthetic data advantage refers to mean the benefits that the firms obtain from generating and using such artificially created datasets that replicate or simulate the properties of real-world data. Since they are not captured or taken from real-world interactions, they are highly useful in filling the gaps where real-world data is scarce, sensitive or incomplete1, Such usage also helps the firms with privacy compliances and can be applied across sectors where real-world data handling is prone to risks.
Commentary
Origin of the term
The term ‘synthetic data’ was coined by Donald Rubin in a 1993 article titled “Discussion: Statistical Disclosure Limitation,"2 published in The Journal of Official Statistics. The paper proposed releasing artificially constructed microdata in place of actual microdata for statistical analysis to maintain confidentiality. Besides privacy, the concept holds immense relevance today in AI training, as such simulated data can be used to train the AI models to improve them, without compromising real-world data integrity. The concept is also referred to as Simulated Data Advantage or Fabricated Data Advantage and is particularly relevant in sectors such as Healthcare3, Drones, Consumer electronics, Fraud Detection, Autonomous Cars among others.
Operation in Practice
Synthetic data functions by allowing firms to create data for situations that may not exist in real-world datasets. This is especially useful for reducing the cost and time involved in data collection, cleaning and storage and can be of immense help in edge cases4 such rare road conditions for autonomous vehicles, unusual fraud patterns in finance, or underrepresented medical scenarios in healthcare. Since it can be automatically labelled and may include rare but important case scenarios, it can at times be of greater use than even the real-world data for training AI systems. However, care must be taken to ensure that the quality of the base data and the generation method is such that the risks of contamination and bias are avoided while training foundational models. Relevance vis-a-vis Competition Law
It is important from an antitrust lens because of the potential two-sided effect on data-based market power. While it may lower entry barriers by reducing dependence on large proprietary datasets5 enabling smaller firms to generate useful training data and compete with data-rich incumbents, it may also reinforce and strengthen the data-rich incumbents in the market who have better quality data, computational and technical capacity to generate better-quality synthetic data. This raises significant antitrust concerns regarding data-driven-foreclosure, particularly when dominant firms refuse to grant competitors access to ‘essential data’6, which cannot be plausibly replicated by firms with lower market power.
Michal S Gal and Orla Lynskey, ‘Synthetic Data: Legal Implications of the Data-Generation Revolution’ (2024) 109 Iowa L Rev 1087.↩︎
Donald B Rubin, ‘Statistical Disclosure Limitation’ (1993) 9(2) Journal of Official Statistics 461.↩︎
Arman Koul, Deborah Duran and Tina Hernandez-Boussard, ‘Synthetic Data, Synthetic Trust: Navigating Data Challenges in the Digital Revolution’ (2025) 7(11) Lancet Digit Health 10092↩︎
‘Synthetic Data and Its Role in the World of Ai’ (Appen, 23 July 2026) <https://www.appen.com/blog/synthetic-data-and-its-role-in-the-world-of-ai> accessed 9 August 2026↩︎
“Synthetic Data and Its Role in the World of AI’ (n 3).↩︎
Zachary Abrahamson, 'Essential Data' (2014) 124 Yale LJ 867.↩︎


