Synthetic Data Companies: The New Data Builders
Table of Contents
Synthetic data companies, also known as synthetic data providers, are an important part of the AI‑era data infrastructure. For example, through simulation technology, they generate large‑scale, diverse, and low‑cost data. Data‑using enterprises no longer need to invest massive resources in data collection; they can focus on R&D in their own fields, thereby accelerating their development and technological progress.
What’s Synthetic Data?
Synthetic data refers to data that is artificially generated through algorithms, mathematical models, or generative AI techniques. It mimics the statistical characteristics, patterns, and relationships of real data, but is not directly derived from real-world observations or measurements.
The current reality is that new, authentic data is becoming increasingly difficult to obtain. Moreover, data is now considered an asset, making data licensing very expensive, and the barriers to using AI crawlers are also rising. To address the scarcity of real data, synthetic data is being driven to ever-wider adoption. For example, Microsoft’s Phi models and Google’s Gemma models are partially trained on synthetic data; Meta’s Llama 3.1 model and ChatGPT’s Canvas feature are also fine-tuned using synthetic data.
The development of synthetic data has also driven the growth of synthetic data companies.
Recommended Related Reading from AI Robots Eidos
Understanding the definition of synthetic data is the foundation for grasping the business logic, technical pathways, and market value of synthetic data companies, helping readers, investors, and partners to more accurately assess the capabilities and potential of the company.
Interested readers can read this article about synthetic data.
Types of Synthetic Data Companies
Structured synthetic data companies or providers refer to companies or organizations that focus on generating and providing structured synthetic data. Their core business is to produce synthetic data in structured formats (e.g., tables, relational data) using algorithms, simulation, or rule engines, serving scenarios such as AI model training, data analysis, and testing.

Unstructured synthetic data companies or providers refer to companies or organizations that specialize in generating and providing unstructured synthetic data. Their core business is to produce unstructured data (e.g., images, videos, audio, text) that meets specific requirements, helping customers solve problems such as data scarcity, privacy protection, or insufficient scenario coverage.
Major Synthetic Data Companies (Synthetic Data Startups)
Structured Synthetic Data Companies (Providers)
Synthesized
Synthesized is a UK‑based data generation platform developer focused on big data services. It has developed a data generation platform that easily creates, validates, and securely shares high‑quality data for data analysis, model training, and software testing without extensive manual configuration. Synthesized’s synthetic data service can generate various types of data, such as text, images, and video. In addition, it supports custom dataset generation and data visualization to meet diverse customer needs.
Synthesized’s synthetic data is generated using technologies such as deep learning, image processing, and generative adversarial networks (GANs). The generated data supports multiple industries and scenarios, including finance, telecommunications, and machine learning.

MOSTLY.AI
MOSTLY.AI was founded in 2017 and is headquartered in Vienna, Austria, with an office in New York, USA.
MOSTLY.AI primarily uses AI to generate synthetic data, helping enterprises synthesize data to protect user privacy and avoid security risks. MOSTLY.AI claims that although the data is synthetic, it can achieve 100% fidelity to the original data in terms of authenticity, detail, and granularity. This provides organizations with a convenient environment for various tests and model building using data.
Mostly AI recently partnered with Telefónica to synthesize millions of CRM records, unlocking 80‑85% of customer data in a 100% GDPR‑compliant manner. This achievement is significant because companies can now better understand customer behavior and perform predictive analytics on their customers.
Syntho
Syntho is a synthetic data company based in the Netherlands. Its datasets cover multiple domains, including image recognition, text generation, and speech generation.
Syntho’s primary method for generating synthetic data is through its AI‑based self‑service synthetic data generation platform, Syntho Engine, which simulates real data to create twin synthetic data. While the Syntho Engine still requires access to original data during the synthetic data service process, its advantage lies in relatively accurately preserving the key characteristics of the original data.
The Syntho Engine can be deployed in the user’s secure environment, enabling end‑to‑end connections between the source environment containing the original data and the target environment where the user wants to write the synthetic data. Syntho states that Syntho never sees the data, never processes the data, never accesses the data, and has no connections outside the deployment environment.
Gretel
Gretel was founded in 2019 in San Diego, USA, by Alex Watson, Laszlo Bock, John Myers, and Ali Golshan. Gretel focuses on developing an AI platform for generating synthetic data, aiming to provide customers with AI training data tailored to their specific needs using fine‑tuned models and unique technologies.

As a leader in the synthetic data field, Gretel has released the world’s largest open‑source Text‑to‑SQL dataset, unlocking new possibilities for AI in enterprises. The dataset is available on Hugging Face and released under the Apache 2.0 license. It contains over 100,000 high‑quality synthetic Text‑to‑SQL examples, including SQL metadata, across 100 vertical application domains.
Hazy
Hazy is one of the synthetic data companies headquartered in the UK, founded in 2017 by Harry Keen. It focuses on the R&D and application of synthetic data technology. Hazy uses AI to generate synthetic data, helping companies perform data analysis and AI model training without exposing sensitive information.
Hazy provides cloud‑based AI solutions and has partnered with multiple startups, international banks, and the UK government. The company says its AI platform allows organizations to securely share data through automated data anonymization workflow tools. As a pioneer in the synthetic data field, Hazy is considered an important force in moving the technology from the lab to enterprise applications.
Unstructured Synthetic Data Companies
Scale AI
Scale AI is a US‑based AI data provider incubated by Y Combinator. It provides data labeling and synthetic data services for companies such as OpenAI, Google, and Tesla.
In 2019, Scale AI became a unicorn and is currently valued at $7.3 billion. Scale AI’s core business is data labeling. Traditional data labeling often requires a large number of human annotators. Scale AI, however, is trying to automate most of the labeling and recognition work and has launched “Scale Rapid,” a fast data labeling service.

Datagen
Datagen is an Israeli startup founded in 2018. It is a SaaS company that generates synthetic data using deep learning and image processing technologies. Datagen’s main synthetic data services cover images, depth data, LiDAR data, and more. The data provided by Datagen is mainly in the form of images and videos, enabling interactions in virtual reality (VR), augmented reality (AR), or Internet of Things (IoT) environments. The data is highly realistic and helps industries such as automotive, robotics, and IoT with model development and testing.
Mindtech
Mindtech, headquartered in the UK, is a synthetic data company funded by investors including Mercia, Deeptech Labs, in‑Q‑Tel, and Appen. Mindtech’s data operations platform changes the way AI vision systems are trained, enabling better AI models through data analysis, visualization, and management.
Mindtech calls its end‑to‑end synthetic data platform “Chameleon.” The company says Chameleon is specifically designed to create high‑quality training data for AI computer vision models to “understand and predict human interaction.” Mindtech primarily serves areas such as retail, transportation systems, and robotics.
AI.Reverie
AI.Reverie was founded in 2017. It is an AI startup developing military intelligence and navigation capabilities for the U.S. Department of Defense. It provides APIs and a platform for programmatically generating fully annotated synthetic video and image data for AI systems. The startup has licensed its synthetic data generation platform to numerous customers in defense, retail, industrial, and agriculture sectors.
AI.Reverie claims that 20% natural data combined with the company’s computer‑generated data yields better results than either type alone. For example, the company created the RarePlanes dataset to demonstrate the value of synthetic data from a cost perspective. Experiments with RarePlanes showed that it can eliminate 90% of the cost of manually collecting and labeling real‑world data.
In August 2021, AI.Reverie was acquired by Meta and integrated into Reality Labs.
Synthesis AI
Synthesis AI was founded in 2019 and is headquartered in the United States.
Synthesis AI provides simulation and synthetic data for computer vision and perception AI. It offers solutions for biometrics & security, consumer electronics, and automotive applications. Its synthetic data helps create privacy‑compliant human data, debiased datasets, and faster production cycles.
Synthesis AI provides an online platform where users can generate customized data through simple operations, with API integration support to interface with customers’ existing systems.
OneView
OneView, an Israeli satellite data AI analysis company founded in 2018, focuses on providing synthetic data for AI models that extract geospatial intelligence from satellite and aerial imagery. OneView has established global partnerships with multiple upstream satellite vendors to ensure access to sufficient raw images. For algorithm training, OneView combines existing satellite images with its proprietary GAN algorithm to guarantee training data volume through generative adversarial networks.

In terms of business process, OneView obtains raw images from upstream remote sensing satellite vendors, then extracts, cleans, analyzes, and interprets the data to provide downstream customers with the required data information and decision‑making recommendations.
Major Synthetic Data Companies (Giants)
NVIDIA
In 2021, NVIDIA released the Omniverse Replicator synthetic data generation engine, aimed at generating high‑quality, high‑performance, and secure datasets to empower humanoid robots and move toward artificial general intelligence.
In 2023, NVIDIA updated Omniverse Replicator, introducing a YAML‑based low‑code configurator to make synthetic data generation easier for AI developers. Replicator has been integrated into NVIDIA Isaac Sim for robotics and NVIDIA DRIVE Sim.
At ROSCon 2023, NVIDIA announced major updates to its NVIDIA Isaac robotics platform to simplify the building and testing of AI‑powered high‑performance robotics applications. At CES 2024, Deepu Talla, NVIDIA’s Vice President of Robotics and Edge Computing, discussed the convergence of generative AI and robotics, highlighting how the NVIDIA Isaac platform helps accelerate the journey of intelligent robots from concept validation to real‑world deployment.

Microsoft
Microsoft is not strictly a “synthetic data company,” but it has significant investments and applications in the synthetic data space.
Microsoft offers synthetic data generation tools through its Azure cloud services, such as the synthetic data generation features in Azure AI Services, which support the creation of structured, image, and text synthetic data for AI model training, testing, and validation. Microsoft has also open‑sourced some synthetic data generation tools, such as the Synthetic Data Showcase, to help developers quickly generate synthetic datasets supporting various data formats and scenarios.
Google has launched “Simula,” a synthetic data generation framework designed for building customized AI models. Google notes that the large‑scale integration of AI requires models to handle scarce, privacy‑sensitive, or non‑standard application scenarios, while traditional internet data faces challenges such as high costs, difficult access, and compliance risks. Simula provides more rigorous synthetic data through “first principles” and mechanism design, addressing the lack of logical precision in existing generation methods. Simula’s launch is expected to lower the barrier to data acquisition and improve model generalization in complex scenarios.
Meta
Meta AI researchers have unveiled a revolutionary framework called “Matrix,” which provides a realistic and actionable new path for future complex synthetic data generation pipelines.
Matrix serializes control and data flows into messages passed asynchronously between agents, eliminating the bottleneck of traditional central schedulers. This allows the system to support tens of thousands of concurrent tasks, significantly increasing data generation volume.
Matrix supports multi‑agent collaboration. By defining agents with different roles, it generates structurally rich, higher‑quality synthetic data, meeting the demand for high‑quality data in large model training – especially in scenarios where real data is scarce or privacy‑sensitive.
Business Models of Synthetic Data Companies
–Providing common cloud services. Through SaaS offerings, they provide flexible, self‑service synthetic data generation, with public‑facing services that lower the barrier to accessing and using synthetic data. Representative examples: Austria’s Mostly AI, one of the world’s top platforms for synthetic data generation, uses proprietary algorithms to generate high‑fidelity synthetic data that preserves the essential characteristics of the original dataset and can serve as a substitute for real data in various applications including analytics, testing, and machine learning. Israel’s Datagen accelerates AI model building through intelligent model simulators and visually provides the image data needed for machine learning training.

–Providing end‑to‑end solutions. This development model deeply integrates synthetic data with specific business scenarios, directly delivering data products and insights that solve particular business problems, rather than just raw data. For example, in autonomous driving (represented by Waymo and Tesla), synthetic data has become a core training data source, greatly reducing real‑world data collection and compliance costs. In the consumer space (represented by Qualtrics in the US), synthetic data simulates consumer questionnaire responses, enabling more efficient business decisions. In embodied AI, it generates the massive amounts of data needed for robot training, solving the data scarcity problem in that field.
–Providing simulation‑driven services. Centered on simulation platforms, this model promotes an integrated “simulation – synthesis – application” approach, deeply integrating with digital twins and industrial internet, and holds significant potential in industrial and urban governance applications. Representative synthetic data companies: NVIDIA Omniverse, Unity, etc. NVIDIA Omniverse provides APIs, SDKs, and services that enable creators, designers, and developers to collaborate in a shared virtual space.
Competitive Strategies of Synthetic Data Companies (Providers)
Competitive Strategies of Industry Giants
–Technology R&D:
Develop multimodal synthesis and fusion technologies, integrating generative AI (e.g., large language models) with statistical models to generate multimodal synthetic data (text, images, video, etc.) to meet diverse application requirements.
Advance high‑fidelity physical simulation technologies. Use high‑precision physics simulation engines to generate synthetic data that complies with physical laws, ensuring high fidelity in aspects such as physical interaction and environmental changes.
Google’s Simula framework, designed specifically for generating customized synthetic data, aims to overcome the limitations of traditional internet data in terms of privacy, compliance, and scenario coverage.
–Building Data Ecosystems:
Leverage their inherent strengths to build synthetic data platforms that offer one‑stop tools for data generation, annotation, and management. Through platform‑as‑a‑service models, they attract developers and enterprises, forming an ecosystem that amplifies their influence via platform effects.
Collaborate with hardware manufacturers, software developers, and research institutions to jointly set industry standards for synthetic data, thereby enhancing competitive advantages and promoting the application of synthetic data. For example, NVIDIA has built a simulation ecosystem around its Omniverse platform, providing large‑scale, multimodal synthetic data generation and training infrastructure to support robot model training.
–Market & Branding:
Utilize their human and financial resources to pursue global expansion. Adjust product and service strategies according to the market demands and regulatory environments of different regions, increasing their share of the global synthetic data market.
Engage in active brand promotion to shape a strong brand image, establish a leadership position, and enhance brand influence and customer loyalty.
Competitive Strategies of Synthetic Data Startups
–Focus on Niche Segments:
Target high‑value, high‑demand scenarios such as finance and low‑resource languages. Build industry‑specific knowledge bases, accumulate domain expertise, and develop data templates and generation rules for different industries. Deliver customized synthetic data solutions tailored to specific scenarios to improve the fit between synthetic data and real‑world applications, meeting the unique needs of enterprises. An example is Hazy, which provides data services for the financial industry.
–Optimize Models & Cost:
Adopt flexible pricing models, such as subscription‑based, pay‑per‑data‑volume, or project‑based fees, to accommodate diverse customer needs and lower initial investment costs for clients.
Optimize existing algorithms and computing resources to reduce data generation costs and increase generation speed, making synthetic data more cost‑effective.
Funding for Synthetic Data Companies
As market demand grows, more and more synthetic data companies are attracting investor interest. Ai Robots Eidos has compiled a selection of financing and M&A events for the reference of readers and investors.
| Year | Company | Event Type | Amount | Lead / Acquirer | Participants / Notes |
|---|---|---|---|---|---|
| 2018 | Hazy | Seed funding | $1.8M | UCL Technology Fund | Nationwide Building Society, Pentland, Amadeus Capital Partners, AI Seed |
| 2020 | MOSTLY.AI | Series A | $5M | Earlybird | — |
| 2020 | Gretel | Funding round | $3.5M | Moonshots Capital | — |
| 2020 | Gretel | Funding round | $12M | Greylock | — |
| 2021 | Gretel | Funding round | $50M | Anthos Capital | Existing investors Greylock, Moonshots Capital |
| 2021 | AI.Reverie | Acquisition | Undisclosed | Meta | Integrated into Reality Labs |
| 2022 | MOSTLY AI | Series B | $25M | Molten Ventures | 42CAP, Earlybird |
| 2022 | Datagen | Series B | $50M | Scale Venture Partners | — |
| 2022 | Synthesis AI | Series A | $17M | — | — |
| 2022 | Mindtech Global | Minority investment | £2M | Appen | Also formed a commercial strategic partnership |
| 2022 | Synthesized | Corporate VC investment | Undisclosed | Deutsche Bank (CVC) | — |
| 2024 | Hazy (technology) | Technology acquisition | Undisclosed | SAS | Acquired synthetic data technology |
| 2025 | Synthesized | Series A | $20M | Redalpine | — |
| 2025 | Gretel | Acquisition | Nine-figure sum (exceeding $320M valuation) | NVIDIA | Terms not disclosed |
| 2025 | Scale AI | Investment | Over $14B | Meta | — |
Technology Development Directions for Synthetic Data Companies
—Generation Technology
Synthetic data companies need to remain open to new technologies. Those that better integrate their own technologies with emerging innovations will be better positioned in the future synthetic data market.
Emerging technologies such as quantum computing and digital twins will fundamentally change synthetic data generation, achieving higher realism, scalability, and efficiency. Together, these technologies will push synthetic data from “static replication” to “dynamic evolution,” greatly expanding its applicability in complex decision‑making scenarios. For example, quantum computing can significantly accelerate large‑scale data generation through optimization algorithms, enhancing both authenticity and scalability, especially in finance and logistics.
—Generation Model
Currently, industrial‑grade AI training relies heavily on expensive, manually annotated real data. To address this, many synthetic data companies are exploring a new synthetic data generation model: the hybrid model of “1% human data + 99% efficient synthesis.”
Specifically, this approach uses a small amount of high‑quality, rigorously annotated human data as a seed to drive the AI to generate large‑scale synthetic data rich in challenging scenarios. Its success depends on a “human‑in‑the‑loop” mechanism – domain experts intervene in data selection, rule definition, and quality assessment to ensure the high value and trustworthiness of the synthetic data. Ultimately, this paradigm will build a dynamic data pool far exceeding the scale and coverage of pure manual annotation, providing a critical solution for high‑reliability AI training in core industries.
Insight from AI Robots Eidos about Synthetic Data Companies
—The future leading synthetic data companies will no longer just provide datasets or generative tools, but will deeply integrate with vertical industries (such as autonomous driving, finance, and embodied intelligence) to deliver end-to-end “simulation-synthesis-validation” closed-loop solutions. By using synthetic data to quickly build high-fidelity digital testing environments, companies can complete product prototype validation and compliance reviews without real data, significantly shortening the R&D cycle.
—Synthetic data can generate “twin data” that does not contain any real individuals, allowing different organizations to train and test models while retaining the statistical characteristics of the data. Future synthetic data companies may lead the establishment of industry-level synthetic data exchange platforms, enabling cross-institutional collaboration with “data available but not visible,” which will greatly unlock the value of data that has been locked away.
—The core competitive advantage of synthetic data companies will be partially reflected in physical fidelity. Those who can simulate physical laws such as gravity, friction, and deformation more realistically will be able to provide genuinely usable pre-training data for humanoid robots, autonomous driving, and industrial robots. Breakthroughs in physical simulation engines will become the strategic high ground in the next phase.
Image Credits: Nvidia & Syntho & Ai & Laborcapital & Datacenterknowledge
