Artificial intelligence is entering a stage where better models are no longer the only priority. The quality, complexity, and specialization of the data used to train and evaluate those models are becoming increasingly important. That shift is putting AI training data at the center of a rapidly growing technology business—and Snorkel AI is emerging as one of the companies at the heart of it.
In September, Snorkel AI raised $350 million in Series E funding at a valuation of $3.5 billion, highlighting the growing investor interest in the infrastructure behind advanced AI systems. The round comes as AI developers increasingly look beyond traditional data labeling toward specialized datasets, expert-generated information, evaluation data, and simulated environments for increasingly capable AI models.
For Technology Moment readers, the story is bigger than another billion-dollar AI funding announcement. It offers a closer look at how the economics of AI are changing. Snorkel AI has evolved from its earlier focus on data-labeling software toward a broader data-as-a-service model, supplying finished datasets and environments designed for modern AI development.
So, why is AI training data becoming such a valuable business? What is driving Snorkel AI’s rapid growth, and what does its $3.5 billion valuation tell us about the future of AI infrastructure? In this Technology Moment analysis, we explore Snorkel AI’s latest funding, its business model, the rise of “Data 2.0,” and why high-quality training data could become one of the most important assets in the next phase of artificial intelligence.
Snorkel AI Raises $350M as Its Valuation Hits $3.5B
Snorkel AI has raised $350 million in Series E funding at a $3.5 billion valuation, marking a major milestone for the company and highlighting the growing importance of AI training data. The funding round was announced on September 22, 2026, and was co-led by Insight Partners and S32, with participation from existing and new investors. The new valuation is nearly three times the $1.3 billion valuation Snorkel AI reached when it raised $100 million in its Series D round in 2025.
The latest Snorkel AI funding comes at a time when AI companies are facing a different kind of data challenge. Earlier AI development relied heavily on large quantities of relatively simple labeled datasets. Today, frontier AI labs increasingly require specialized training data, evaluation datasets, expert-generated information, and simulated environments for increasingly complex AI systems.
Snorkel AI has also changed its business model. Instead of focusing primarily on software for automating data labeling, the company has moved toward a data-as-a-service approach, providing customers with completed datasets and environments designed for AI training and evaluation. TechCrunch reported that the company’s annualized revenue run rate had reached $375 million, while Reuters reported it had crossed $350 million, showing the rapid expansion of the business.
For the wider AI industry, the Snorkel AI $350M funding round is therefore more than another startup investment. It reflects increasing investor and industry attention toward the infrastructure, datasets, and specialized expertise required to build the next generation of AI systems.
Who Invested in Snorkel AI?
The latest Snorkel AI Series E funding was co-led by Insight Partners and S32, with significant participation from existing investor Addition. The round also attracted a group of new investors, including March Capital, Blumberg Capital, Allegis Capital, Frontline, Standard, and Third Point Ventures. Existing investors including Greylock, Lightspeed, GV, Factory, Prosperity7, Walden Catalyst, and Wells Fargo also participated, according to Snorkel AI’s announcement.
The investor mix is significant because it shows continued financial interest in companies operating around AI infrastructure, AI data, machine learning datasets, and frontier AI. Rather than simply investing in another application layer, the funding is supporting a company focused on the data and environments used to develop and evaluate advanced AI systems.
Insight Partners has also described the changing requirements of AI training data. According to the investment firm, AI systems are moving toward longer-running and higher-stakes tasks, increasing demand for specialized data, environments, and evaluations that can be difficult to create. The funding will help Snorkel AI expand what it calls its agentic data factory, which provides data and environments for advanced AI development. The company says its work includes datasets, benchmarks, evaluations, and custom solutions for frontier and agentic AI systems.
For readers following Snorkel AI investors, Snorkel AI funding news, Snorkel AI valuation, and AI startups, the Series E round provides an important signal about where capital is moving within the broader AI ecosystem. The investment also places greater attention on the business potential of high-quality training data and specialized AI data development.
What Is Snorkel AI?
Snorkel AI is an AI data company that helps organizations develop the datasets, benchmarks, evaluations, and environments needed to build and assess advanced AI systems. The company was founded in 2019 out of the Stanford AI Lab, initially around research into data-centric approaches to machine learning. Its early work focused on improving the way organizations create and label training data. Snorkel’s technology used approaches such as weak supervision and data programming to reduce reliance on manually labeling every individual example. Over time, the company expanded beyond traditional data labeling toward specialized AI data development.
Today, Snorkel describes itself as a frontier AI data lab rather than simply an AI data-labeling software provider. Its offerings include datasets, evaluation frameworks, benchmarks, simulated environments, and custom data solutions. Its Data-as-a-Service (DaaS) model is designed to help AI teams obtain specialized data for domain-specific and high-consequence applications.
This evolution reflects a broader change in AI development. As models become more capable, simply collecting larger quantities of generic information may not solve every training or evaluation problem. Developers may instead need high-quality expert training data, carefully designed tasks, reliable evaluation methods, and environments where AI agents can be tested.
Snorkel AI calls this emerging approach “Data 2.0.” The company argues that advanced AI requires a research- and engineering-driven approach to data development rather than treating data creation simply as a large-scale labeling exercise. Understanding what Snorkel AI does therefore requires looking beyond traditional data annotation. Its current business sits at the intersection of data-centric AI, AI model training, AI evaluation, machine learning datasets, agentic AI, and AI infrastructure.
Why AI Training Data Is Becoming Big Business
The growing value of AI training data is closely connected to the increasing complexity of AI models. Earlier data-generation workflows often involved straightforward tasks such as labeling images, ranking responses, or categorizing information. As AI systems move toward reasoning, tool use, autonomous actions, and specialized applications, developers need more sophisticated forms of data.
That includes expert training data, AI evaluation data, synthetic training data, benchmarks, reinforcement-learning environments, and agentic tasks. These resources can require specialized knowledge and careful design rather than simple high-volume annotation. Snorkel AI and its investors describe this transition as a move from “Data 1.0” toward “Data 2.0,” where quality, complexity, and domain expertise become increasingly important.
This creates an opportunity for companies that can help AI developers produce and validate such data at scale. Snorkel AI’s shift toward data-as-a-service is one example. The company says its offering provides ready-to-use and custom data development for specialized AI applications, including datasets, benchmarks, evaluations, and environments.
The trend also explains why AI data startups are attracting increasing attention. As frontier AI labs work on more capable systems, the challenge is not necessarily finding more raw information. It can be finding the right data for a particular task, defining meaningful evaluation criteria, and creating environments where complex AI behavior can be measured. That makes AI data development an increasingly important part of the broader AI infrastructure market. Snorkel AI’s $350 million funding round does not by itself prove how large the entire AI training data market will become, but it demonstrates that investors are placing substantial capital behind companies addressing this emerging requirement.
From Data Labeling to “Data 2.0”
The AI data industry is moving beyond the traditional idea of simply collecting and labeling large datasets. For years, AI development depended heavily on data labeling, annotation, and large volumes of relatively simple training examples. This approach helped machine learning systems recognize patterns, classify information, and improve through supervised learning. Snorkel AI’s early work was closely connected to this data-centric AI approach, including techniques such as weak supervision and programmatic data development. But as AI models have become more capable, the requirements for useful training data have also changed.
Snorkel AI describes this transition as “Data 2.0.” In the company’s explanation, the earlier “Data 1.0” era focused largely on volume and straightforward labeling tasks. Data 2.0, by contrast, focuses on highly targeted, complex, and high-quality data that can help advanced models improve in areas where they still struggle.
This means AI data development can involve much more than labeling a dataset. Modern AI data development may require expert training data, AI evaluation data, benchmarks, synthetic training data, and reinforcement learning environments. For example, a coding dataset for an advanced AI model may need realistic software-development tasks, detailed evaluation criteria, and safeguards against models exploiting weaknesses in the task itself.
Snorkel’s approach combines human-in-the-loop AI with specialized AI models and agents. The company says these systems can assist experts with creating and reviewing complex data while human feedback continues to evaluate and improve the AI tools. The shift from traditional data labeling to Data 2.0 therefore reflects a broader change in the AI industry: as models become more sophisticated, the value of data increasingly depends on its quality, complexity, relevance, and ability to address specific model limitations.
How Snorkel AI Makes Money
Snorkel AI’s business model has expanded significantly from its earlier software-focused approach. The company now operates a data-as-a-service (DaaS) model designed to help organizations develop specialized AI training data, datasets, evaluations, and environments. Reuters reports that this business, launched in September 2025, has been a major driver of Snorkel AI’s rapid revenue growth.
Instead of simply selling access to a traditional data-labeling platform, Snorkel AI increasingly provides customers with finished AI data products. These can include specialized datasets and reinforcement-learning environments designed around specific AI development requirements. The company also works with expert networks in areas such as coding, law, medicine, and other specialized domains.
This approach is important because advanced AI training data can be difficult to produce using conventional crowdsourcing. A simple classification task may be completed by a large pool of general workers, but a complex legal, financial, scientific, or software-engineering task may require people with specialized knowledge. Snorkel AI’s model combines that human expertise with AI tooling to develop and review more complicated data.
The company also describes an agentic data platform in which AI systems assist human experts with creating datasets and environments, performing quality checks, and improving the overall data-development process. Snorkel says this creates a feedback loop in which AI helps experts and human feedback subsequently helps improve specialized AI systems.
For companies developing frontier AI, enterprise AI, or specialized AI agents, the appeal of this model is the ability to obtain data tailored to particular applications rather than relying only on generic datasets. That puts Snorkel AI at the intersection of AI data development, data annotation, AI model evaluation, machine learning datasets, and AI infrastructure. The business therefore represents a broader shift in the AI economy: companies are increasingly paying not simply for access to data, but for specialized data development and the technology required to create, evaluate, and maintain it.
Snorkel AI’s Rapid Revenue Growth
Snorkel AI’s latest funding round has attracted attention not only because of its $3.5 billion valuation, but also because of the company’s reported revenue growth. In its September 2026 funding announcement, Snorkel AI said its annualized revenue run rate had reached $375 million, representing more than 18 times growth since launching its data-as-a-service offering nearly a year earlier. Reuters separately reported that annualized revenue had risen to more than $350 million, compared with $20 million the previous year.
Those figures illustrate how quickly the company’s business has expanded alongside demand for advanced AI training data. Snorkel AI says it now works with frontier AI labs, hyperscalers, AI-native companies, enterprises, and U.S. government agencies. The important point is that this growth is connected to a changing AI data market. As AI models become more capable, companies need increasingly specialized datasets, evaluation systems, and environments. Insight Partners, which co-led Snorkel AI’s Series E, says demand is shifting from general-purpose data that humans can easily write down toward specialized data, environments, and evaluations that have to be engineered.
Snorkel AI’s revenue model is therefore closely connected to its data-as-a-service strategy. Rather than treating expert knowledge simply as labor hours, the company says it sells data products generated through a combination of human expertise and AI-powered development. Reuters reported that this approach is intended to support expert compensation while maintaining the economics of a technology business.
The company’s growth also helps explain the attention surrounding its latest $350M funding round. The financing gives Snorkel additional resources to expand its agentic data factory, grow enterprise and government operations, and explore additional industries and data types. Reuters reported that the company aims to reach profitability by the end of 2026. For the broader AI industry, Snorkel’s revenue growth provides a concrete example of how AI data businesses are evolving from traditional labeling services toward specialized data platforms and AI infrastructure.
Why Frontier AI Labs Need Better Training Data
The demand for better frontier AI training data comes from a simple problem: making an AI model more capable becomes increasingly difficult once it has already learned the basics. When models are relatively limited, large amounts of ordinary training data can produce substantial improvements. But as models approach higher levels of capability, developers increasingly need data that targets specific weaknesses and difficult real-world tasks. Snorkel AI describes this as the difference between the volume-focused “Data 1.0” era and the more complex “Data 2.0” era.
For frontier AI labs, that can mean creating datasets that test advanced reasoning, coding, research, professional expertise, or multi-step decision-making. It can also mean building reinforcement learning environments in which AI agents can interact with tools and attempt complex tasks. Insight Partners notes that these environments need to be engineered around realistic operating contexts rather than simply crowdsourced as basic labeling exercises.
Coding provides a useful example. Snorkel AI says advanced coding datasets may need to represent software-development problems that can take experienced engineers days or weeks to solve. They also need meaningful reward signals, detailed quality controls, and protection against AI systems exploiting weaknesses in the task or evaluation process. This is where expert training data, AI evaluation data, AI benchmarking, data curation, and agentic AI become increasingly connected. Training data teaches a model, while evaluation data and benchmarks help determine whether that model actually improved. Environments allow AI agents to practice complex tasks and generate measurable outcomes.
Another important factor is the role of humans. Snorkel AI argues that advanced data development cannot rely entirely on AI-generated information because human expertise and judgment remain important for identifying what models do not yet know and for maintaining oversight. Its platform therefore combines human experts with specialized AI agents. For frontier AI labs, the challenge is consequently shifting from simply acquiring more data to developing the right data at the right level of difficulty. That change is helping turn AI training data from a supporting component of model development into a significant part of the broader AI infrastructure and AI data market.
Snorkel AI vs. the Traditional AI Data Labeling Market
The AI data industry is changing from a model centered primarily on large-scale data labeling toward a broader approach focused on specialized datasets, evaluation systems, synthetic data, and environments for advanced AI models. Snorkel AI is a useful example of this transition. The company originally developed software for automating aspects of data labeling, including programmatic labeling and weak supervision. Today, however, Snorkel positions itself as a frontier AI data lab, providing datasets, benchmarks, evaluations, and environments for advanced AI systems.
Traditional AI data labeling generally involves assigning labels or annotations to existing information so machine learning models can learn from it. This remains important, particularly for high-volume datasets. However, frontier AI development increasingly requires more complicated forms of AI data development, including expert-generated tasks, detailed rubrics, evaluation datasets, reinforcement-learning environments, and agentic workflows. Snorkel calls this broader transition Data 2.0.
| Traditional AI Data Labeling | Snorkel AI / Data 2.0 Approach |
|---|---|
| Focus on labeling and annotation | Focus on complete data development |
| Primarily volume-driven | Quality, complexity and specialization |
| Human labeling workflows | Human experts + AI agents |
| Static datasets | Datasets + environments + evaluations |
| Basic supervised-learning data | Training, post-training and evaluation data |
| Manual quality checks | Automated and expert quality controls |
| General-purpose tasks | Specialized and domain-specific tasks |
| Data labeling software | Data-as-a-service and data-lab model |
This does not mean traditional labeling is disappearing. Instead, the market is expanding into additional layers of AI infrastructure. Snorkel’s current model combines human expertise with specialized AI models and agents, while its competitors and other AI training data companies are also developing approaches for supplying increasingly sophisticated data. The distinction is therefore less about “labeling versus no labeling” and more about simple data production versus engineered data systems designed around the needs of increasingly capable AI models.
The Growing AI Training Data Market
The growing AI training data market is being shaped by the increasing complexity of modern AI systems. As frontier models become more capable, developers need data that can address specific weaknesses rather than relying exclusively on larger quantities of generic information. This is creating demand for high-quality training data, expert training data, AI evaluation data, machine learning datasets, synthetic training data, benchmarks, and reinforcement-learning environments. Reuters reports that demand from frontier AI labs for increasingly complex training data and simulated environments has been a major factor behind Snorkel AI’s recent growth.
One important change is the movement toward specialized data. A general dataset may help a model learn broad patterns, but advanced systems working in coding, law, medicine, finance, science, or other professional fields can require domain-specific knowledge and carefully designed evaluation criteria. Snorkel says its expert network covers specialized areas and that its platform supports domains including software engineering, STEM, healthcare, finance, legal and regulatory work, manufacturing, and government.
Another major development is the growing role of synthetic data and AI-assisted data development. Snorkel’s model combines human experts with AI models and agents, rather than relying entirely on either manual work or automatically generated information. Reuters reports that the company sees human input remaining important while synthetic and automated approaches are needed to scale data production.
This creates a broader opportunity for AI data startups and AI training data companies. The market now extends beyond traditional annotation providers into data platforms, evaluation companies, synthetic-data providers, expert networks, and AI infrastructure businesses. The exact future size of the market remains uncertain, but Snorkel’s funding provides a clear example of investor interest in this category. Its $350 million Series E funding at a $3.5 billion valuation demonstrates that specialized AI data has become a significant part of the wider AI investment landscape.
What the $350M Funding Means for Snorkel AI
Snorkel AI’s $350 million funding round gives the company additional capital to expand its data-development business at a time when demand for specialized AI training data is increasing. The Series E round was co-led by Insight Partners and S32, with participation from Addition and a broader group of existing and new investors. The funding valued Snorkel AI at $3.5 billion, almost three times its $1.3 billion valuation from its previous $100 million financing in 2025.
According to Snorkel AI, the new capital will support the expansion of its agentic data factory, which is designed to produce the data and environments required by advanced AI systems. The company also plans to invest in researchers and engineers, expand enterprise and government operations, support third-party AI model evaluations, and enter additional industry verticals and data modalities.
The funding also comes as Snorkel’s business model is changing. Its data-as-a-service offering, launched in September 2025, has become an important growth engine. Reuters reports that the company’s annualized revenue run rate has surpassed $350 million, compared with approximately $20 million a year earlier. Snorkel itself reported an annualized revenue run rate of $375 million after more than 18x growth since launching its new offering.
For Snorkel AI, the investment therefore supports more than simple expansion. It provides resources to develop the company’s AI data platform, agentic AI capabilities, evaluation infrastructure, and specialized training-data products. The wider significance is that investors are backing a business positioned around one of the increasingly important bottlenecks in advanced AI: producing data that is difficult to create, verify, and scale. Snorkel’s next phase will show how effectively its Data 2.0 approach can translate growing demand into sustainable business growth.
Frequently Asked Questions About Snorkel AI
What is Snorkel AI?
Snorkel AI is a frontier AI data lab founded out of the Stanford AI Lab in 2019. It develops datasets, benchmarks, evaluations, environments, and custom agents for advanced AI systems. The company has evolved from its earlier focus on programmatic data labeling toward a broader data-as-a-service and AI data development model.
Why did Snorkel AI raise $350 million?
Snorkel AI raised $350 million in Series E funding to expand its agentic data factory, hire researchers and engineers, grow enterprise and government operations, support AI model evaluations, and expand into new industries and data modalities.
What is Snorkel AI’s valuation?
Following its latest Series E funding, Snorkel AI is valued at $3.5 billion. That represents a substantial increase from the $1.3 billion valuation it reached after its previous $100 million funding round in 2025. The latest valuation reflects the company’s growth and investor interest in complex AI training data.
How does Snorkel AI make money?
Snorkel AI increasingly operates through a data-as-a-service model. Rather than focusing only on labeling software, it provides customers with specialized datasets, benchmarks, evaluations, and environments. Its current model combines expert human input with AI-assisted and synthetic approaches to develop data for advanced AI applications.













