# Pre-Run Anomaly Detection

*/Opportunities/Pre-Run_Anomaly_Detection*

## Opportunity Overview

**Wedge**: The beachhead focuses strictly on dbt users running daily batch transformations in Snowflake or BigQuery. This niche experiences acute, measurable compute waste from failed downstream models and integrates rapidly via standard open-source APIs. Expansion proceeds from batch dbt transformations into pre-flight checks for distributed machine learning training jobs and real-time streaming infrastructure.
**Timing**: Large language models now reliably parse complex, unstructured log states and data distributions without rigid manual rule-writing, making zero-shot, context-aware anomaly detection viable before a pipeline run begins.
**Why This I C P**: Data engineering and MLOps teams endure the immediate pain of failed pipelines through off-hours pager alerts and wasted cloud compute spend, creating immediate budget authority for preventative tooling.
**Size Of Prize**: Roughly 40,000 mid-market and enterprise data teams spend an average of $20,000 annually on pipeline observability and compute waste recovery, yielding an $800M addressable prize.
**Gap Narrative**: Data engineering teams execute data pipelines and machine learning models, only discovering structural anomalies after expensive compute or downstream data corruption occurs. They lack a pre-flight system that evaluates incoming data schemas, checks statistical boundaries, and identifies structural deviations before execution to prevent failed jobs and costly rollbacks.
**Defensibility**: Defensibility compounds through accumulated baseline models of an organization's specific data operations. As the system observes more pre-run states and subsequent execution outcomes, its contextual accuracy becomes hyper-tailored to the proprietary data velocity and shape of the enterprise, creating high mathematical switching costs.
**Why This Thesis**: An agentic approach fits this problem shape because anomaly detection requires contextual reasoning over thousands of distinct schemas and historical data distributions, which agents execute autonomously rather than forcing engineers to manually maintain static rule sets.

## Opportunity Linked Thesis

**Thesis**: [Software](/Theses/Software)

## Opportunity Linked I C P

**Icp**: [Data Analytics Firm](/CompanyTypes/Data_Analytics_Firm)

## Opportunity Market Sizing

_Illustrative — target and order-of-magnitude estimate figures, not an achieved track record (this Thing is concept-stage)._

**S A M**: ~$400M - $600M (cloud-native data analytics firms utilizing consumption-based data warehouses)
**S O M**: ~$15M - $25M
**T A M**: ~75k global data analytics firms and enterprise data teams × ~$20k/yr ≈ $1.5B
**Growth Rate**: ~20-25%/yr, driven by escalating cloud compute usage-based pricing and increasing data pipeline complexity
**Paid Comparable Spend**: ~$40k - $70k/yr per firm in wasted cloud compute for failed queries and manual data engineering review hours

## Opportunity Incumbents

- [Great Expectations](/Products/Great_Expectations) — Open-Source
- [Monte Carlo Data](/Products/Monte_Carlo_Data) — Tool
- [Custom Python Scripts](/Products/Custom_Python_Scripts) — DIY
- [SonarQube Code Quality](/Products/SonarQube_Code_Quality) — Tool
- [Datadog CI Visibility](/Products/Datadog_CI_Visibility) — Tool
- [Manual Pre-Flight Checks](/Products/Manual_Pre-Flight_Checks) — DIY

## Opportunity Win Conditions

**Kill Thresholds**:
- False positive rate > 15% after 14 days of pipeline integration
- Time-to-first-value > 48 hours for standard CI/CD setups
- Average verifiable compute savings < $1,500 per month during trial periods
- Conversion rate from trial to paid < 25% at the 90-day mark
**Leading Metrics**:
- Time-to-first-value (hours from signup to first intercepted query)
- False positive rate (percentage of flagged queries manually overridden by engineers)
- Percentage of total warehouse queries routed through the pre-run checker
- Compute credits saved per week per active deployment
- Engineer weekly active usage (reviewing flagged queries)
**What Proves Right**: Data engineering teams integrate the pre-run checker into their deployment pipelines and block at least 30 percent of malformed queries prior to execution. Early cohorts exhibit a 60 percent reduction in weekly wasted compute spend on consumption-based warehouses like Snowflake or BigQuery. Customers sign $15,000 annual contracts after a 30-day proof of value demonstrates a direct reduction in their monthly cloud infrastructure bill.
**What Proves Wrong**: Data engineers bypass the pre-run anomaly checks because the false positive rate exceeds 10 percent and blocks valid pipeline executions. The tool fails to flag structural errors for complex multi-join queries, limiting its value to trivial syntax errors already caught by basic linters. Prospects refuse to pay standalone subscription fees because they view pre-run validation as a feature that should exist natively within their data warehouse compilers.

## Opportunity Build Profile

**Hardest Part**: Distinguishing genuine data entry errors from expected operational variance such as seasonal bonuses, prorated schedules, or shift differentials without triggering overwhelming false positive alerts.
**Min Viable Scope**: Focus exclusively on flagging gross variance in total payout amounts and missing line items for a single system of record integration. Deliberately exclude automated correction, tax anomaly detection, and complex benefits reconciliation, outputting only a prioritized review list for a human operator.
**Cold Start Problem**: The system lacks a baseline for normal variance without access to historical run data. Break this by requiring read-only API access to the customer system of record to ingest and analyze the last twelve months of historical runs during initial onboarding.
**Time To First Value**: 1-2 weeks of onboarding to ingest historical data and surface anomalies during the very first live prep-run cycle.
**Data Moat Available**: true
**Technical Difficulty**: High

## Neighborhood

### Surfaced from

- [Cloud-Native SMB Payroll SaaS](/CompanyTypes/Cloud-Native_SMB_Payroll_SaaS) — surfaces · CompanyTypes

### Incumbent in

- [Bespoke Python Scripts](/Products/Bespoke_Python_Scripts) — incumbent in · Products
- [SonarQube Code Quality](/Products/SonarQube_Code_Quality) — incumbent in · Products
- [Manual Pre-Flight Checks](/Products/Manual_Pre-Flight_Checks) — incumbent in · Products
- [Monte Carlo Data](/Products/Monte_Carlo_Data) — incumbent in · Products
- [Datadog CI Visibility](/Products/Datadog_CI_Visibility) — incumbent in · Products
- [Great Expectations](/Products/Great_Expectations) — incumbent in · Products

### Applies thesis

- [Data Analytics Firm](/CompanyTypes/Data_Analytics_Firm) — applies thesis · CompanyTypes

### Embodies

- [Software](/Theses/Software) — embodies · Theses

### Similar Opportunities

- [Data Pipeline Repair](/Opportunities/Data_Pipeline_Repair) — similar · Opportunities
- [Training Data Sanitizer](/Metrics/Data_Accuracy_Rate/Occupations/Machine_Learning_Engineers/Opportunities/Training_Data_Sanitizer) — similar · Opportunities
- [Metric Triage Agent](/Opportunities/Metric_Triage_Agent) — similar · Opportunities
- [Semantic Auditing for Data](/Opportunities/Semantic_Auditing_for_Data) — similar · Opportunities
- [Automated Schema Reconciliation for Enterprises](/Opportunities/Automated_Schema_Reconciliation_for_Enterprises) — similar · Opportunities
- [Algorithm Medic](/Opportunities/Algorithm_Medic) — similar · Opportunities
- [Predictive Hibernation For DevOps](/Opportunities/Predictive_Hibernation_For_DevOps) — similar · Opportunities
- [Predictive Telemetry Engine](/Opportunities/Predictive_Telemetry_Engine) — similar · Opportunities
- [Automated Review for DevOps Teams](/Opportunities/Automated_Review_for_DevOps_Teams) — similar · Opportunities
- [Context Enrichment Pipeline](/Opportunities/Context_Enrichment_Pipeline) — similar · Opportunities
- [AI Systems Engineering](/Skills/Systems_Evaluation/Opportunities/AI_Systems_Engineering) — similar · Opportunities
- [Incident Prevention API](/Opportunities/Incident_Prevention_API) — similar · Opportunities
- [Schema Inference for Data Engineers](/Opportunities/Schema_Inference_for_Data_Engineers) — similar · Opportunities
- [Feature Generation Agent](/Opportunities/Feature_Generation_Agent) — similar · Opportunities
- [Automated Fault Triage](/Opportunities/Automated_Fault_Triage) — similar · Opportunities
- [AI Compliance Auditing](/Metrics/Information_Accuracy/Opportunities/AI_Compliance_Auditing) — similar · Opportunities
- [Continuous Anomaly Detection](/CompanyTypes/Accounting_Firm/Opportunities/Continuous_Anomaly_Detection) — similar · Opportunities
- [AI Pipeline Configuration for Enterprise DevOps](/Opportunities/AI_Pipeline_Configuration_for_Enterprise_DevOps) — similar · Opportunities
- [Cloud Cost Remediation](/Opportunities/Cloud_Cost_Remediation) — similar · Opportunities
- [Root Cause Investigator](/Opportunities/Root_Cause_Investigator) — similar · Opportunities
