# AI Incident Triage

*/Opportunities/AI_Incident_Triage*

## Opportunity Overview

**Wedge**: Target on-call engineers at mid-market B2B SaaS companies by handling non-critical warning-tier alerts. This low-risk beachhead allows the product to prove accuracy and reduce out-of-hours interruptions without risking production stability. Once trust is established, expand coverage to Sev-1 critical incidents and eventually generate automated remediation pull requests.
**Timing**: Recent advancements in long-context window models allow LLMs to ingest thousands of raw server logs, trace spans, and recent git commits simultaneously. Previous generations lacked the context capacity and retrieval-augmented reasoning required to map generic infrastructure alerts to specific code deployments.
**Why This I C P**: Site Reliability Engineers and DevOps teams operate under strict service level agreements where every minute of downtime equates to measurable revenue loss. They are highly technical buyers who readily adopt API-driven workflows to reduce severe alert fatigue.
**Size Of Prize**: There are approximately 150,000 mid-market and enterprise software organizations globally. At an estimated annual spend of $25,000 per organization for automated incident triage software, the total addressable prize is roughly $3.75B.
**Gap Narrative**: Current application performance monitoring tools generate high volumes of alert noise, forcing Site Reliability Engineers to manually correlate metrics, logs, and traces across multiple dashboards during an active outage. The missing component is an autonomous reasoning layer that parses the initial alert, executes diagnostic queries, and synthesizes a root-cause hypothesis before human intervention. This replaces the initial 30-minute manual investigation phase with instant context.
**Defensibility**: The product accumulates a proprietary knowledge graph mapping the customer's specific microservice architecture, historical incident resolutions, and unwritten diagnostic steps. As the agent observes more outages, its diagnostic accuracy becomes deeply entrenched in the idiosyncrasies of the customer's environment, creating a high switching cost.
**Why This Thesis**: An Agent-based approach fits the operational reality of incident triage, which inherently consists of sequential fetch-and-read actions. The agent executes read-only queries against observability platforms and code repositories, replicating the exact investigative workflow of a human responder.

## Opportunity Linked Thesis

**Thesis**: [Agent](/Theses/Agent)

## Opportunity Linked I C P

**Icp**: [Managed Service Provider](/CompanyTypes/Managed_Service_Provider)

## Opportunity Market Sizing

_Illustrative — target and order-of-magnitude estimate figures, not an achieved track record (this Thing is concept-stage)._

**S A M**: ~$500-800M (English-speaking mid-market MSPs managing 5k+ endpoints)
**S O M**: ~$10-25M
**T A M**: ~120k global MSPs × ~$25k/yr allocated to L1 triage automation ≈ ~$3B
**Growth Rate**: ~15-20%/yr, driven by rising offshore labor costs and exponential growth in endpoint security alerts
**Paid Comparable Spend**: ~$45k-65k/yr per junior L1 dispatch technician or outsourced NOC seat

## Opportunity Incumbents

- [PagerDuty Incident Response](/Products/PagerDuty_Incident_Response) — Tool
- [ServiceNow ITOM](/Products/ServiceNow_ITOM) — Tool
- [BigPanda AIOps](/Products/BigPanda_AIOps) — Tool
- [Jira Service Management](/Products/Jira_Service_Management) — Tool
- [Manual Slack Routing](/Products/Manual_Slack_Routing) — DIY
- [Shared Excel Trackers](/Products/Shared_Excel_Trackers) — Spreadsheet
- [Custom Python Scripts](/Products/Custom_Python_Scripts) — DIY

## Opportunity Win Conditions

**Kill Thresholds**:
- Technician override rate > 20% after 30 days of tuning
- Integration time > 40 hours for standard PSA tools
- L1 ticket automation rate < 25% at day 60
- Willingness to pay < $15k ACV after the initial 30-day pilot
**Leading Metrics**:
- Time-to-first-automated-dispatch
- Percentage of L1 tickets routed without human intervention
- False-positive escalation rate
- Technician override rate per day
**What Proves Right**: Mid-market MSPs deploy the triage engine and route 40% of incoming L1 endpoint alerts to automated resolution or direct L2 escalation within 14 days of onboarding. Cohorts retain at over 90% beyond the first quarter, paying $25k ACV to directly offset outsourced NOC seats. Dispatch technicians actively reduce their manual ticket touches by half.
**What Proves Wrong**: The model fails to accurately categorize edge-case security alerts, driving a false-positive escalation rate above 15% that destroys technician trust. MSPs abandon the pilot and revert to manual routing because integration with legacy PSA tools requires more than 20 hours of custom configuration. Customers refuse to pay more than $5k annually because they classify the tool as a basic workflow script rather than a headcount replacement.

## Opportunity Build Profile

**Hardest Part**: Correlating massive volumes of noisy, unstructured log data with APM traces to pinpoint a root cause without hallucinating false anomalies. High-stress, time-critical environments mean the acceptable error rate for misdirection is near zero.
**Min Viable Scope**: Limit the v1 to reading Datadog alerts and AWS Kubernetes logs to generate a root cause hypothesis and paste the relevant runbook link into Slack. Deliberately exclude automated remediation, script execution, or multi-cloud support.
**Cold Start Problem**: The model requires rich, context-specific historical incident data like post-mortems and Slack threads which companies guard fiercely. Overcome this by offering a read-only ingest tool for a single design partner to process their last 50 resolved P1 incidents.
**Time To First Value**: 1-2 weeks to ingest and index historical runbooks and previous incident post-mortems before the first live triage.
**Data Moat Available**: true
**Technical Difficulty**: High

## Neighborhood

### Where the gap lives

- [Occupational Health and Safety Technicians](/Occupations/Occupational_Health_and_Safety_Technicians) — latent gap · Occupations
- [Cloud Infrastructure Managers](/Customers/Cloud_Infrastructure_Managers) — latent gap · Customers
- [Volume Of Data Compromised](/Metrics/Volume_Of_Data_Compromised) — latent gap · Metrics
- [Reported Compliance Violations](/Metrics/Reported_Compliance_Violations) — latent gap · Metrics
- [Site Reliability Engineers](/JobTypes/Site_Reliability_Engineers) — latent gap · JobTypes
- [Resolve technical user escalations](/Tasks/Resolve_technical_user_escalations) — latent gap · Tasks
- [Defect Leakage Rate](/Metrics/Defect_Leakage_Rate) — latent gap · Metrics
- [Reliability engineers](/Customers/Reliability_engineers) — latent gap · Customers
- [Response Protocol Compliance Rate](/Metrics/Response_Protocol_Compliance_Rate) — latent gap · Metrics
- [Testing Irregularity Rate](/Metrics/Testing_Irregularity_Rate) — latent gap · Metrics
- [Interface Clash Rate](/Metrics/Interface_Clash_Rate) — latent gap · Metrics
- [Root Cause Identification Rate](/Metrics/Root_Cause_Identification_Rate) — latent gap · Metrics
- [Software Development](/Industries/Software_Development) — latent gap · Industries

### Incumbent in

- [Bespoke Python Scripts](/Products/Bespoke_Python_Scripts) — incumbent in · Products
- [BigPanda](/Products/BigPanda) — incumbent in · Products
- [Jira Service Management](/Software/Jira_Service_Management) — incumbent in · Software
- [ServiceNow ITOM](/Products/ServiceNow_ITOM) — incumbent in · Products
- [Shared Excel Trackers](/Products/Shared_Excel_Trackers) — incumbent in · Products
- [Manual Slack Routing](/Products/Manual_Slack_Routing) — incumbent in · Products
- [PagerDuty Incident Response](/Products/PagerDuty_Incident_Response) — incumbent in · Products

### Applies thesis

- [Managed Service Provider](/CompanyTypes/Managed_Service_Provider) — applies thesis · CompanyTypes

### Embodies

- [Agent](/Theses/Agent) — embodies · Theses

### Similar Opportunities

- [Incident Triage Agent](/Opportunities/Incident_Triage_Agent) — similar · Opportunities
- [Automated Incident Dispatch](/Opportunities/Automated_Incident_Dispatch) — similar · Opportunities
- [Incident Resolution Automation](/Opportunities/Incident_Resolution_Automation) — similar · Opportunities
- [Autonomous Incident Responder](/Opportunities/Autonomous_Incident_Responder) — similar · Opportunities
- [Outage Detection Automation](/Opportunities/Outage_Detection_Automation) — similar · Opportunities
- [SLA Degradation Triage](/Opportunities/SLA_Degradation_Triage) — similar · Opportunities
- [Root Cause Investigator](/Opportunities/Root_Cause_Investigator) — similar · Opportunities
- [Autonomous SRE Responder](/Opportunities/Autonomous_SRE_Responder) — similar · Opportunities
- [Automated Incident Resolution](/Opportunities/Automated_Incident_Resolution) — similar · Opportunities
- [Automated Fault Triage](/Opportunities/Automated_Fault_Triage) — similar · Opportunities
- [Signal Node](/Opportunities/Signal_Node) — similar · Opportunities
- [AI Alert Aggregation](/Opportunities/AI_Alert_Aggregation) — similar · Opportunities
- [Root Cause Analyst](/Skills/Complex_Problem_Solving/Opportunities/Root_Cause_Analyst) — similar · Opportunities
- [Root Cause Analyst](/Opportunities/Root_Cause_Analyst) — similar · Opportunities
- [Staged Runbook Retrieval](/Opportunities/Staged_Runbook_Retrieval) — similar · Opportunities
- [Downtime Recovery Agent](/Opportunities/Downtime_Recovery_Agent) — similar · Opportunities
- [Troubleshooting as a Service](/Opportunities/Troubleshooting_as_a_Service) — similar · Opportunities
- [Automated SLA Recovery](/Skills/Systems_Evaluation/Opportunities/Automated_SLA_Recovery) — similar · Opportunities
- [Incident Context Synthesizer](/Opportunities/Incident_Context_Synthesizer) — similar · Opportunities
- [Bug Triage Automation](/CompanyTypes/White-Label_Software_Provider/Opportunities/Bug_Triage_Automation) — similar · Opportunities
