# Staged Runbook Retrieval

*/Opportunities/Staged_Runbook_Retrieval*

## Opportunity Overview

**Wedge**: The initial beachhead targets Kubernetes infrastructure alerts, specifically out-of-memory and crash loop events for mid-market SaaS companies. This niche provides a highly standardized alert structure and predictable diagnostic steps, allowing for immediate proof of value. Expansion follows by ingesting custom internal microservice documentation, moving from standardized infrastructure alerts to proprietary application logic.
**Timing**: Large context window LLMs and fast vector embeddings now allow systems to continuously map live telemetry and complex alert payloads to unstructured documentation in milliseconds. Previously, deterministic retrieval was too fragile and AI retrieval was too slow to trust during high-stakes outages.
**Why This I C P**: DevOps teams and SREs are strictly measured on MTTR metrics and already operate in heavily instrumented environments like Datadog and PagerDuty. The data required to trigger context-aware retrieval is already structured and accessible via webhooks, enabling immediate deployment.
**Size Of Prize**: There are roughly 40,000 mid-to-large enterprise engineering organizations globally that invest heavily in incident management. At an estimated average annual software spend of $15,000 per organization dedicated to mean-time-to-recovery (MTTR) reduction tools, this represents a $600M addressable prize.
**Gap Narrative**: SREs and IT operations teams lose critical minutes during high-severity incidents manually searching for and contextualizing the correct runbook for a specific alert state. Existing knowledge bases require active, keyword-based search that pulls operators out of their diagnostic flow. A system that automatically retrieves and surfaces the exact procedural steps based on real-time telemetry and incident stage eliminates this cognitive friction.
**Defensibility**: Defensibility compounds through workflow lock-in and a proprietary feedback loop. As the system observes which runbook steps operators actually execute, skip, or modify during an active incident, it refines its retrieval ranking for that specific company. Over time, it captures the undocumented tribal knowledge of the engineering team, creating severe switching costs.
**Why This Thesis**: A Software approach that embeds directly into existing incident communication channels like Slack fits the SRE workflow perfectly. Operators require the exact runbook steps delivered into the chat interface where the triage is actively happening, rather than logging into a separate dashboard.

## Opportunity Linked Thesis

**Thesis**: [Software](/Theses/Software)

## Opportunity Linked I C P

**Icp**: [Managed Service Provider](/CompanyTypes/Managed_Service_Provider)

## Opportunity Market Sizing

_Illustrative — target and order-of-magnitude estimate figures, not an achieved track record (this Thing is concept-stage)._

**S A M**: ~$400-600M global managed service providers
**S O M**: ~$15-30M
**T A M**: ~100k global IT service providers and internal help desks × ~$15k/yr ≈ $1.5B
**Growth Rate**: ~12-18%/yr, driven by increasing IT stack complexity and high Tier 1 technician turnover rates
**Paid Comparable Spend**: ~$15k-30k/yr per MSP on static knowledge base subscriptions and wasted Tier 1 to Tier 2 escalation labor

## Opportunity Incumbents

- [PagerDuty Runbook Automation](/Products/PagerDuty_Runbook_Automation) — Tool
- [Atlassian Confluence Documentation](/Products/Atlassian_Confluence_Documentation) — DIY
- [AWS Systems Manager](/Products/AWS_Systems_Manager) — Tool
- [FireHydrant Incident Management](/Products/FireHydrant_Incident_Management) — Tool
- [Rundeck Open Source](/Products/Rundeck_Open_Source) — Open-Source
- [Internal Notion Wikis](/Products/Internal_Notion_Wikis) — DIY
- [ServiceNow IT Operations](/Products/ServiceNow_IT_Operations) — Tool

## Opportunity Win Conditions

**Kill Thresholds**:
- Tier 1 to Tier 2 escalation rate drops by less than 15 percent after 45 days
- Runbook adoption rate by Tier 1 agents remains under 40 percent in month one
- Initial runbook ingestion and staging takes longer than 14 days per customer
**Leading Metrics**:
- Tier 1 resolution rate for runbook-assisted tickets
- Runbook retrieval relevance score via agent acceptance
- Mean time to resolution for supported incident categories
- Volume of manual documentation searches post-retrieval
**What Proves Right**: Staged Runbook Retrieval identifies incident parameters and surfaces the exact execution steps for Tier 1 agents. Tier 1 agents resolve these tickets directly, reducing the Tier 2 escalation volume. Customers commit to the $15k annual price point because the system clearly eliminates the equivalent Tier 2 labor waste.
**What Proves Wrong**: Tier 1 agents ignore the surfaced steps and escalate the ticket to Tier 2 anyway. The system maps incident payloads to generic wiki links rather than actionable, staged procedures. Setup stalls because Tier 3 engineers refuse to convert their existing documentation into the required structured format.

## Opportunity Build Profile

**Hardest Part**: Mapping noisy machine alerts to informally structured runbook steps with near-perfect precision to prevent destructive actions during Sev1 incidents.
**Min Viable Scope**: Connect PagerDuty webhooks to a single GitHub repository of Markdown runbooks and post the top matched step into a Slack incident channel. Deliberately leave out automated remediation execution, conversational debugging, and multi-source enterprise search.
**Cold Start Problem**: The system lacks baseline mapping between unique infrastructure alerts and human-written documentation. Break this by ingesting a static export of recent resolved incident tickets and retroactively matching their payload data to the runbook repository.
**Time To First Value**: 1 week to index historical documentation and trigger the first contextual retrieval during a live incident.
**Data Moat Available**: true
**Technical Difficulty**: High

## Neighborhood

### Incumbent in

- [ServiceNow IT Operations](/Products/ServiceNow_IT_Operations) — incumbent in · Products
- [PagerDuty Runbook Automation](/Products/PagerDuty_Runbook_Automation) — incumbent in · Products
- [Rundeck Open Source](/Products/Rundeck_Open_Source) — incumbent in · Products
- [AWS Systems Manager](/Products/AWS_Systems_Manager) — incumbent in · Products
- [Atlassian Confluence Documentation](/Products/Atlassian_Confluence_Documentation) — incumbent in · Products
- [FireHydrant Incident Management](/Products/FireHydrant_Incident_Management) — incumbent in · Products
- [Internal Notion Wikis](/Products/Internal_Notion_Wikis) — incumbent in · Products

### Applies thesis

- [Managed Service Provider](/CompanyTypes/Managed_Service_Provider) — applies thesis · CompanyTypes

### Embodies

- [Software](/Theses/Software) — embodies · Theses

### Similar Opportunities

- [Automated Fault Triage](/Opportunities/Automated_Fault_Triage) — similar · Opportunities
- [Automated Incident Dispatch](/Opportunities/Automated_Incident_Dispatch) — similar · Opportunities
- [AI Incident Triage](/Opportunities/AI_Incident_Triage) — similar · Opportunities
- [Incident Context Synthesizer](/Opportunities/Incident_Context_Synthesizer) — similar · Opportunities
- [Autonomous SRE Responder](/Opportunities/Autonomous_SRE_Responder) — similar · Opportunities
- [Automated Incident Reporter](/Opportunities/Automated_Incident_Reporter) — similar · Opportunities
- [Predictive Telemetry Engine](/Opportunities/Predictive_Telemetry_Engine) — similar · Opportunities
- [Autonomous Incident Responder](/Opportunities/Autonomous_Incident_Responder) — similar · Opportunities
- [Downtime Recovery Agent](/Opportunities/Downtime_Recovery_Agent) — similar · Opportunities
- [Incident Triage Agent](/Opportunities/Incident_Triage_Agent) — similar · Opportunities
- [Root Cause Investigator](/Opportunities/Root_Cause_Investigator) — similar · Opportunities
- [Outage Detection Automation](/Opportunities/Outage_Detection_Automation) — similar · Opportunities
- [Automated Log Reconciliation](/Opportunities/Automated_Log_Reconciliation) — similar · Opportunities
- [Incident Resolution Automation](/Opportunities/Incident_Resolution_Automation) — similar · Opportunities
- [Root Cause Analyst](/Opportunities/Root_Cause_Analyst) — similar · Opportunities
- [SLA Degradation Triage](/Opportunities/SLA_Degradation_Triage) — similar · Opportunities
- [Signal Node](/Opportunities/Signal_Node) — similar · Opportunities
- [AI Alert Aggregation](/Opportunities/AI_Alert_Aggregation) — similar · Opportunities
- [Incident Prevention API](/Opportunities/Incident_Prevention_API) — similar · Opportunities
- [Incident Narrative Desk](/Opportunities/Incident_Narrative_Desk) — similar · Opportunities
