AI-Native Platform Operations
B2B SaaS Platform Team
15–30 min → minutes to diagnose
Replaced tribal on-call runbooks with 11 MCP observability playbooks — health checks, incident triage, and trace investigation without shell access.
The challenge
On-call engineers relied on tribal knowledge and ad-hoc shell commands to diagnose production issues. New team members struggled during incidents, and granting broad shell access raised security concerns for a compliance-conscious SaaS platform.
Our approach
- Interviewed on-call engineers to capture the top 11 recurring diagnostic workflows.
- Encoded each workflow as an MCP observability playbook with explicit guardrails.
- Hermes Agent executes health checks, log queries, and trace investigation via MCP — no direct shell.
- Added confirmation gates before any write or restart actions.
- Integrated playbooks into the existing incident channel for one-click triage.
Results
- Typical incident diagnosis dropped from 15–30 minutes to minutes for known patterns.
- New engineers can run the same triage playbooks without months of tribal knowledge.
- Reduced need for broad production shell access — MCP tools are scoped and auditable.
- 11 playbooks cover the majority of recurring on-call scenarios.
Technology
More case studies
GrowthDesk Expansion Platform
6-stage pipeline · one platform
Unified multi-site expansion workspace — financial modeling, interactive maps, AI feasibility research, and multi-stakeholder pipelines in one secure platform.
Local AI InfrastructureMac Studio Hermes Agent Server
Private AI · secure remote access
Dedicated Mac Studio AI server — Ollama local LLMs, Hermes Agent, Obsidian MCP, and secure mesh VPN remote access with isolated macOS accounts per team.
Ready for similar outcomes?
Tell us about your platform, ops, or content goals — we'll map a path from assessment to production.