← Home
Home/Command Center
Command Center
Live incident lifecycle, safety gate, approval, and audit trail
Mission Dark
OpsPilot AI

AI-Powered Incident Command Center

Transform operational alerts into structured investigations, root-cause hypotheses, remediation plans, and postmortems using specialized Qwen AI agents deployed on Alibaba Cloud.

live
7 AI Agents

Triage to postmortem

gated
Human Approval

Guarded remediation

deployed
Alibaba Cloud ECS

Live backend deployment

Priority metrics

Current state
Live signals
0
slate
Active incident
No active incident yet
slate
Current state
standby
slate
Assigned owner
unassigned
amber
Backend
not checked
slate

Live alert stream

Server-sent events stream live operational alerts from the backend.
connecting

Incoming operational alerts from the backend stream. Promote a signal when it needs incident command.

SSE feed

Waiting for the first live alert...

The stream will populate automatically when the backend emits operational signals.

Command Centerstandbyhuman approval gated

Incident lifecycle console

Trace evidence, review policy decisions, approve remediation, and generate an auditable postmortem.

Incident state machine

Step 1
triaging
Step 2
investigating
Step 3
hypothesis
Step 4
approval gate
Step 5
remediating
Step 6
monitoring
Step 7
resolved

Command intelligence

Operational impact, service topology, SLA risk, and guarded response actions in one compact view.

human approval gated

Topology map

1 affected
1
edge-gateway
healthy
2
auth-service
healthy
3
checkout-api
affected
4
cache-layer
healthy
5
payment-api
healthy
Blast radius
Users impacted
amber
0
Affected services
Service
cyan
1
SLA tracker
SLA remaining
green
116m
Incident cost
Revenue at risk
violet
$0
MTTR / MTTD
MTTR
green
30m
MTTR / MTTD
MTTD
green
5m

Rollback trigger

approval required
Last known good
cache-config-v18

Runbook executor

deploymentspending
Rollback planpending
Validationpending

Advanced AI cockpit

Predictive signals, learned patterns, capacity risk, safe automation, controlled failure checks, and generated runbook guidance.

confidence

Predictive alerting

green
watch mode
checkout-api · confidence 81%

Auto-remediation guardrail

green
low-risk autonomous
Risk: P3

Pattern library

cyan
3 learned patterns
similarity: 72%

Capacity forecast

green
68% next limit
forecast window: 45m

Chaos readiness

slate
controlled experiment
pending

Runbook generation

green
generated runbook
Rollback cache configuration safely

Incident summary

No active incident yet

standby
Incident
No active incident yet
P3
Service
checkout-api
Status
standby
Assignee
unassigned
Age
live

Severity: P3 · Status: standby

Service
checkout-api
Environment
production
Severity
high
Assignee
unassigned
Business impact
checkout degraded
Confidence
0.86
Risk
medium

Leading hypothesis

Cache configuration regression introduced in the latest config change

Collaboration hub

Keep the incident channel, roles, on-call routing, and stakeholder message in one place.

commander

War room mode

Dedicated incident channel with current context and responder ownership.

connecting
Channel
#ops-standby
Participants
commander, unassigned

Role-based access

Switch role to preview permissions without changing authentication.

read-onlycan annotatecan assigncan approve

On-call schedule

Primary and secondary responders for the current operational window.

Primary
sre-primary
Secondary
platform-responder

Shared timeline

Add a responder note to the incident event history.

Evidence console

Evidence lineage connects metrics, logs, deployments, and runbooks to the recommendation.

Evidence lineage behind the recommendation.

traceable
source 1
p95 latency increased from 420ms to 2.8s
source 2
cache hit ratio dropped from 91% to 41%
source 3
database latency increased by 63%

Why multi-agent?

Each agent owns one decision boundary: triage, evidence, hypothesis, risk, approval, execution review, and postmortem.

AI useful layer

Natural language summary

AI useful layer

Start or select an incident to generate AI insights.

Live agent timeline

Agent progress plus recorded incident events.

0 events
Triage
Agent step 1 / 9
waiting
Observability
Agent step 2 / 9
waiting
Runbook
Agent step 3 / 9
waiting
Hypothesis
Agent step 4 / 9
waiting
Planner
Agent step 5 / 9
waiting
Safety
Agent step 6 / 9
waiting
Approval Gate
Agent step 7 / 9
waiting
Execution Review
Agent step 8 / 9
waiting
Postmortem
Agent step 9 / 9
waiting

Incident event log

Incident event log

Create an incident from the live alert stream to populate event history.

Execution Review

What happened after approval and how the platform verified recovery.

pending
Action
Waiting approval
Validation
Not started
Risk
medium
p95 latency
2.8s
Cache hit ratio
41%

Postmortem preview

The incident was likely caused by cache config regression and mitigated with approval-gated rollback.

draft

Stakeholder update

One-click operational update for leadership and customer-facing teams.

No active incident yet | checkout-api | Severity: P3 | Status: standby | Assignee: unassigned. Recommended action: Rollback cache configuration safely.

Post-mortem template

Structured review fields populated from the current incident context.

Impact
pending
Root cause
Cache configuration regression introduced in the latest config change
Mitigation
Rollback cache configuration safely
Follow-ups
pending