Master technical document — generalized and de-branded. Owner: architect.
System context
Sources: Splunk, ServiceNow, AppDynamics, PagerDuty, Grafana, Datadog, New Relic, SNMP traps, custom API.
Downstream: incident handoff into the customer’s ITSM.
Dependency: the CMDB/CSDM truth graph.
The seven-stage pipeline
For each stage document: purpose, technology, scaling unit, failure mode.
Streaming backbone
Kafka (KRaft) topics & partitioning strategy · Flink jobs with RocksDB state and checkpointing · Debezium CDC · MinIO object storage · DLQ design and replay path.
Control tower (self-monitoring)
The platform monitors its own health so a degraded component is detected before it corrupts incident accuracy. Watched: Flink, Kafka, Redis, Neo4j, ClickHouse, MinIO, Debezium. 18 self-metrics feed the go-live dashboards.
Runbook matching → approval thresholds (auto-approve low-risk; sign-off for critical systems) → ITSM handoff → outcome feedback into the graph. ORCA recommends; you decide.
Non-functional targets
T1 ingestion→incident p99 < 5s · component recovery targets are measured in chaos testing · availability model.