Skip to main content
Master technical document — generalized and de-branded. Owner: architect.

System context

Sources: Splunk, ServiceNow, AppDynamics, PagerDuty, Grafana, Datadog, New Relic, SNMP traps, custom API. Downstream: incident handoff into the customer’s ITSM. Dependency: the CMDB/CSDM truth graph.

The seven-stage pipeline

For each stage document: purpose, technology, scaling unit, failure mode.

Streaming backbone

Kafka (KRaft) topics & partitioning strategy · Flink jobs with RocksDB state and checkpointing · Debezium CDC · MinIO object storage · DLQ design and replay path.

Control tower (self-monitoring)

The platform monitors its own health so a degraded component is detected before it corrupts incident accuracy. Watched: Flink, Kafka, Redis, Neo4j, ClickHouse, MinIO, Debezium. 18 self-metrics feed the go-live dashboards.

Human-in-the-loop remediation

Runbook matching → approval thresholds (auto-approve low-risk; sign-off for critical systems) → ITSM handoff → outcome feedback into the graph. ORCA recommends; you decide.

Non-functional targets

T1 ingestion→incident p99 < 5s · component recovery targets are measured in chaos testing · availability model.