DBRaven
Replay StormCritical

Event Replay Storm Recovery

A large-scale event replay: triggered by a projection rebuild, a bug fix requiring reprocessing, or a new consumer catching up from offset zero: floods the event store and downstream systems with a burst of historical events. The replay storm saturates database write capacity, overloads downstream services not designed for burst replay, and surfaces non-idempotent handlers that produce duplicate side effects.

Event store (Kafka or EventStoreDB) + projection processor + downstream OLTP database + external API consumers

Degradation Replay

Stage 1

Nominal: Production Event Rate

Nominal
Trigger

Consumer processing at normal production event rate; no replay in progress

Operational Metrics
Event Processing Rate
800 events/s
warn 5,000crit 15,000
Downstream DB Write Rate
450 writes/s
warn 2,000crit 5,000
Duplicate Event Rate
0 events/s
warn 10crit 100
Symptoms
  • ·Event consumption rate at normal production level
  • ·Downstream DB write rate within normal operating range
  • ·No duplicate events or double-processing in handler logs
Topology Effects
  • ·Consumers tracking near-current offset on topic partitions
  • ·Downstream services processing events within their provisioned capacity
Operational Consequences
  • !Normal operation: consumers processing production events in real time

Operational simulation model only, not a production forecast. Degradation stages are derived from structured operational knowledge, not measured telemetry. Do not use for capacity planning or incident response.

Run With Your Parameters

Adjust the parameters below to see how metric values shift across degradation stages. Formulas are deterministic: same inputs always produce the same output.

Simulation Parameters

Rate multiplier during dead-letter replay burst

Computed Degradation Stages

nominal·Nominal: Production Event Rate

Consumer processing at normal production event rate; no replay in progress

Message Produce Rate
10,000msg/s
warn: 4,500crit: 6,750
Consumer Throughput
4,500msg/s
warn: 8,000crit: 5,000
Message Backlog
0messages
warn: 100,000crit: 300,000
Consumer Health
100%
warn: 80crit: 60
Message Processing Lag
0seconds
warn: 30crit: 120
degraded·Replay Initiated: Storm Ramp-Up

Replay started: consumer offset reset to historical point; processing at maximum throughput

Message Produce Rate
16,250msg/s
warn: 4,500crit: 6,750
Consumer Throughput
4,500msg/s
warn: 13,000crit: 8,125
Message Backlog
352,500messages
warn: 100,000crit: 300,000
Consumer Health
60%
warn: 80crit: 60
Message Processing Lag
78.33seconds
warn: 30crit: 120
warning·Warning: Downstream Saturation

Replay write burst saturates downstream database; OLTP transactions queuing

Message Produce Rate
21,250msg/s
warn: 4,500crit: 6,750
Consumer Throughput
4,500msg/s
warn: 17,000crit: 10,625
Message Backlog
2,010,000messages
warn: 100,000crit: 300,000
Consumer Health
60%
warn: 80crit: 60
Message Processing Lag
446.67seconds
warn: 30crit: 120
critical·Critical: Data Integrity Failure + Downstream Cascade

Non-idempotent handlers have produced data corruption; external APIs triggering circuit breakers

Message Produce Rate
27,500msg/s
warn: 4,500crit: 6,750
Consumer Throughput
4,500msg/s
warn: 22,000crit: 13,750
Message Backlog
6,900,000messages
warn: 100,000crit: 300,000
Consumer Health
60%
warn: 80crit: 60
Message Processing Lag
1,533.33seconds
warn: 30crit: 120
recovery·Recovery: Controlled Replay with Throttling

Handlers fixed for idempotency; replay resumed with throttling; OLTP recovering

Message Produce Rate
12,500msg/s
warn: 4,500crit: 6,750
Consumer Throughput
4,500msg/s
warn: 10,000crit: 6,250
Message Backlog
480,000messages
warn: 100,000crit: 300,000
Consumer Health
90%
warn: 80crit: 60
Message Processing Lag
106.67seconds
warn: 30crit: 120

Threshold Events

Backlog Exceeds Warning Thresholdcritical

Message backlog at 352,500 (3.5× threshold). Consumers processing at 4,500 msg/s vs 16,250 produce rate.

threshold: 100,000actual: 352,500
Consumer Health Degradedcritical

Consumer health at 60%: sustained backlog causes GC pressure and memory growth.

threshold: 80actual: 60

Interpretation

critical

Produce rate at 5× storm rate exceeds consumer capacity (4,500 msg/s). Backlog peaks at 6,900,000 messages. Consumer health: 60%.

Bottleneck

Producer throughput exceeds consumer capacity: backlog accumulates unboundedly

Recommendation

Scale consumers to at least 17 instances. Add consumer health monitoring with backpressure. Set dead-letter queue with bounded retry budget.

Parameterized simulation: not a production forecast. Values derived from deterministic formulas applied to your parameters. Do not use for capacity planning or operational decisions without validation.

Propagation Model

Exponentialsevere amplification
Replay ProcessorDownstream Database, External API Consumers

Replay at maximum throughput produces N× normal event rate; downstream systems receive burst they were not provisioned for

Stabilizes: Storm subsides when replay reaches current offset: returns to normal production event rate

Cascademoderate amplification
Non-Idempotent HandlerOLTP Database, External Services

Non-idempotent handlers create duplicate records, double-counting, or repeated API calls for each replayed event

Stabilizes: Resolved only by idempotency fix + compensating data cleanup: cannot auto-heal

Thresholdmoderate amplification
Downstream Database Write RateOLTP Transaction Queue, Read Performance

Replay-driven write burst saturates database IOPS; OLTP transactions queue behind replay writes; user-facing reads degrade

Stabilizes: OLTP performance recovers once replay rate is throttled or replay completes

Recovery Patterns

Stop, fix idempotency, resume with throttling

2-8 hours for idempotency fix + data reconciliation + controlled replay
Tradeoffs
  • ·Delayed replay completion: business feature depending on rebuilt projection has longer wait
  • ·Throttled replay impacts replay completion time proportionally
Residual Risks
  • !Data reconciliation may be incomplete: audit trail must confirm all duplicates identified

Replay via separate dedicated cluster

Replay time unchanged; OLTP cluster protected from burst
Tradeoffs
  • ·Requires provisioning a temporary replay cluster: cost and operational overhead
  • ·Results must be merged back to production cluster
Residual Risks
  • !Merge step introduces its own data integrity risk: must be carefully validated

Operational Summary

Event replay storms occur when historical event reprocessing runs at maximum throughput, flooding downstream systems with a burst N× normal production rate. The critical failure mode is non-idempotent event handlers: replaying events creates duplicate records, double-counted aggregates, or repeated API calls that cannot auto-heal. Prevention requires idempotent handlers (event ID deduplication), dedicated replay consumer groups, and strict rate throttling relative to downstream write capacity.