DBRaven
Partition HotspotCritical

Partition Hotspot Amplification

Sequential or low-cardinality partition keys concentrate writes on a single partition or shard, creating a hotspot. The hot partition absorbs disproportionate I/O, CPU, and memory pressure while other partitions sit idle. As load grows, the hotspot becomes a hard throughput ceiling for the entire table or topic.

Partitioned PostgreSQL table or Kafka topic with sequential key pattern

Degradation Replay

Stage 1

Nominal: Even Distribution

Nominal
Trigger

Write load evenly distributed across partitions

Operational Metrics
Hot Partition Load Share
18 %
warn 40crit 70
Partition Size Variance
8 %
warn 50crit 200
Symptoms
  • ·Partition sizes within 10% of each other
  • ·No single partition consuming >30% of total table I/O
Topology Effects
  • ·All partition host nodes at similar load
  • ·Buffer cache pages evenly spread across partitions
Operational Consequences
  • !All partitions serving reads and writes at equal throughput

Operational simulation model only, not a production forecast. Degradation stages are derived from structured operational knowledge, not measured telemetry. Do not use for capacity planning or incident response.

Run With Your Parameters

Adjust the parameters below to see how metric values shift across degradation stages. Formulas are deterministic: same inputs always produce the same output.

Simulation Parameters

Fraction of traffic hitting the hottest partition

Computed Degradation Stages

nominal·Nominal: Even Distribution

Write load evenly distributed across partitions

Total RPS
4,000req/s
warn: 8,000crit: 10,000
Hot Partition RPS
1,600req/s
warn: 1,000crit: 2,000
Partition Skew Ratio
3.2×
warn: 2crit: 5
Hot Partition Pool Utilization
16%
warn: 70crit: 90
Balanced Partition RPS (expected)
500req/s
degraded·Hotspot Forming: Sequential Key Concentration

Sequential insert pattern emerging: all new rows targeting last partition

Total RPS
6,500req/s
warn: 8,000crit: 10,000
Hot Partition RPS
2,600req/s
warn: 1,625crit: 3,250
Partition Skew Ratio
3.2×
warn: 2crit: 5
Hot Partition Pool Utilization
26%
warn: 70crit: 90
Balanced Partition RPS (expected)
812.5req/s
warning·Warning: I/O Ceiling Approaching

Hot partition host I/O utilization exceeds 75%; write queue forming

Total RPS
8,500req/s
warn: 8,000crit: 10,000
Hot Partition RPS
3,400req/s
warn: 2,125crit: 4,250
Partition Skew Ratio
3.2×
warn: 2crit: 5
Hot Partition Pool Utilization
34%
warn: 70crit: 90
Balanced Partition RPS (expected)
1,062.5req/s
critical·Critical: Write Throughput Ceiling Reached

Hot partition I/O saturated; all writes queuing behind disk bottleneck

Total RPS
11,000req/s
warn: 8,000crit: 10,000
Hot Partition RPS
4,400req/s
warn: 2,750crit: 5,500
Partition Skew Ratio
3.2×
warn: 2crit: 5
Hot Partition Pool Utilization
44%
warn: 70crit: 90
Balanced Partition RPS (expected)
1,375req/s

Threshold Events

Severe Partition Skewwarning

Hot partition receiving 1,600 RPS (3.2× balanced expectation). Balanced load would be 500 RPS per partition.

threshold: 3actual: 3.2

Interpretation

warning

40% of traffic (10,000 RPS) concentrates on one partition. Hot partition receives 3.2× balanced load. Connection pool: 44%.

Bottleneck

Traffic skew 3.2× exceeds partition capacity: hot shard saturation

Recommendation

Re-key the hot partition using a high-cardinality composite key. Add a dedicated read replica for the hot partition. Consider hash-range hybrid partitioning to redistribute hot key space.

Parameterized simulation: not a production forecast. Values derived from deterministic formulas applied to your parameters. Do not use for capacity planning or operational decisions without validation.

Propagation Model

Step
Hot PartitionPartition Host Node, Buffer Cache

I/O and CPU pressure concentrates on partition host; buffer cache fills with hot partition pages, evicting other data

Stabilizes: Stabilizes only when write load is redistributed via key salting or repartitioning

Threshold
Hot Partition I/OWrite Throughput Ceiling

Hot partition IOPS reaches disk throughput limit; writes begin queuing behind I/O bottleneck

Stabilizes: No stabilization possible without architectural change: throughput ceiling is hard

Cascademild amplification
Partition HostQuery Planner, Index Scans

Range queries on sequential key require full hot-partition scans; query planner estimates stale for hot partition

Stabilizes: Resolved by ANALYZE + histogram updates

Recovery Patterns

Key salting for new writes

Immediate for new writes; existing data remains imbalanced until repartitioned
Tradeoffs
  • ·Range queries on salted key require scatter-gather across all salt buckets
  • ·Application must understand salt prefix for direct lookups
Residual Risks
  • !Historical data remains on original partition: read hotspot persists for old data queries

Full table repartitioning

Hours to days depending on table size; requires online DDL or pg_partman
Tradeoffs
  • ·Requires downtime or dual-write during migration
  • ·Temporary storage overhead for parallel partition structure
Residual Risks
  • !New partition key must be validated against all query patterns: poor choice creates new hotspot

Operational Summary

Partition hotspots arise from poor key choice: sequential, low-cardinality, or monotonically increasing keys concentrate writes on one partition. This creates a hard throughput ceiling that cannot be resolved by vertical scaling or adding more instances. The fix requires redistributing the key space, which is a significant operational undertaking on large tables.

Partition Hotspot Amplification: DBRaven