Goblin Core Pub/Sub Performance Benchmark

Generated by the native C++ harness on the dedicated benchmark host on July 15, 2026.

Summary

With one publisher and one subscriber, Goblin Core's typed SBE path over a 4 KiB shared-memory ring delivers a 16-byte message in 0.301 us at p50 and 0.391 us at p99. That is 39.8x and 39.3x faster, respectively, than the fastest incumbent result. The closed-loop workload completes 2.63 million fully acknowledged and delivered publications per second, 32.7x the fastest incumbent.

The advantage survives fanout: at 64 listeners, the ring reaches the last listener in 12.704 us at p50 and completes 77,042 publications/s, 12.8x and 13.7x the fastest incumbent results. Literal-channel lookup stays flat at 0.311-0.321 us from one through 16,384 subscriptions.

Pattern matching is a second win, not just a transport win. With 8,192 patterns, Goblin's RESP2/UDS path reaches the matching subscriber in 62.709 us, 3.95x faster than the fastest incumbent; SBE/ring reduces that to 43.512 us, 5.69x faster. Goblin's conventional RESP2/UDS path remains within 5% of the fastest incumbent on the one-listener literal cases, so the ring result is not hiding a slow server behind a fast transport.

Goblin also has the lowest total server RSS in every comparable large subscription topology here. Its incremental subscription structures are not always the smallest: Redis 7.2.4 has the lowest delta at 16,384 literal channels, while Redis 8.8 and Valkey lead the 32-client deltas. Both total RSS and baseline-subtracted deltas are reported below.

Method

All latency columns below are microseconds. Lower latency is better; higher publication rate is better.

One Publisher, One Subscriber

The payload is delivered to one literal subscriber. p99 and p99.9 are shown for the 16-byte hot-path case; the final columns show payload scaling at p50.

Engine 16 B p50 16 B p99 16 B p99.9 16 B pub/s 256 B p50 4,096 B p50
goblin-sbe-ring-4kb 0.301 0.391 0.812 2,630,080 0.321 n/a
goblin-resp-uds 12.504 15.219 17.974 77,942 12.634 14.928
redis-7.2.4 12.213 15.820 18.505 80,354 12.464 15.980
redis-8.8 12.474 15.930 19.317 77,446 13.916 15.900
valkey-9.1 12.514 16.120 19.166 76,610 14.017 15.910
dragonfly 11.953 15.369 3,527.150 33,700 12.303 15.069
mini-redis-go-55178df 13.015 21.381 27.282 70,109 13.315 20.128

The ring's median changes by only 0.020 us between 16- and 256-byte payloads. On the shared UDS path, Goblin is 4.6% behind Dragonfly's fastest 16-byte median, but has the lowest p99.9 of every socket row.

Literal Fanout

Each publication targets one literal channel shared by every listener. Delivery latency ends when the last listener receives the message.

Engine 1 listener 8 listeners 32 listeners 64 listeners 64 p99 64 p99.9 64 pub/s
goblin-sbe-ring-4kb 0.301 1.533 6.713 12.704 16.171 26.350 77,042
goblin-resp-uds 12.424 29.005 90.541 176.594 188.156 210.108 5,643
redis-7.2.4 12.143 29.075 94.349 174.791 186.132 1,000.070 5,607
redis-8.8 12.474 29.245 95.571 177.556 189.859 220.357 5,617
valkey-9.1 12.764 29.335 99.298 180.371 193.356 225.697 5,522
dragonfly 12.023 27.181 86.003 161.987 3,746.550 3,932.040 2,998
mini-redis-go-55178df 13.075 38.633 117.322 220.447 388.475 708.651 4,450

At 64 listeners, Goblin SBE/ring is 12.8x faster at p50 and 11.5x faster at p99 than the best incumbent cells. Goblin RESP2/UDS tracks the Redis-family implementations closely and has the tightest socket p99.9 in this row.

Literal Routing Scale

One client subscribes to the listed number of distinct literal channels. Every publication targets the same subscribed channel, so this isolates lookup cost as the channel table grows.

Engine 1 channel 1,024 channels 16,384 channels 16,384 p99 16,384 pub/s
goblin-sbe-ring-4kb 0.321 0.311 0.321 0.441 2,488,380
goblin-resp-uds 12.444 12.464 12.474 15.840 78,408
redis-7.2.4 12.133 12.153 12.273 15.740 78,763
redis-8.8 12.444 12.434 12.564 16.341 76,762
valkey-9.1 13.155 13.836 13.335 16.151 75,198
dragonfly 11.943 11.963 11.973 15.589 33,562
mini-redis-go-55178df 13.065 13.035 13.085 22.873 67,589

Goblin's literal lookup is flat on both transports. At 16,384 channels, the ring is 37.3x faster at p50 and 35.4x faster at p99 than the best incumbent.

Pattern Routing

Pattern subscriptions use Redis-compatible glob matching. Pattern count is the scaling variable, the hit target matches one pattern, and Goblin examines every registered pattern for each publication. mini-redis-go at the tested revision does not implement PSUBSCRIBE, so its pattern cells are n/a.

Matching Publication

Engine 1 pattern 64 patterns 1,024 patterns 8,192 patterns 8,192 p99 8,192 pub/s
goblin-sbe-ring-4kb 0.351 0.731 5.390 43.512 50.205 22,764
goblin-resp-uds 12.464 13.856 18.826 62.709 68.670 15,841
redis-7.2.4 12.273 14.518 44.394 367.355 389.397 2,705
redis-8.8 13.916 15.279 49.975 393.976 438.239 2,530
valkey-9.1 13.836 14.898 45.446 330.686 363.939 3,003
dragonfly 12.013 13.806 40.346 247.789 4,022.980 2,096
mini-redis-go-55178df n/a n/a n/a n/a n/a n/a

Unmatched Publication

With no delivery to wait for, this table reports the PUBLISH acknowledgement latency. It directly exposes the cost of scanning patterns and finding no match.

Engine 0 patterns 64 patterns 1,024 patterns 8,192 patterns 8,192 p99
goblin-sbe-ring-4kb 0.180 0.561 5.871 43.011 48.712
goblin-resp-uds 7.584 10.550 14.808 58.541 64.492
redis-7.2.4 10.640 10.440 30.838 282.905 321.017
redis-8.8 9.648 11.101 33.403 266.584 300.559
valkey-9.1 10.490 11.943 32.181 241.347 271.273
dragonfly 12.484 13.435 29.716 153.921 173.879
mini-redis-go-55178df 10.249 n/a n/a n/a n/a

Pattern cost is still linear; Goblin has a substantially lower constant factor. At 8,192 patterns, Goblin RESP2/UDS is 3.95x faster than the best incumbent hit median and 2.63x faster on a miss. The ring extends the hit lead to 5.69x.

Many Clients and Subscriptions

Thirty-two subscribers each hold 32 private subscriptions plus one shared subscription. The publication targets the shared literal channel or shared pattern, producing 32 validated deliveries.

Engine literal p50 literal p99 literal pub/s pattern p50 pattern p99 pattern pub/s
goblin-sbe-ring-4kb 6.623 7.344 145,663 12.283 16.151 78,749
goblin-resp-uds 90.551 97.765 11,016 98.266 108.014 10,097
redis-7.2.4 94.098 100.861 10,631 144.724 151.797 6,911
redis-8.8 95.140 101.652 10,509 151.066 158.580 6,605
valkey-9.1 97.835 104.067 10,170 147.539 155.204 6,718
dragonfly 85.973 3,024.900 5,838 126.850 3,725.870 3,995
mini-redis-go-55178df 116.781 149.082 8,338 n/a n/a n/a

SBE/ring completes the literal topology 13.0x faster at p50 and the pattern topology 10.3x faster than the best incumbent. Goblin RESP2/UDS is 5.3% behind Dragonfly's literal p50 while completing 1.89x as many closed-loop publications; on patterns, Goblin RESP2/UDS leads the best incumbent p50 by 1.29x.

Server Memory

Each cell is total RSS / RSS delta in MiB after constructing the topology. The delta subtracts that fresh server's empty baseline. Total RSS is the fair view of the complete process; the delta helps isolate subscription growth. Goblin ring mappings exist before the baseline, while client output buffers are activated by subscription, so neither number should be read without the other. Small RSS deltas are also subject to page granularity and allocator reuse.

Engine 16,384 literals 8,192 patterns 32x32 literals 32x32 patterns
goblin-sbe-ring-4kb 9.60 / 4.41 7.96 / 2.77 6.64 / 0.55 6.73 / 0.64
goblin-resp-uds 9.56 / 4.43 7.93 / 2.79 6.01 / 0.90 6.17 / 1.04
redis-7.2.4 10.66 / 3.86 8.89 / 2.11 7.29 / 0.52 7.23 / 0.52
redis-8.8 11.95 / 4.08 10.05 / 2.19 8.29 / 0.41 8.28 / 0.41
valkey-9.1 12.44 / 4.83 10.26 / 2.71 7.96 / 0.41 7.92 / 0.43
dragonfly 27.67 / 5.32 25.33 / 3.03 24.55 / 2.19 24.39 / 1.95
mini-redis-go-55178df 17.59 / 6.66 n/a 12.66 / 1.74 n/a

Goblin has the lowest total server RSS in all four topologies. Redis and Valkey sometimes allocate less incremental memory after their larger baselines, which is why both sides of the slash are retained.

What This Benchmark Does Not Claim

Reproducing

Build the server and native driver, then run the matrix:

cmake -S . -B build-rel -DCMAKE_BUILD_TYPE=Release \
  -DGOBLIN_CORE_ARCH=avx2 -DGOBLIN_CORE_BUILD_BENCHMARKS=ON
cmake --build build-rel -j --target \
  goblin_core_server goblin_core_pubsub_benchmark
bash benchmarks/pubsub_benchmark.sh

The launcher paths are environment-overridable; no machine-specific binary paths are required by the benchmark source. It writes the complete 26-column CSV, including minimum, p50, p90, p99, p99.9, mean, publication rate, baseline RSS, subscribed RSS, and RSS delta.

The command behavior is documented in the Pub/Sub reference. The low-latency transport and wire format are documented in ring buffers and the SBE protocol.

Tested Servers