Goblin Core Pub/Sub Performance Benchmark
Generated by the native C++ harness on the dedicated benchmark host on July 15, 2026.
Summary
With one publisher and one subscriber, Goblin Core's typed SBE path over a 4 KiB shared-memory ring delivers a 16-byte message in 0.301 us at p50 and 0.391 us at p99. That is 39.8x and 39.3x faster, respectively, than the fastest incumbent result. The closed-loop workload completes 2.63 million fully acknowledged and delivered publications per second, 32.7x the fastest incumbent.
The advantage survives fanout: at 64 listeners, the ring reaches the last listener in 12.704 us at p50 and completes 77,042 publications/s, 12.8x and 13.7x the fastest incumbent results. Literal-channel lookup stays flat at 0.311-0.321 us from one through 16,384 subscriptions.
Pattern matching is a second win, not just a transport win. With 8,192 patterns, Goblin's RESP2/UDS path reaches the matching subscriber in 62.709 us, 3.95x faster than the fastest incumbent; SBE/ring reduces that to 43.512 us, 5.69x faster. Goblin's conventional RESP2/UDS path remains within 5% of the fastest incumbent on the one-listener literal cases, so the ring result is not hiding a slow server behind a fast transport.
Goblin also has the lowest total server RSS in every comparable large subscription topology here. Its incremental subscription structures are not always the smallest: Redis 7.2.4 has the lowest delta at 16,384 literal channels, while Redis 8.8 and Valkey lead the 32-client deltas. Both total RSS and baseline-subtracted deltas are reported below.
Method
benchmarks/pubsub_benchmark.cppis a native C++23 publisher and subscriber harness. Python,redis-cli, andredis-benchmarkare absent from the measured path.- The publisher issues one
PUBLISH, waits for its integer acknowledgement and every expected delivery, then issues the next publication. There is no pipelining or backlog. The reported publication rate is this complete closed-loop cycle rate. - Acknowledgement latency runs from the publisher's send until the
PUBLISHreply. Delivery latency runs from the same start timestamp until the final subscriber receives and validates the channel and payload. A delivery may precede the publisher acknowledgement; Dragonfly exhibits this in several rows. - The one-to-one group measures
100,000publications after10,000warmups. Fanout and ordinary routing rows measure20,000after2,000warmups. The 8,192-pattern rows use5,000samples after500warmups. - Every scenario starts a fresh server, and engines run one at a time. Goblin, Redis, Valkey, and mini-redis-go use server CPU
2; Dragonfly manages its own affinity with one proactor. The publisher uses CPU3, and subscriber reader threads start on CPU4. The 64-listener case avoids the server and publisher's SMT siblings. - The host is a quiet AMD Ryzen Threadripper PRO 5995WX with 64 cores. Results were collected from Release AVX2 builds on Linux. Timing uses
rdtscp, calibrated against the monotonic clock. - Goblin is measured twice: typed SBE over shared-memory rings and RESP2 over a Unix-domain socket. Every incumbent uses RESP2/UDS through the same native RESP encoder, parser, validation, and timer.
- Every Goblin ring is exactly
4 KiB. Large subscription setup requests are split into 32-name batches outside the timed region. A 4,096-byte payload plus SBE and ring framing cannot fit in a 4 KiB ring, so that cell isn/arather than a smaller, incomparable payload. - Goblin uses an
8 KiBper-client unsolicited-output allocation in this run. Messages are drained synchronously, so the benchmark never relies on queueing up to that limit. - Redis and Valkey use
benchmarks/redis-parity.conf. Dragonfly uses one proactor thread for single-core parity. mini-redis-go usesGOMAXPROCS=1, with AOF and metrics disabled. - RSS is read from the launched server PID, never from a server-reported memory field. mini-redis-go uses
ps -o rss=; the other engines use Linux/proc/<pid>/statusasVmRSS + HugetlbPages. - Incumbents are treated strictly as black-box servers. Their source code is not inspected.
All latency columns below are microseconds. Lower latency is better; higher publication rate is better.
One Publisher, One Subscriber
The payload is delivered to one literal subscriber. p99 and p99.9 are shown for the 16-byte hot-path case; the final columns show payload scaling at p50.
| Engine | 16 B p50 | 16 B p99 | 16 B p99.9 | 16 B pub/s | 256 B p50 | 4,096 B p50 |
|---|---|---|---|---|---|---|
goblin-sbe-ring-4kb |
0.301 | 0.391 | 0.812 | 2,630,080 | 0.321 | n/a |
goblin-resp-uds |
12.504 | 15.219 | 17.974 | 77,942 | 12.634 | 14.928 |
redis-7.2.4 |
12.213 | 15.820 | 18.505 | 80,354 | 12.464 | 15.980 |
redis-8.8 |
12.474 | 15.930 | 19.317 | 77,446 | 13.916 | 15.900 |
valkey-9.1 |
12.514 | 16.120 | 19.166 | 76,610 | 14.017 | 15.910 |
dragonfly |
11.953 | 15.369 | 3,527.150 | 33,700 | 12.303 | 15.069 |
mini-redis-go-55178df |
13.015 | 21.381 | 27.282 | 70,109 | 13.315 | 20.128 |
The ring's median changes by only 0.020 us between 16- and 256-byte payloads. On the shared UDS path, Goblin is 4.6% behind Dragonfly's fastest 16-byte median, but has the lowest p99.9 of every socket row.
Literal Fanout
Each publication targets one literal channel shared by every listener. Delivery latency ends when the last listener receives the message.
| Engine | 1 listener | 8 listeners | 32 listeners | 64 listeners | 64 p99 | 64 p99.9 | 64 pub/s |
|---|---|---|---|---|---|---|---|
goblin-sbe-ring-4kb |
0.301 | 1.533 | 6.713 | 12.704 | 16.171 | 26.350 | 77,042 |
goblin-resp-uds |
12.424 | 29.005 | 90.541 | 176.594 | 188.156 | 210.108 | 5,643 |
redis-7.2.4 |
12.143 | 29.075 | 94.349 | 174.791 | 186.132 | 1,000.070 | 5,607 |
redis-8.8 |
12.474 | 29.245 | 95.571 | 177.556 | 189.859 | 220.357 | 5,617 |
valkey-9.1 |
12.764 | 29.335 | 99.298 | 180.371 | 193.356 | 225.697 | 5,522 |
dragonfly |
12.023 | 27.181 | 86.003 | 161.987 | 3,746.550 | 3,932.040 | 2,998 |
mini-redis-go-55178df |
13.075 | 38.633 | 117.322 | 220.447 | 388.475 | 708.651 | 4,450 |
At 64 listeners, Goblin SBE/ring is 12.8x faster at p50 and 11.5x faster at p99 than the best incumbent cells. Goblin RESP2/UDS tracks the Redis-family implementations closely and has the tightest socket p99.9 in this row.
Literal Routing Scale
One client subscribes to the listed number of distinct literal channels. Every publication targets the same subscribed channel, so this isolates lookup cost as the channel table grows.
| Engine | 1 channel | 1,024 channels | 16,384 channels | 16,384 p99 | 16,384 pub/s |
|---|---|---|---|---|---|
goblin-sbe-ring-4kb |
0.321 | 0.311 | 0.321 | 0.441 | 2,488,380 |
goblin-resp-uds |
12.444 | 12.464 | 12.474 | 15.840 | 78,408 |
redis-7.2.4 |
12.133 | 12.153 | 12.273 | 15.740 | 78,763 |
redis-8.8 |
12.444 | 12.434 | 12.564 | 16.341 | 76,762 |
valkey-9.1 |
13.155 | 13.836 | 13.335 | 16.151 | 75,198 |
dragonfly |
11.943 | 11.963 | 11.973 | 15.589 | 33,562 |
mini-redis-go-55178df |
13.065 | 13.035 | 13.085 | 22.873 | 67,589 |
Goblin's literal lookup is flat on both transports. At 16,384 channels, the ring is 37.3x faster at p50 and 35.4x faster at p99 than the best incumbent.
Pattern Routing
Pattern subscriptions use Redis-compatible glob matching. Pattern count is the scaling variable, the hit target matches one pattern, and Goblin examines every registered pattern for each publication. mini-redis-go at the tested revision does not implement PSUBSCRIBE, so its pattern cells are n/a.
Matching Publication
| Engine | 1 pattern | 64 patterns | 1,024 patterns | 8,192 patterns | 8,192 p99 | 8,192 pub/s |
|---|---|---|---|---|---|---|
goblin-sbe-ring-4kb |
0.351 | 0.731 | 5.390 | 43.512 | 50.205 | 22,764 |
goblin-resp-uds |
12.464 | 13.856 | 18.826 | 62.709 | 68.670 | 15,841 |
redis-7.2.4 |
12.273 | 14.518 | 44.394 | 367.355 | 389.397 | 2,705 |
redis-8.8 |
13.916 | 15.279 | 49.975 | 393.976 | 438.239 | 2,530 |
valkey-9.1 |
13.836 | 14.898 | 45.446 | 330.686 | 363.939 | 3,003 |
dragonfly |
12.013 | 13.806 | 40.346 | 247.789 | 4,022.980 | 2,096 |
mini-redis-go-55178df |
n/a | n/a | n/a | n/a | n/a | n/a |
Unmatched Publication
With no delivery to wait for, this table reports the PUBLISH acknowledgement latency. It directly exposes the cost of scanning patterns and finding no match.
| Engine | 0 patterns | 64 patterns | 1,024 patterns | 8,192 patterns | 8,192 p99 |
|---|---|---|---|---|---|
goblin-sbe-ring-4kb |
0.180 | 0.561 | 5.871 | 43.011 | 48.712 |
goblin-resp-uds |
7.584 | 10.550 | 14.808 | 58.541 | 64.492 |
redis-7.2.4 |
10.640 | 10.440 | 30.838 | 282.905 | 321.017 |
redis-8.8 |
9.648 | 11.101 | 33.403 | 266.584 | 300.559 |
valkey-9.1 |
10.490 | 11.943 | 32.181 | 241.347 | 271.273 |
dragonfly |
12.484 | 13.435 | 29.716 | 153.921 | 173.879 |
mini-redis-go-55178df |
10.249 | n/a | n/a | n/a | n/a |
Pattern cost is still linear; Goblin has a substantially lower constant factor. At 8,192 patterns, Goblin RESP2/UDS is 3.95x faster than the best incumbent hit median and 2.63x faster on a miss. The ring extends the hit lead to 5.69x.
Many Clients and Subscriptions
Thirty-two subscribers each hold 32 private subscriptions plus one shared subscription. The publication targets the shared literal channel or shared pattern, producing 32 validated deliveries.
| Engine | literal p50 | literal p99 | literal pub/s | pattern p50 | pattern p99 | pattern pub/s |
|---|---|---|---|---|---|---|
goblin-sbe-ring-4kb |
6.623 | 7.344 | 145,663 | 12.283 | 16.151 | 78,749 |
goblin-resp-uds |
90.551 | 97.765 | 11,016 | 98.266 | 108.014 | 10,097 |
redis-7.2.4 |
94.098 | 100.861 | 10,631 | 144.724 | 151.797 | 6,911 |
redis-8.8 |
95.140 | 101.652 | 10,509 | 151.066 | 158.580 | 6,605 |
valkey-9.1 |
97.835 | 104.067 | 10,170 | 147.539 | 155.204 | 6,718 |
dragonfly |
85.973 | 3,024.900 | 5,838 | 126.850 | 3,725.870 | 3,995 |
mini-redis-go-55178df |
116.781 | 149.082 | 8,338 | n/a | n/a | n/a |
SBE/ring completes the literal topology 13.0x faster at p50 and the pattern topology 10.3x faster than the best incumbent. Goblin RESP2/UDS is 5.3% behind Dragonfly's literal p50 while completing 1.89x as many closed-loop publications; on patterns, Goblin RESP2/UDS leads the best incumbent p50 by 1.29x.
Server Memory
Each cell is total RSS / RSS delta in MiB after constructing the topology. The delta subtracts that fresh server's empty baseline. Total RSS is the fair view of the complete process; the delta helps isolate subscription growth. Goblin ring mappings exist before the baseline, while client output buffers are activated by subscription, so neither number should be read without the other. Small RSS deltas are also subject to page granularity and allocator reuse.
| Engine | 16,384 literals | 8,192 patterns | 32x32 literals | 32x32 patterns |
|---|---|---|---|---|
goblin-sbe-ring-4kb |
9.60 / 4.41 | 7.96 / 2.77 | 6.64 / 0.55 | 6.73 / 0.64 |
goblin-resp-uds |
9.56 / 4.43 | 7.93 / 2.79 | 6.01 / 0.90 | 6.17 / 1.04 |
redis-7.2.4 |
10.66 / 3.86 | 8.89 / 2.11 | 7.29 / 0.52 | 7.23 / 0.52 |
redis-8.8 |
11.95 / 4.08 | 10.05 / 2.19 | 8.29 / 0.41 | 8.28 / 0.41 |
valkey-9.1 |
12.44 / 4.83 | 10.26 / 2.71 | 7.96 / 0.41 | 7.92 / 0.43 |
dragonfly |
27.67 / 5.32 | 25.33 / 3.03 | 24.55 / 2.19 | 24.39 / 1.95 |
mini-redis-go-55178df |
17.59 / 6.66 | n/a | 12.66 / 1.74 | n/a |
Goblin has the lowest total server RSS in all four topologies. Redis and Valkey sometimes allocate less incremental memory after their larger baselines, which is why both sides of the slash are retained.
What This Benchmark Does Not Claim
- It is a synchronous latency and complete-delivery-rate test, not an asynchronous maximum-ingress test. A publisher allowed to build a queue would measure buffering policy and backpressure as much as Pub/Sub execution.
- SBE/ring is a Goblin-specific low-latency capability. Goblin RESP2/UDS is the apples-to-apples compatibility row for the incumbent comparison.
- The 4 KiB ring deliberately trades maximum message size for a tiny mapped footprint. Larger rings can carry the 4,096-byte case; this report holds the requested 4 KiB capacity fixed.
- Pattern matching scans every registered pattern. The tables show its linear scaling rather than implying constant-time glob routing.
- These are single-host Unix-domain-socket and shared-memory results. TCP, cross-host networking, subscriber application work, and client deserialization beyond the native harness are outside this measurement.
Reproducing
Build the server and native driver, then run the matrix:
cmake -S . -B build-rel -DCMAKE_BUILD_TYPE=Release \
-DGOBLIN_CORE_ARCH=avx2 -DGOBLIN_CORE_BUILD_BENCHMARKS=ON
cmake --build build-rel -j --target \
goblin_core_server goblin_core_pubsub_benchmark
bash benchmarks/pubsub_benchmark.sh
The launcher paths are environment-overridable; no machine-specific binary paths are required by the benchmark source. It writes the complete 26-column CSV, including minimum, p50, p90, p99, p99.9, mean, publication rate, baseline RSS, subscribed RSS, and RSS delta.
The command behavior is documented in the Pub/Sub reference. The low-latency transport and wire format are documented in ring buffers and the SBE protocol.
Tested Servers
goblin-sbe-ring-4kbgoblin-resp-udsredis-7.2.4redis-8.8valkey-9.1dragonflymini-redis-go-55178df