MQTT v3.1.1 and v5 connect handshakes with the CONNACK checked, plus reconnect-storm ramps.
Category MessagingStack Any MQTT brokerJobs 16
About this MQTT broker load test
Load tests an MQTT broker (Mosquitto, EMQX, HiveMQ, VerneMQ): MQTT 3.1.1 and 5 connect handshakes with the CONNACK checked, a TCP accept-only test and a reconnect-storm ramp. It models thousands of devices reconnecting at once after an outage.
It holds 16 ready-made jobs: 4 single scenarios and an enterprise test plan of 12 stages to run in order, 8 of them with pass/fail targets (SLOs). Each job is a plain request pattern you can change before running.
Scenarios include CONNECT v3.1.1, CONNECT v5, reconnect storm ramp and TCP accept only.
How to load test MQTT broker
Open the template in BLASTA.
Set host to point at your own MQTT broker system, ideally a staging copy.
Pick a job and choose the rate and duration.
Start the run and watch requests per second, latency percentiles and errors live; the result is saved to your history.
What you set before running
host
Broker host name
port
MQTT port (1883; 8883 is TLS and is not supported here)
A real MQTT CONNECT (clean session) with the 'connection accepted' CONNACK checked: the cost of every device reconnect. Every request uses the same client id, so a broker that enforces unique ids will drop the previous connection: that is the session-takeover path, and it is expected.
MQTT 5 CONNECT with empty properties. Only the CONNACK packet type is checked: a v5 CONNACK carries a variable-length property list, so its reason code cannot be matched byte for byte. A broker without v5 support closes the connection or answers with another packet, counted as a failure.
Enterprise plan, step 1 of 10. One request a second for 30 seconds. Run this first, every time: it proves the address, credentials and headers are right and that the environment is up before any real load is applied. Gate: zero errors. Reference request: CONNECT v3.1.1.
Step 2 of 10. About a fifth of normal traffic for 5 minutes: the uncontended latency of this request. Every later result is judged against it, so record p50 and p95. Gate: at most 0.5% errors and the default latency targets. Reference request: CONNECT v3.1.1.
Step 3 of 10. Normal busy-hour traffic for 10 minutes. The rate is the reference job's rate: raise it to your measured production peak-hour rate. This is the run that proves (or breaks) your SLO. Gate: at most 1% errors, p95 and p99 inside the targets. Reference request: CONNECT v3.1.1.
Step 4 of 10. Twice the average for 10 minutes: the busiest hour of the year plus headroom. Latency may rise, but must stay in SLO; if it does not, you have no headroom. Gate: at most 2% errors, latency targets doubled. Reference request: CONNECT v3.1.1.
Step 5 of 10. Ramps to four times average over 10 minutes, then holds for 2. Finds where it degrades and HOW: gracefully (latency rises, errors stay low) or badly (errors, timeouts, crashes, restarts). Observation only, no gate. Reference request: CONNECT v3.1.1.
Step 6 of 10. Reaches ten times average within 10 seconds and holds for 2 minutes: a campaign email, a news link, a failover. Checks autoscaling, queue limits and load shedding. Gate: at most 5% errors, because shedding load is acceptable and crashing is not. Reference request: CONNECT v3.1.1.
Step 7 of 10. Run IMMEDIATELY after the spike, at average load for 5 minutes. Latency and errors must return to the baseline from step 2. If they do not, something is stuck: queues, connection pools, GC, an autoscaler cool-down. Gate: same as average load. Reference request: CONNECT v3.1.1.
Step 8 of 10. One hour of steady load. Finds leaks and slow decay in memory, connections, file descriptors, disk, log volume and cache churn. Watch the resource graphs: any line that climbs and never flattens is a finding. Gate: at most 0.5% errors. Reference request: CONNECT v3.1.1.
Step 9 of 10. Ramps to twenty times average over 20 minutes. Stop it when errors pass about 5%: the rate at that moment is your ceiling, and ceiling divided by peak is your capacity margin. Use a production-like environment, never production. Reference request: CONNECT v3.1.1.
Step 10 of 10. Average load for 15 minutes. About 5 minutes in, cause the event you are testing: kill a pod or node, fail over the database, roll out a new version, drain a zone. Errors in the window are your real availability loss. Gate: at most 1% errors overall; read the time series for how long the dip lasted. Reference request: CONNECT v3.1.1.
Ramps new connections per second to ten times normal: a fleet reconnecting after an outage, a deploy that restarts every pod, a network blip. Shows accept-queue, file descriptor and handshake limits. Reference request: CONNECT v3.1.1.
Thousands of devices reconnect after an outage, all at once. Ramps to twenty times normal over 3 minutes. Reference request: CONNECT v3.1.1.
Frequently asked questions
What does the MQTT broker load test cover?
The MQTT broker template has 16 jobs: 4 single scenarios and an enterprise test plan of 12 stages (smoke, baseline, load, stress, spike, soak, breakpoint and failover window). Scenarios include CONNECT v3.1.1, CONNECT v5, reconnect storm ramp and TCP accept only. 8 of them have pass/fail targets (SLOs), so a run can be judged against limits you set.
How do I load test MQTT broker with BLASTA?
Open the template in BLASTA and set host, then pick a job and start it. Results stream live: requests per second, latency percentiles and errors, and the run is kept in your history. To run from the command line, use blasta preset new mqtt with your address.
Is it safe to run the MQTT broker load test against production?
All 16 jobs in this template are read-only: they request pages or data and do not change anything. Even so, a load test can slow a live system down, so start with a low rate and prefer a staging copy. Only test systems you own or have permission to test.
Set up a target above and press Start test. Charts and numbers appear here as the test runs.
Load generator resources
CPU used by BLASTA % of capacity
RAM used by BLASTA
Requests per second
Latency p95 lower is better
Response time
Status codes
Errors
Pick a template
Ready-made jobs for common systems. Search by system, protocol or what you want to test
(wordpress, saml, redis, spike, login storm…), pick a job, and it opens on the Test page ready to run.
You are browsing as a visitor: you can read every template and job. Sign in or create an account to use them.
No templates match
Try a shorter search, or a system name such as keycloak, postgres or soap.