health endpoint
Read-onlyLoad balancers poll this constantly. It must stay fast and must not touch the database unless that is the point of it.
Health endpoints, and verifying that rate limits and WAF rules really trigger.
Category ReliabilityStack Load balancer, API gateway, WAFJobs 18
Checks the guard rails in front of your service: a health endpoint, traffic just under a rate limit, traffic well over it, and a burst followed by silence. It proves that rate limits and WAF rules really trigger, and that your service recovers afterwards.
It holds 18 ready-made jobs: 4 single scenarios and an enterprise test plan of 14 stages to run in order, 12 of them with pass/fail targets (SLOs). Each job is a plain request pattern you can change before running.
Scenarios include health endpoint, just under the rate limit, well over the rate limit and burst then idle.
url to point at your own Health checks and rate limiting system, ideally a staging copy.urlhealthlimitedlimitLoad balancers poll this constantly. It must stay fast and must not touch the database unless that is the point of it.
Stays below {{limit}} req/s. Every request should succeed, so any 429 is counted as an error: it means the limit is stricter than documented.
Sends about 3x the limit. 429s are expected, so only 200 and 429 count as success. Errors here (5xx, timeouts) mean the limiter lets overload reach your app.
Short, hard burst: tests token-bucket limits that allow bursts and then refuse. Compare the 429 share with the over-limit job.
Run in order: smoke, baseline, load, stress, spike, soak, breakpoint and failover window, each with pass/fail targets.
Enterprise plan, step 1 of 10. One request a second for 30 seconds. Run this first, every time: it proves the address, credentials and headers are right and that the environment is up before any real load is applied. Gate: zero errors. Reference request: just under the rate limit.
Step 2 of 10. About a fifth of normal traffic for 5 minutes: the uncontended latency of this request. Every later result is judged against it, so record p50 and p95. Gate: at most 0.5% errors and the default latency targets. Reference request: just under the rate limit.
Step 3 of 10. Normal busy-hour traffic for 10 minutes. The rate is the reference job's rate: raise it to your measured production peak-hour rate. This is the run that proves (or breaks) your SLO. Gate: at most 1% errors, p95 and p99 inside the targets. Reference request: just under the rate limit.
Step 4 of 10. Twice the average for 10 minutes: the busiest hour of the year plus headroom. Latency may rise, but must stay in SLO; if it does not, you have no headroom. Gate: at most 2% errors, latency targets doubled. Reference request: just under the rate limit.
Step 5 of 10. Ramps to four times average over 10 minutes, then holds for 2. Finds where it degrades and HOW: gracefully (latency rises, errors stay low) or badly (errors, timeouts, crashes, restarts). Observation only, no gate. Reference request: just under the rate limit.
Step 6 of 10. Reaches ten times average within 10 seconds and holds for 2 minutes: a campaign email, a news link, a failover. Checks autoscaling, queue limits and load shedding. Gate: at most 5% errors, because shedding load is acceptable and crashing is not. Reference request: just under the rate limit.
Step 7 of 10. Run IMMEDIATELY after the spike, at average load for 5 minutes. Latency and errors must return to the baseline from step 2. If they do not, something is stuck: queues, connection pools, GC, an autoscaler cool-down. Gate: same as average load. Reference request: just under the rate limit.
Step 8 of 10. One hour of steady load. Finds leaks and slow decay in memory, connections, file descriptors, disk, log volume and cache churn. Watch the resource graphs: any line that climbs and never flattens is a finding. Gate: at most 0.5% errors. Reference request: just under the rate limit.
Step 9 of 10. Ramps to twenty times average over 20 minutes. Stop it when errors pass about 5%: the rate at that moment is your ceiling, and ceiling divided by peak is your capacity margin. Use a production-like environment, never production. Reference request: just under the rate limit.
Step 10 of 10. Average load for 15 minutes. About 5 minutes in, cause the event you are testing: kill a pod or node, fail over the database, roll out a new version, drain a zone. Errors in the window are your real availability loss. Gate: at most 1% errors overall; read the time series for how long the dip lasted. Reference request: just under the rate limit.
Sends Connection: close, so every request pays for a new TCP and TLS handshake: the cost for clients that do not reuse connections (scripts, some mobile SDKs, health checkers). Shows load balancer and TLS termination limits. Reference request: just under the rate limit.
Same request identified as a search crawler. Tests WAF and bot-management rules and whether crawlers get cached or origin responses. Crawlers can easily be a third of all traffic. Reference request: just under the rate limit.
Adds a 4 KB Cookie header, like a logged-in user with many tracking cookies. Proxies and servers reject headers around 8 KB, so this shows how close you are to that limit. Reference request: just under the rate limit.
A fleet of synthetic monitors checking the URL with HEAD requests every 30 seconds from 20 locations, plus load balancer probes. Cheap each, constant in total. Reference request: just under the rate limit.
The Health checks and rate limiting template has 18 jobs: 4 single scenarios and an enterprise test plan of 14 stages (smoke, baseline, load, stress, spike, soak, breakpoint and failover window). Scenarios include health endpoint, just under the rate limit, well over the rate limit and burst then idle. 12 of them have pass/fail targets (SLOs), so a run can be judged against limits you set.
Open the template in BLASTA and set url, then pick a job and start it. Results stream live: requests per second, latency percentiles and errors, and the run is kept in your history. To run from the command line, use blasta preset new health-and-limits with your address.
All 18 jobs in this template are read-only: they request pages or data and do not change anything. Even so, a load test can slow a live system down, so start with a low rate and prefer a staging copy. Only test systems you own or have permission to test.
Welcome back. Sign in to run and review load tests.
Send the confirmation email again
Just looking? Try a quick test without an account
Point BLASTA at something you own, choose how hard to hit it, and press Start test. Results stream in live.
Queries are read-only unless Allow writes is on.
Multi-statement input and writable CTEs (WITH d AS (DELETE…)) are always refused.
For MQTT, LDAP, AMQP and other binary protocols. The templates fill this in for you.
The address accepts host:port or tcp://host:port.
Leave this empty to only test the connection handshake.
No results yet
Set up a target above and press Start test. Charts and numbers appear here as the test runs.
Ready-made jobs for common systems. Search by system, protocol or what you want to test (wordpress, saml, redis, spike, login storm…), pick a job, and it opens on the Test page ready to run.
You are browsing as a visitor: you can read every template and job. Sign in or create an account to use them.
No templates match
Try a shorter search, or a system name such as keycloak, postgres or soap.
You are browsing as a visitor: you can read every job here. Sign in or create an account to use them.
Fill in your system's address. The jobs below update as you type.
Your tests. Open one to see it the way it looked live, with its charts, load settings and server usage.
| When | Job | Target | Requests | Avg req/s | p95 | Errors | CPU avg | RAM avg | Result |
|---|
Charts were not recorded for this run (it was saved by an older version).