BLASTA

Grafana and Prometheus load testing template

Dashboards API, health and PromQL queries.

Category OperationsStack Grafana, PrometheusJobs 19

About this Grafana and Prometheus load test

Load tests Grafana and Prometheus: the Grafana health API and login page, Prometheus health, an instant query and a one-hour range query. It shows how dashboards and queries behave when many people open them at once.

It holds 19 ready-made jobs: 5 single scenarios and an enterprise test plan of 14 stages to run in order, 12 of them with pass/fail targets (SLOs). Each job is a plain request pattern you can change before running.

Scenarios include grafana /api/health, grafana login page, prometheus /-/healthy, prometheus instant query and prometheus range query (1h).

How to load test Grafana and Prometheus

  1. Open the template in BLASTA.
  2. Set url to point at your own Grafana and Prometheus system, ideally a staging copy.
  3. Pick a job and choose the rate and duration.
  4. Start the run and watch requests per second, latency percentiles and errors live; the result is saved to your history.

What you set before running

url
Base URL of the service, no trailing slash
promql
A PromQL expression

Test scenarios (5)

prometheus range query (1h)

Read-only

What a dashboard panel does. Heavy: one dashboard with 30 panels is 30 of these at once. Keep the rate low.

Enterprise test plan (14)

Run in order: smoke, baseline, load, stress, spike, soak, breakpoint and failover window, each with pass/fail targets.

Grafana and Prometheus: 01 smoke

Read-only

Enterprise plan, step 1 of 10. One request a second for 30 seconds. Run this first, every time: it proves the address, credentials and headers are right and that the environment is up before any real load is applied. Gate: zero errors. Reference request: prometheus instant query.

Grafana and Prometheus: 02 baseline (20% load)

Read-only

Step 2 of 10. About a fifth of normal traffic for 5 minutes: the uncontended latency of this request. Every later result is judged against it, so record p50 and p95. Gate: at most 0.5% errors and the default latency targets. Reference request: prometheus instant query.

Grafana and Prometheus: 03 average load (SLO check)

Read-only

Step 3 of 10. Normal busy-hour traffic for 10 minutes. The rate is the reference job's rate: raise it to your measured production peak-hour rate. This is the run that proves (or breaks) your SLO. Gate: at most 1% errors, p95 and p99 inside the targets. Reference request: prometheus instant query.

Grafana and Prometheus: 04 peak load (2x average)

Read-only

Step 4 of 10. Twice the average for 10 minutes: the busiest hour of the year plus headroom. Latency may rise, but must stay in SLO; if it does not, you have no headroom. Gate: at most 2% errors, latency targets doubled. Reference request: prometheus instant query.

Grafana and Prometheus: 05 stress (ramp to 4x)

Read-only

Step 5 of 10. Ramps to four times average over 10 minutes, then holds for 2. Finds where it degrades and HOW: gracefully (latency rises, errors stay low) or badly (errors, timeouts, crashes, restarts). Observation only, no gate. Reference request: prometheus instant query.

Grafana and Prometheus: 06 spike (10x in 10 s)

Read-only

Step 6 of 10. Reaches ten times average within 10 seconds and holds for 2 minutes: a campaign email, a news link, a failover. Checks autoscaling, queue limits and load shedding. Gate: at most 5% errors, because shedding load is acceptable and crashing is not. Reference request: prometheus instant query.

Grafana and Prometheus: 07 recovery after the spike

Read-only

Step 7 of 10. Run IMMEDIATELY after the spike, at average load for 5 minutes. Latency and errors must return to the baseline from step 2. If they do not, something is stuck: queues, connection pools, GC, an autoscaler cool-down. Gate: same as average load. Reference request: prometheus instant query.

Grafana and Prometheus: 08 soak (1 hour at 60%)

Read-only

Step 8 of 10. One hour of steady load. Finds leaks and slow decay in memory, connections, file descriptors, disk, log volume and cache churn. Watch the resource graphs: any line that climbs and never flattens is a finding. Gate: at most 0.5% errors. Reference request: prometheus instant query.

Grafana and Prometheus: 09 breakpoint (find the ceiling)

Read-only

Step 9 of 10. Ramps to twenty times average over 20 minutes. Stop it when errors pass about 5%: the rate at that moment is your ceiling, and ceiling divided by peak is your capacity margin. Use a production-like environment, never production. Reference request: prometheus instant query.

Grafana and Prometheus: 10 resilience window (failover / deploy)

Read-only

Step 10 of 10. Average load for 15 minutes. About 5 minutes in, cause the event you are testing: kill a pod or node, fail over the database, roll out a new version, drain a zone. Errors in the window are your real availability loss. Gate: at most 1% errors overall; read the time series for how long the dip lasted. Reference request: prometheus instant query.

Grafana and Prometheus: no keep-alive (new connection per request)

Read-only

Sends Connection: close, so every request pays for a new TCP and TLS handshake: the cost for clients that do not reuse connections (scripts, some mobile SDKs, health checkers). Shows load balancer and TLS termination limits. Reference request: prometheus instant query.

Grafana and Prometheus: large cookies (~4 KB header)

Read-only

Adds a 4 KB Cookie header, like a logged-in user with many tracking cookies. Proxies and servers reject headers around 8 KB, so this shows how close you are to that limit. Reference request: prometheus instant query.

Frequently asked questions

What does the Grafana and Prometheus load test cover?

The Grafana and Prometheus template has 19 jobs: 5 single scenarios and an enterprise test plan of 14 stages (smoke, baseline, load, stress, spike, soak, breakpoint and failover window). Scenarios include grafana /api/health, grafana login page, prometheus /-/healthy and prometheus instant query. 12 of them have pass/fail targets (SLOs), so a run can be judged against limits you set.

How do I load test Grafana and Prometheus with BLASTA?

Open the template in BLASTA and set url, then pick a job and start it. Results stream live: requests per second, latency percentiles and errors, and the run is kept in your history. To run from the command line, use blasta preset new observability with your address.

Is it safe to run the Grafana and Prometheus load test against production?

All 19 jobs in this template are read-only: they request pages or data and do not change anything. Even so, a load test can slow a live system down, so start with a low rate and prefer a staging copy. Only test systems you own or have permission to test.

Run a load test

Point BLASTA at something you own, choose how hard to hit it, and press Start test. Results stream in live.

1 What do you want to test? Use a template

Request headers

2 How hard should it hit?

Advanced limits
Pass / fail targets (SLO) optional
The result is marked SLO met or SLO missed. The same targets live in a job file, where blasta run exits 2 on a miss so CI or Kubernetes can gate a release.

Set up this job

Quick check

Please confirm you are not a robot to start your free test.

Clear history

Delete every finished test in your history. Tests still running are kept. This cannot be undone.

Add identity provider

Add user

The account is active at once. Share the password with them securely; they can change it from their menu.

Sign in to use templates

Templates and test history are part of the full app. Sign in, or create a free account, to use them.

Sign inCreate account

Change password