Checking whether Vault is running is easy. Knowing whether it is safe to touch the infrastructure around it is not.
/sys/health answers one question: what state is this node in? It says nothing
about how much redundancy the Raft cluster has left, whether the HA nodes agree
on who is active, how far a DR secondary has fallen behind, or whether anyone
has taken a snapshot recently. Each of those lives in a different endpoint, and
any of them can be the reason a maintenance window goes wrong.
This note collects them: what each endpoint reports, what the fields in its response mean, and a short script that collects evidence for review before updating a Kubernetes cluster, draining a node, or starting VMware maintenance.
One caveat up front: the health and seal endpoints apply to any Vault, but the Raft, Autopilot, and snapshot commands assume Integrated Storage. On a Consul or other storage backend, those sections do not apply.
Contents↗
- Why Health Checks Matter
- Prerequisites and Scope
- The /sys/health Endpoint
- Seal Status
- Raft Autopilot State
- HA Status
- DR Replication Status
- Snapshot Creation
- Putting It All Together
- Script Breakdown
- What This Does Not Cover
Why Health Checks Matter↗
A health check that only distinguishes “up” from “down” is not enough for Vault. Vault has states that are neither: sealed but running, unsealed but standby, active but disconnected from its replication primary. Each one calls for a different response, and a binary check collapses them into noise.
The second reason is routing. In an HA configuration, a load balancer needs to send traffic to the active node and keep standby nodes in rotation as candidates, not as failures. Vault is designed for this: it returns distinct status codes per state so the balancer can tell “up but not the one you want” apart from “down”.
The third is timing. Clock drift between nodes and replication lag degrade before they break. They are visible in these endpoints while they are still cheap to fix.
Vault is rarely either up or down. The useful question is which state it is in, and whether that state is the one you expected.
Prerequisites and Scope↗
To collect all the evidence and save a snapshot, the example needs:
- Vault configured with Integrated Storage;
- the Vault CLI on the operator host, plus
curlfor the status-code examples; - a Vault token with the capabilities required for Raft, Autopilot, HA status, DR status where available, and snapshot creation;
- an existing, writable
/opt/vaultdirectory on the operator host; - to be run from the root namespace, against the cluster you intend to maintain.
DR replication is Enterprise-only. A failed command remains visible in the output while the script continues collecting the other evidence.
The scope is control-plane state on the cluster you address: seal status, HA role, Raft topology, replication, and a snapshot. What it leaves out is covered at the end.
The /sys/health Endpoint↗
/sys/health is the primary endpoint for operational status. It reports
whether Vault is initialized, unsealed, active, standby, or replicating — and
it encodes the answer in the HTTP status code, so monitoring tools can act on
the response without parsing the body.
Status Codes↗
| Code | State | What it means |
|---|---|---|
| 200 | Initialized, unsealed, active | Ready to process requests allowed by authentication and policy |
| 429 | Unsealed, standby | Functional but not the active node. The odd choice of code is deliberate: it keeps standby nodes out of a balancer’s active pool without marking them unhealthy |
| 472 | DR replication secondary | A healthy state, not a failure. A DR secondary is supposed to exist and to refuse normal traffic |
| 473 | Performance standby | Enterprise. Serves read-only requests locally and forwards writes to the active node |
| 474 | Standby, no connection to the active node | The node is not a usable failover candidate, even if the cluster still serves traffic. Worth alerting on |
| 501 | Not initialized | Expected on a new instance until vault operator init runs. Later, it usually means the request reached the wrong instance |
| 503 | Sealed | No secret can be read or written until Vault is unsealed. The critical condition for most monitoring |
| 530 | Removed from the HA cluster | The node is no longer a cluster member |
Two of those are routinely misread. 429 and 472 are healthy states: a
standby node and a DR secondary are both working as designed. Monitoring that
treats them as outages pages on a correctly configured cluster.
If the defaults do not suit your load balancer, the codes are remappable
through query parameters: activecode, standbycode, sealedcode,
drsecondarycode, performancestandbycode, haunhealthycode, uninitcode,
and removedcode. The standbyok and perfstandbyok flags can also map those
healthy standby states to the active code. Setting standbycode=200, for
example, makes standby nodes indistinguishable from active ones to a balancer
that only understands 2xx.
Accessing the /sys/health Endpoint↗
The endpoint is part of Vault’s HTTP API and requires no authentication, which makes it usable from load balancers and probes that hold no token.
Because the state lives in the status code, a plain curl is not enough — it
prints only the body. Ask for the code explicitly:
# Status code only, suitable for scripting
curl -s -o /dev/null -w '%{http_code}\n' "$VAULT_ADDR/v1/sys/health"
# Code and body together
curl -i "$VAULT_ADDR/v1/sys/health"
The CLI can display fields from the same endpoint, but it does not expose the
raw HTTP status as directly and may report non-2xx states as request errors.
Use curl when the status code is the value being tested:
vault read sys/health
Example JSON response:
{
"initialized": true,
"sealed": false,
"standby": false,
"performance_standby": false,
"replication_dr_mode": "disabled",
"replication_performance_mode": "disabled",
"server_time_utc": 1730844191,
"version": "1.14.4",
"cluster_name": "vault-cluster-7cb803e8",
"cluster_id": "c792cdf2-2ff3-7ca1-0835-31667b14a9b5"
}
Example CLI response:
Key Value
--- -----
cluster_id c792cdf2-2ff3-7ca1-0835-31667b14a9b5
cluster_name vault-cluster-7cb803e8
initialized true
performance_standby false
replication_dr_mode disabled
replication_performance_mode disabled
sealed false
server_time_utc 1730844191
standby false
version 1.14.4
Note that standby and performance_standby are separate fields, and that
replication_dr_mode tells you whether the 472 code is even reachable on
this cluster.
Seal Status↗
/sys/seal-status answers the question /sys/health compresses into a single
code: if Vault is sealed, how far has the unseal progressed. Vault data remains
encrypted at rest in either state. When sealed, Vault cannot recover the key
material needed to decrypt that data, so no secret is accessible until enough
key shares have been supplied or the auto-unseal mechanism succeeds.
vault read -format=json sys/seal-status
Response example:
{
"request_id": "",
"lease_id": "",
"lease_duration": 0,
"renewable": false,
"data": {
"build_date": "2023-09-22T21:29:05Z",
"cluster_id": "c792cdf2-2ff3-7ca1-0835-31667b14a9b5",
"cluster_name": "vault-cluster-7cb803e8",
"initialized": true,
"migration": false,
"n": 1,
"nonce": "",
"progress": 0,
"recovery_seal": false,
"sealed": false,
"storage_type": "inmem",
"t": 1,
"type": "shamir",
"version": "1.14.4"
},
"warnings": null
}
| Field | Meaning |
|---|---|
sealed | true means Vault is locked and inaccessible; false means it is unsealed and serving |
t | Threshold: the number of key shares required to unseal. Set at initialization |
n | Total number of key shares issued, typically distributed across several holders |
progress | Shares entered so far in the current unseal attempt. Vault stays sealed while progress < t, and unseals when it reaches t |
nonce | Identifies the current unseal attempt. Shares from different nonces do not combine |
type | The seal mechanism: shamir for key shares, or awskms, azurekeyvault, gcpckms and others for auto-unseal |
recovery_seal | true when the instance uses auto-unseal and is reporting recovery-key state rather than unseal-key state |
migration | true while a seal migration is in progress, for example moving from Shamir to auto-unseal |
storage_type | The configured storage backend, such as raft, consul, or inmem |
version | Vault version, useful when behavior differs across releases |
cluster_name, cluster_id | Cluster identifiers, useful in multi-cluster and DR setups |
type matters operationally more than it looks. With shamir, unsealing needs
t humans and their key shares. With an auto-unseal backend, it needs the KMS
to be reachable — which makes the KMS a dependency of your secret store, and a
sealed Vault a possible symptom of a cloud IAM problem rather than a Vault one.
Raft Autopilot State↗
/sys/storage/raft/autopilot/state reports the state of the integrated storage
cluster: each node, its role, and how much redundancy remains. Autopilot
manages voter stabilization and promotion. It can also remove failed servers
when dead-server cleanup is explicitly enabled. These controls reduce manual
peer management, but they do not make quorum self-healing under every failure.
vault operator raft autopilot state -format=json
Response example:
{
"healthy": true,
"failure_tolerance": 1,
"servers": {
"raft1": {
"id": "raft1",
"name": "raft1",
"address": "127.0.0.1:8201",
"node_status": "alive",
"last_contact": "0s",
"last_term": 3,
"last_index": 459,
"healthy": true,
"stable_since": "2021-03-19T20:14:11.831678-04:00",
"status": "leader",
"meta": null
},
"raft2": {
"id": "raft2",
"name": "raft2",
"address": "127.0.0.2:8201",
"node_status": "alive",
"last_contact": "516.49595ms",
"last_term": 3,
"last_index": 459,
"healthy": true,
"stable_since": "2021-03-19T20:14:19.831931-04:00",
"status": "voter",
"meta": null
},
"raft3": {
"id": "raft3",
"name": "raft3",
"address": "127.0.0.3:8201",
"node_status": "alive",
"last_contact": "196.706591ms",
"last_term": 3,
"last_index": 459,
"healthy": true,
"stable_since": "2021-03-19T20:14:25.83565-04:00",
"status": "voter",
"meta": null
}
},
"leader": "raft1",
"voters": ["raft1", "raft2", "raft3"],
"non_voters": null
}
Top-level fields:
| Field | Meaning |
|---|---|
healthy | Whether every node is healthy and able to participate in quorum |
failure_tolerance | How many nodes can fail while the cluster still holds quorum |
leader | The name of the current Raft leader — a string, not a per-node flag |
voters | Nodes that participate in quorum decisions |
non_voters | Nodes currently replicating without voting. New nodes may appear here while stabilizing; permanent non-voters for read scaling are Enterprise |
Per-node fields inside servers:
| Field | Meaning |
|---|---|
id, name | Node identifier within the Raft cluster |
address | Cluster address of the node, on the cluster port (8201 by default) |
node_status | alive, failed, or left |
status | The node’s Raft role: leader, voter, or non-voter |
last_contact | Time since the leader last heard from this node. On the leader itself this is 0s |
last_term | Raft term the node last observed. A node lagging in term is behind on elections |
last_index | Last applied log index. A gap against the leader’s index is replication lag |
healthy | Whether autopilot considers this node healthy |
stable_since | When the node last entered its current state. Recent values on an old cluster suggest flapping |
Gate maintenance on both healthy and failure_tolerance. The first describes
whether Autopilot considers every node healthy. The second describes how many
additional healthy nodes the current topology can lose without losing quorum.
A cluster reporting healthy: true and failure_tolerance: 0 may not have lost
a node, but it has no remaining failure budget and should not lose another one
during maintenance.
healthy: truedescribes the present.failure_tolerancedescribes what the cluster can still survive.
Autopilot itself is available in all editions. Its redundancy zones and automated upgrade features are Enterprise.
HA Status↗
/sys/ha-status lists every node in the HA cluster and identifies the active
one. Where /sys/health reports the state of the node you asked, ha-status
gives you the cluster’s view: useful both for routing and for detecting a node
that unexpectedly lost its active role.
vault read -format=json sys/ha-status
Response example:
{
"Nodes": [
{
"active_node": true,
"api_address": "http://10.0.0.2:8200",
"clock_skew_ms": 0,
"cluster_address": "https://10.0.0.2:8201",
"echo_duration_ms": 0,
"hostname": "node1",
"last_echo": null,
"version": "1.17.0"
},
{
"active_node": false,
"api_address": "http://10.0.0.3:8200",
"clock_skew_ms": 0,
"cluster_address": "https://10.0.0.3:8201",
"echo_duration_ms": 20,
"hostname": "node2",
"last_echo": "2024-03-04T08:05:48.403148-05:00",
"version": "1.17.0"
},
{
"active_node": false,
"api_address": "http://10.0.0.4:8200",
"clock_skew_ms": -1,
"cluster_address": "https://10.0.0.4:8201",
"echo_duration_ms": 17,
"hostname": "node3",
"last_echo": "2024-03-04T08:05:48.657318-05:00",
"version": "1.17.0"
}
]
}
| Field | Meaning |
|---|---|
active_node | Whether this node currently holds the active role. Exactly one node should report true |
api_address | HTTP API address used by clients |
cluster_address | Address used for intra-cluster traffic: request forwarding and replication |
hostname | Node hostname, for correlating with logs and metrics |
clock_skew_ms | Difference between this node’s clock and the active node’s |
echo_duration_ms | Round-trip time of the last echo check, a latency signal between nodes |
last_echo | Timestamp of the last echo from the active node. null on the active node itself |
version | Vault version on this node, useful for spotting drift mid-upgrade |
On Enterprise, the response also carries replication_primary_canary_age_ms
(replication lag against the primary), upgrade_version (target version during
an automated upgrade), and redundancy_zone (the node’s zone in a
zone-aware HA setup).
A cluster view reporting anything other than exactly one active node is the
finding. It indicates an unavailable or inconsistent HA view that must be
investigated before maintenance. clock_skew_ms is also worth watching: drift
can affect lease timing and make health and replication observations harder to
interpret.
DR Replication Status↗
Enterprise only. Disaster recovery replication requires a Vault Enterprise license. On Community edition the read returns an error, which is itself a valid observation — the script below leaves it in the output rather than hiding it.
/sys/replication/dr/status reports the state of the replication link, the
connected secondaries, and how far they lag behind the primary. Checking it
before maintenance confirms the DR cluster can actually take over.
The endpoint answers from both sides of the relationship: on a primary it describes the secondaries, and on a secondary it describes its own sync state against the primary.
vault read -format=json sys/replication/dr/status
Response example, queried on a primary:
{
"data": {
"cluster_id": "eef2a5ab-51e2-1c05-407c-8b4dc8d09ebf",
"corrupted_merkle_tree": false,
"known_secondaries": [
"4ca6b639-046b-5bb1-8043-6788ddf09121"
],
"last_corruption_check_epoch": "-62135596800",
"last_dr_wal": 223,
"last_reindex_epoch": "0",
"last_wal": 223,
"merkle_root": "2494830f1a1c304829b5742a232d39b5457bce9a",
"mode": "primary",
"primary_cluster_addr": "",
"secondaries": [
{
"api_address": "https://127.0.0.1:65531",
"clock_skew_ms": "0",
"cluster_address": "https://127.0.0.1:65534",
"connection_status": "connected",
"last_heartbeat": "2024-03-04T10:05:56-05:00",
"last_heartbeat_duration_ms": "0",
"node_id": "4ca6b639-046b-5bb1-8043-6788ddf09121",
"replication_primary_canary_age_ms": "696"
}
],
"ssct_generation_counter": 0,
"state": "running"
}
}
| Field | Meaning |
|---|---|
mode | This cluster’s replication role: primary, secondary, or disabled |
state | State of the replication process. running is steady state; merkle-diff, merkle-sync and stream-wals indicate work in progress |
last_wal | Most recent Write-Ahead Log index generated on the primary |
last_dr_wal | Most recent WAL index replicated for DR. The gap against last_wal is the lag |
known_secondaries | IDs of the secondary clusters registered with this primary |
merkle_root | Root hash of the Merkle tree, a fingerprint of the replicated data state |
corrupted_merkle_tree | true indicates a data integrity problem, not a lag problem |
primary_cluster_addr | Address of the primary. Empty on the primary itself, populated on a secondary |
Per secondary, inside secondaries:
| Field | Meaning |
|---|---|
node_id | Identifier of the secondary cluster |
api_address, cluster_address | Addresses used to reach the secondary |
connection_status | connected or otherwise. The first thing to check on a failover drill |
last_heartbeat | Timestamp of the last heartbeat received from the secondary |
last_heartbeat_duration_ms | Round-trip time of that heartbeat |
clock_skew_ms | Clock difference between primary and this secondary |
replication_primary_canary_age_ms | Age of the last canary value observed, a direct lag measure in milliseconds |
Connectivity alone does not clear a DR setup before maintenance. Four things
have to hold: state is running, corrupted_merkle_tree is false, every
expected secondary appears as connected, and the last_wal / last_dr_wal
gap stays below a threshold chosen for the system’s recovery-point objective. A
connected secondary can still be too far behind to satisfy that objective.
Snapshot Creation↗
vault operator raft snapshot save captures a point-in-time copy of the
integrated storage: every secret, policy, mount, and piece of cluster state,
written as a compressed archive. Taken before maintenance, it is a recovery
artifact if something goes wrong. Restoring it is a separate, disruptive
procedure, not an automatic rollback.
vault operator raft snapshot save \
"/opt/vault/snapshot-$(date -u +%Y%m%dT%H%M%SZ).snap"
A UTC timestamp in YYYYMMDDTHHMMSSZ form sorts chronologically in any
listing; %m-%d-%Y does not, which becomes annoying the first time you need
the most recent snapshot in a hurry.
Snapshots are valuable for three distinct reasons:
- Disaster recovery. Stored offsite, a snapshot is a recovery path that does not depend on the cluster that produced it.
- Point-in-time recovery. A retained and verified snapshot provides a known recovery point when corruption or operator error appears later.
- Recovery testing. Restoring into an isolated recovery environment verifies that the procedure, key custody, and artifact are usable. A production snapshot must not be treated as an ordinary staging fixture.
A snapshot contains the cluster’s encrypted data and still inherits the sensitivity of Vault itself. Store it outside the cluster’s failure domain, restrict access, record a checksum, define retention, and test restoration with the required unseal or recovery-key holders.
Putting It All Together↗
This script gathers evidence for a human to analyze. It prints the available operational state and attempts to save a snapshot. If a command fails, the remaining commands still run, so one unavailable endpoint does not prevent collecting the other observations.
Replace the address and token placeholders before running it. The example disables TLS verification for local development, as its comment states.
#!/bin/bash
# Disable TLS verification for local development (Not recommended for production)
export VAULT_SKIP_VERIFY=true
# Set the Vault server address (replace 'your_vault_ip' with the actual IP address of your Vault server)
export VAULT_ADDR=https://your_vault_ip:8200
# Set your Vault token for authentication (replace 'your_token' with your actual token)
export VAULT_TOKEN=your_token
# Display the list of Raft peers to verify cluster membership and health
vault operator raft list-peers
# Print general status of the Vault server, including seal status and HA state
echo "### General status"
vault status
echo ""
# Display the Raft autopilot state to check cluster health and node roles
echo "### Raft autopilot state"
vault operator raft autopilot state
echo ""
# List all peers again to ensure accurate cluster information after checking autopilot state
echo "### List Peers"
vault operator raft list-peers
echo ""
# Display Disaster Recovery (DR) replication status to confirm synchronization with DR cluster
echo "### DR status"
vault read -format=json sys/replication/dr/status
echo ""
# Display High Availability (HA) status to check the active and standby nodes in the cluster
echo "### HA status"
vault read -format=json sys/ha-status
echo ""
# Save a Raft snapshot for backup, using a sortable UTC timestamp
echo "### Snapshot"
vault operator raft snapshot save \
"/opt/vault/backup-$(date -u +%Y%m%dT%H%M%SZ).snap"
echo ""
Script Breakdown↗
Environment. The script sets the server address, token, and TLS option at the beginning. The placeholders must be replaced for the target environment.
Command failures. There is no set -e or automatic health verdict. Review
both command output and errors. The final exit status does not summarize the
success of earlier commands.
General status. vault status reports seal state, storage type, version,
and HA information for the addressed node.
Raft peers and Autopilot. The first peer listing records membership before the other observations. The second provides another observation later in the run. Compare both with the expected topology and review Autopilot’s member health, voter count, contact times, log indexes, and failure tolerance.
DR status. The script always attempts this read. On Community edition or when the endpoint is unavailable, an error becomes part of the evidence and the remaining commands continue. On Enterprise, interpret replication role, state, connectivity, and synchronization in context.
HA status. Review the active node and recent peer contact. The endpoint lists the active node and peers it has heard from since becoming active. Compare that view with the expected membership.
Snapshot. The CLI attempts to save the file under /opt/vault on the
operator host. That directory must exist and be writable. The filename carries
a sortable UTC timestamp, so repeated runs never collide and the most recent
snapshot is the last one listed. Confirm snapshot creation from the command
output. Retention, secure storage, and restore testing belong to the backup
procedure.
What This Does Not Cover↗
Storage capacity, telemetry, and audit device health are separate concerns and are not checked here. An audit device that cannot write will block requests while every endpoint above still reports a healthy cluster.
Nor does any of this confirm that applications can authenticate. A healthy, unsealed, quorate Vault with correct replication can still be denying every request your workloads make, because auth method configuration, policies, and token lifecycles are a different layer. These observations assess Vault’s operational state; application correctness requires separate checks.
A status report is evidence. The maintenance decision still needs judgment.