What it means
Redis Cluster splits keys over 16,384 hash slots, and every slot belongs to one primary node. When a node decides some slots have no working node, it reports the cluster as down. By default it then refuses commands for every key, including keys on healthy nodes, so that a cluster missing part of its data doesn’t answer as if it were whole:
(error) CLUSTERDOWN The cluster is down
There are three forms:
| Message | When |
|---|---|
CLUSTERDOWN The cluster is down | The cluster is in the fail state: a primary has failed and no replica replaced it, or slots are unassigned |
CLUSTERDOWN Hash slot not served | The key’s slot has no node at all (never assigned, or removed) |
CLUSTERDOWN The cluster is down and only accepts read commands | The cluster is down, cluster-allow-reads-when-down is on, and you sent a write |
A failed node is noticed after cluster-node-timeout (15 seconds by default): until then, commands
for its slots time out or fail to connect, and other keys still work. Once enough primaries agree
it has failed, the cluster goes to fail and answers CLUSTERDOWN. It comes back on its own as
soon as every slot is covered again.
Common causes
- A primary failed and had no replica, or its replicas were down too. With no replica to promote, its slots have no node.
- Too few primaries to agree. A failover needs a majority of primaries to vote. In a three-primary cluster, losing two at once (for example, two on the same host) leaves no majority.
- A cluster that was never fully built: nodes started with
cluster-enabled yesbut slots not assigned (redis-cli --cluster createnot run, or it stopped half-way). Every key then getsHash slot not served. - Slots lost while resharding or removing nodes: a node removed with
CLUSTER FORGETorCLUSTER DELSLOTSrun without assigning its slots elsewhere. - A network split that leaves the node you’re talking to on the minority side.
How to fix it
See what the cluster thinks
redis-cli -p <port> CLUSTER INFO
redis-cli -p <port> CLUSTER NODES
redis-cli --cluster check <host>:<port>
In CLUSTER INFO, cluster_state:fail with cluster_slots_fail above 0 means slots whose node
has failed; cluster_slots_assigned below 16384 means slots with no node. CLUSTER NODES marks
failed nodes fail (or fail? while it’s still being decided) and lists the slots each node owns.
--cluster check ends with [ERR] Not all 16384 slots are covered by nodes. when slots are
missing.
Bring the failed node back
Restart it with its data and its nodes.conf (the cluster configuration file). It rejoins with
the same ID and slots, and the cluster returns to ok. If it can’t come back, fail over to one of
its replicas:
redis-cli -h <replica-host> -p <replica-port> CLUSTER FAILOVER FORCE
FORCE doesn’t wait for the failed primary, but the other primaries must still agree. When too
few primaries are left to agree, CLUSTER FAILOVER TAKEOVER skips that vote; use it only when
you’re sure the old primary won’t come back with the same slots.
Assign missing slots
If slots have no node and hold no data you need, give them one:
redis-cli --cluster fix <host>:<port>
It lists the uncovered slots and asks before covering them with a node. Their keys are gone; restore them from a backup if you need them.
Make the next failure smaller
Give every primary at least one replica (--cluster-replicas 1 when creating the cluster), and put
primaries and their replicas on different hosts. If your application can work with part of the
data, cluster-require-full-coverage no keeps the healthy slots answering while others are down,
and cluster-allow-reads-when-down yes keeps reads (not writes) going on the nodes that are still
up. Both trade consistency for availability: decide per application.
Reproduce it
A temporary Redis 8.10.2 container running three nodes on ports 6381–6383, each started with
--cluster-enabled yes --cluster-node-timeout 2000 (2 seconds, to speed things up). Before the
slots were assigned, any key:
$ redis-cli -p 6381 SET foo 1
(error) CLUSTERDOWN Hash slot not served
CLUSTER INFO showed cluster_state:fail and cluster_slots_assigned:0. Then
redis-cli --cluster create 127.0.0.1:6381 127.0.0.1:6382 127.0.0.1:6383 --cluster-replicas 0
gave slots 0–5460 to 6381, 5461–10922 to 6382 and 10923–16383 to 6383. bar is in slot 5061 (on
6381) and foo in 12182 (on 6383). The node on 6383 was stopped with SHUTDOWN NOSAVE. One second
later GET bar on 6381 still replied "1"; five seconds later:
$ redis-cli -p 6381 GET bar
(error) CLUSTERDOWN The cluster is down
$ redis-cli -p 6381 CLUSTER INFO
cluster_state:fail
cluster_slots_assigned:16384
cluster_slots_ok:10923
cluster_slots_pfail:0
cluster_slots_fail:5461
…
$ redis-cli -p 6381 CLUSTER NODES
…
1be74a9e7ee0cef72b66b8a77a5352ac3665e16a 127.0.0.1:6383@16383 master,fail - 1791715941154 1791715940132 3 disconnected 10923-16383
…
PING still replied PONG; SET bar 2 and GET foo got the same CLUSTERDOWN. With
cluster-allow-reads-when-down yes set on the nodes before the failure:
GET bar
"3"
SET bar 4
(error) CLUSTERDOWN The cluster is down and only accepts read commands
Unassigned slots: with all three nodes up again, CLUSTER DELSLOTSRANGE 12000 12999 on each node
left slots 12000–12999 (part of 6383’s range) with no node. GET bar then got CLUSTERDOWN The cluster is down and GET foo
(slot 12182) got CLUSTERDOWN Hash slot not served. With cluster-require-full-coverage no,
GET bar worked again and foo still got Hash slot not served. redis-cli --cluster check
reported [ERR] Not all 16384 slots are covered by nodes., and --cluster fix asked:
The following uncovered slots have no keys across the cluster:
[12000-12999]
Fix these slots by covering with a random node? (type 'yes' to accept):
After yes, cluster_state was ok, cluster_slots_assigned was 16384, and SET foo 1 worked.
A temporary Valkey 8.1.10 cluster built the same way gave CLUSTERDOWN Hash slot not served before
--cluster create, CLUSTERDOWN The cluster is down after one node was stopped, and
CLUSTERDOWN The cluster is down and only accepts read commands for a write with
cluster-allow-reads-when-down yes.
In Inlet
Inlet connects to a cluster through any node or a configuration endpoint and follows the whole
cluster (Pro). Activity has a Nodes section listing the cluster’s nodes, and CLUSTER INFO and
CLUSTER NODES run in a query tab. When the cluster answers CLUSTERDOWN, Inlet shows the
message and links to this page.