Removing a node from a cluster
Warning
This article is dedicated to YDB clusters that use configuration V2. This configuration method is currently experimental and is only available for YDB versions starting from v25.1. For production use, we recommend choosing configuration V1 — it is the main method and is officially supported for all YDB clusters.
This article describes how to remove a dynamic or static node from a YDB cluster deployed manually on virtual machines or physical servers. Removing nodes from a Kubernetes deployment is outside the scope of this procedure.
Removing a dynamic node
Removing a dynamic node does not require changing the cluster configuration.
To remove a dynamic node without affecting query processing:
- Drain the tablets from the node and wait for the operation to complete.
- Stop the YDB process on the node. Check first that the process can be stopped safely, then stop it.
After stopping the process, check the Nodes tab on the cluster monitoring page and make sure that the removed node is no longer shown as active. A dynamic node may disappear from the list immediately or remain temporarily with the Disconnected status until its node ID lease in NodeBroker expires. If the node becomes active again, make sure that the process was stopped on the correct host and is not configured to restart automatically, for example, by a systemd unit or a supervisor.
Removing a static node
Static nodes serve the storage system and are listed in the hosts section. A static node can contain VDisks of dynamic and static groups, as well as State Storage, Board, and SchemeBoard replicas. These resources must be moved before the node is removed from the configuration.
Before starting the procedure, use the YDB UI to check that the affected storage groups are healthy, that is, all VDisks of these groups are shown in the Ok state (highlighted in green), with none in Error or Degraded state.
The remaining nodes must have enough free PDisk space and slots for all VDisks from the node being removed. VDisk placement across failure domains and failure realms must comply with the configured erasure coding scheme to preserve group fault tolerance after node removal. For details on calculating the required capacity margin, see Estimating Required Equipment.
SelfHeal is enabled for dynamic groups by default. Before removing the node, make sure it is also enabled for any resources hosted on this node:
- If the node contains a static group VDisk, enable static group SelfHeal. Alternatively, you can move the static group VDisk off the node manually, see Static Group Move.
- If the node contains State Storage, Board, or SchemeBoard replicas, enable Metadata Distribution SelfHeal. Alternatively, you can move these replicas off the node manually, see Configuring the State Storage, Board, and Scheme Board metadata distribution subsystems.
To remove a static node:
-
If tablets are running on the node, drain them.
-
Check that the process can be stopped safely, then stop it.
-
Wait for SelfHeal to move the VDisks from the node. With the default settings, relocation starts approximately one hour after the node is stopped. To start relocation immediately, first obtain the IDs of all PDisks on the node being removed using YDB DSTool:
ydb-dstool -e <bs_endpoint> pdisk list --columns NodeId:PDiskId FQDN Path<bs_endpoint>is the endpoint of any available storage node in the[PROTOCOL://]HOST[:PORT]format, for example,http://node1.example.com:8765.Select every row whose
FQDNmatches the host name of the node being removed. Verify the disk paths in thePathcolumn and save all correspondingNodeId:PDiskIdvalues.Then set all identified PDisks to
BROKENin one command:ydb-dstool -e <bs_endpoint> pdisk set --status BROKEN --unavail-as-offline --pdisk-ids "<pdisk_id_1>" ... "<pdisk_id_N>"Use the same
<bs_endpoint>. Replace"<pdisk_id_1>" ... "<pdisk_id_N>"with a space-separated list of all saved disk IDs in the[NodeId:PDiskId]format, for example,"[3:1]" "[3:2]"for two disks on the node withNodeIdequal to3. The--unavail-as-offlineoption treats PDisks unavailable through the Whiteboard monitoring service as offline.The command runs in the foreground. Wait for it to complete successfully, then verify that data relocation is complete in the next step. For details, see Move VDisks from a broken/missing block store volume.
-
In the YDB UI, check that no VDisks remain on the node and that the affected storage groups are healthy (all VDisks are in the
Okstate). If State Storage, Board, or SchemeBoard replicas were moved from the node, check that the relocation is complete. -
Fetch the current cluster configuration using the ydb admin cluster config fetch command:
ydb [global options...] admin cluster config fetch > config.yaml -
If the node being removed is not the last entry in the
hostslist, removing it shifts the positions of all subsequent entries. To preserve their identifiers, explicitly setnode_idon each subsequent entry that does not already have one, using that entry's current identifier before removal: for an entry that relied on the default, this is its current one-based position in the list. If an entry already has an explicitnode_id, preserve its existing value. If the last entry is being removed, this step is not required.Warning
Correct node numbering is critical for cluster health. An error in
node_idassignment can move VDisks, State Storage, Board, or SchemeBoard replicas to the wrong host and may lead to irreversible data loss.For example, given the following configuration where the node
node3is being removed:hosts: - host: node1 - host: node2 - host: node3 # to be removed - host: node4 - host: node5 node_id: 50Before removing
node3, explicitly setnode_id: 4fornode4so that its ID does not change to3. Keep the existingnode_id: 50fornode5. After removingnode3, the list will look like this:hosts: - host: node1 - host: node2 - host: node4 node_id: 4 - host: node5 node_id: 50 -
Remove the node entry from the
hostssection. -
Apply the configuration using the ydb admin cluster config replace command:
ydb [global options...] admin cluster config replace -f config.yamlIf the command returns an error
If VDisks remain on the PDisks of the node being removed, the command returns an error similar to the following:
failed to remove PDisk# 1:1 as it has active VSlotsIn this case, wait for SelfHeal to move the remaining VDisks. Relocation time depends on the amount of data and disk performance. Monitor the relocation on the Storage tab of the node being removed in the YDB UI. When no VDisks remain on the node, rerun the
config replacecommand with the same file.If the VDisk list is not shrinking and replication is not in progress, move the remaining VDisks manually.
After the configuration is applied successfully, the server and its disks can be decommissioned.