Removing a node from a cluster

Warning

This article is dedicated to YDB clusters that use configuration V2. This configuration method is currently experimental and is only available for YDB versions starting from v25.1. For production use, we recommend choosing configuration V1 — it is the main method and is officially supported for all YDB clusters.

This article describes how to remove a dynamic or static node from a YDB cluster deployed manually on virtual machines or physical servers. Removing nodes from a Kubernetes deployment is outside the scope of this procedure.

Removing a dynamic node

Removing a dynamic node does not require changing the cluster configuration.

To remove a dynamic node without affecting query processing:

  1. Drain the tablets from the node and wait for the operation to complete.
  2. Stop the YDB process on the node. Check first that the process can be stopped safely, then stop it.

After stopping the process, check the Nodes tab on the cluster monitoring page and make sure that the removed node is no longer shown as active. A dynamic node may disappear from the list immediately or remain temporarily with the Disconnected status until its node ID lease in NodeBroker expires. If the node becomes active again, make sure that the process was stopped on the correct host and is not configured to restart automatically, for example, by a systemd unit or a supervisor.

Removing a static node

Static nodes serve the storage system and are listed in the hosts section. A static node can contain VDisks of dynamic and static groups, as well as State Storage, Board, and SchemeBoard replicas. These resources must be moved before the node is removed from the configuration.

Before starting the procedure, use the YDB UI to check that the affected storage groups are healthy, that is, all VDisks of these groups are shown in the Ok state (highlighted in green), with none in Error or Degraded state.

The remaining nodes must have enough free PDisk space and slots for all VDisks from the node being removed. VDisk placement across failure domains and failure realms must comply with the configured erasure coding scheme to preserve group fault tolerance after node removal. For details on calculating the required capacity margin, see Estimating Required Equipment.

SelfHeal is enabled for dynamic groups by default. Before removing the node, make sure it is also enabled for any resources hosted on this node:

To remove a static node:

  1. If tablets are running on the node, drain them.

  2. Check that the process can be stopped safely, then stop it.

  3. Wait for SelfHeal to move the VDisks from the node. With the default settings, relocation starts approximately one hour after the node is stopped. To start relocation immediately, first obtain the IDs of all PDisks on the node being removed using YDB DSTool:

    ydb-dstool -e <bs_endpoint> pdisk list --columns NodeId:PDiskId FQDN Path
    

    <bs_endpoint> is the endpoint of any available storage node in the [PROTOCOL://]HOST[:PORT] format, for example, http://node1.example.com:8765.

    Select every row whose FQDN matches the host name of the node being removed. Verify the disk paths in the Path column and save all corresponding NodeId:PDiskId values.

    Then set all identified PDisks to BROKEN in one command:

    ydb-dstool -e <bs_endpoint> pdisk set --status BROKEN --unavail-as-offline --pdisk-ids "<pdisk_id_1>" ... "<pdisk_id_N>"
    

    Use the same <bs_endpoint>. Replace "<pdisk_id_1>" ... "<pdisk_id_N>" with a space-separated list of all saved disk IDs in the [NodeId:PDiskId] format, for example, "[3:1]" "[3:2]" for two disks on the node with NodeId equal to 3. The --unavail-as-offline option treats PDisks unavailable through the Whiteboard monitoring service as offline.

    The command runs in the foreground. Wait for it to complete successfully, then verify that data relocation is complete in the next step. For details, see Move VDisks from a broken/missing block store volume.

  4. In the YDB UI, check that no VDisks remain on the node and that the affected storage groups are healthy (all VDisks are in the Ok state). If State Storage, Board, or SchemeBoard replicas were moved from the node, check that the relocation is complete.

  5. Fetch the current cluster configuration using the ydb admin cluster config fetch command:

    ydb [global options...] admin cluster config fetch > config.yaml
    
  6. If the node being removed is not the last entry in the hosts list, removing it shifts the positions of all subsequent entries. To preserve their identifiers, explicitly set node_id on each subsequent entry that does not already have one, using that entry's current identifier before removal: for an entry that relied on the default, this is its current one-based position in the list. If an entry already has an explicit node_id, preserve its existing value. If the last entry is being removed, this step is not required.

    Warning

    Correct node numbering is critical for cluster health. An error in node_id assignment can move VDisks, State Storage, Board, or SchemeBoard replicas to the wrong host and may lead to irreversible data loss.

    For example, given the following configuration where the node node3 is being removed:

    hosts:
    - host: node1
    - host: node2
    - host: node3 # to be removed
    - host: node4
    - host: node5
      node_id: 50
    

    Before removing node3, explicitly set node_id: 4 for node4 so that its ID does not change to 3. Keep the existing node_id: 50 for node5. After removing node3, the list will look like this:

    hosts:
    - host: node1
    - host: node2
    - host: node4
      node_id: 4
    - host: node5
      node_id: 50
    
  7. Remove the node entry from the hosts section.

  8. Apply the configuration using the ydb admin cluster config replace command:

    ydb [global options...] admin cluster config replace -f config.yaml
    
    If the command returns an error

    If VDisks remain on the PDisks of the node being removed, the command returns an error similar to the following:

    failed to remove PDisk# 1:1 as it has active VSlots
    

    In this case, wait for SelfHeal to move the remaining VDisks. Relocation time depends on the amount of data and disk performance. Monitor the relocation on the Storage tab of the node being removed in the YDB UI. When no VDisks remain on the node, rerun the config replace command with the same file.

    If the VDisk list is not shrinking and replication is not in progress, move the remaining VDisks manually.

After the configuration is applied successfully, the server and its disks can be decommissioned.