A method for quickly restoring a Ceph node
By monitoring the failed nodes in the ceph distributed storage system, moving the OSD data disk to the recovery node, and synchronizing the configuration, the rapid recovery of ceph node failure is achieved, solving the problem of long data recovery time and risk of loss, and improving recovery efficiency and system reliability.
Patent Information
- Application Number
- CN202510153389.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-12
AI Technical Summary
The data recovery time caused by node failure in ceph distributed storage systems is long, there is a risk of data loss, and the recovery process is complex, affecting read and write performance.
By monitoring the failed nodes in the CEph distributed storage system in real time, move the OSD data disk of the failed node to the recovery node, check and ensure that the network and disk configuration of the recovery node are consistent with the failed node, install the CEph basic package, configure the CEph configuration file, and configure the CEph Osd block soft link, restore the configuration files in the Osd path directory, install the reconfigured CEph Osd disk to the failed node, start the OSD service, and wait for the cluster state to return to normal.
It realizes rapid recovery of ceph node failures, improves data recovery efficiency, shortens node recovery time, and reduces the risk of data loss.
Smart Images

Figure CN119621408B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of distributed storage systems, and particularly relates to a method for quickly recovering ceph nodes. Background Art
[0002] Ceph is a unified distributed storage system composed of multiple nodes; when the cluster scale becomes larger and larger, the number of nodes increases, and the probability of node failures becomes higher and higher;
[0003] When a certain node in the ceph distributed storage system fails, the following problems exist in the processing process of the failed node in the prior art:
[0004] The data recovery time of the cluster is linearly related to the amount of data carried by the failed node. The larger the amount of data, the longer the data recovery time; there is a risk of data loss when the cluster performs data recovery; the read and write performance of the cluster will be affected during the data recovery process of the cluster; the recovery process is complex and inconvenient for operation and maintenance. In the scenario where the node is abnormal due to a system disk or operating system failure, at this time, the ceph OSD disk on the node is normal, and data recovery caused by the above recovery method is unnecessary.
[0005] Therefore, how to improve the fault recovery process when a failed node appears in the prior art, achieve fast recovery of node failures in the ceph distributed storage system, and reduce the risk of data loss is a technical problem that urgently needs to be solved at present. Summary of the Invention
[0006] The purpose of the present invention is to provide a method for quickly recovering ceph nodes to quickly recover node failures in the ceph distributed storage system.
[0007] To solve the above technical problems, the technical solution adopted by the present invention is as follows:
[0008] A method for quickly recovering ceph nodes includes the following steps:
[0009] S1: Real-time collect the specified data of each node in the ceph distributed storage system, analyze the real-time collected specified data of each node, and determine whether there is a failed node among the nodes. If so, execute step S2;
[0010] S2: Move the OSD data disk of the failed node to the recovery node;
[0011] S3: Check whether the ceph osd disks are normal. If so, check the network configuration and disk configuration of the ceph osd disks of the recovery node, and determine whether the network configuration and disk configuration of the recovery node disks are the same as those of the faulty node. If so, execute step S4; if the ceph osd disks are not normal, directly replace the ceph osd disks of the faulty node in the ceph distributed storage system.
[0012] S4: Install the ceph basic packages, configure the ceph configuration file, and configure the ceph osd block soft link.
[0013] S5: Restore the configuration file in the osd path directory, and install the reconfigured ceph osd disks at the faulty node of the ceph distributed storage system.
[0014] S6: Start the osd service and wait for the ceph cluster status to return to normal.
[0015] Preferably, the network configuration of the ceph osd disks of the recovery node in step S3 includes public network configuration, cluster network configuration, messaging type configuration, firewall configuration, and MTU setting.
[0016] The disk configuration includes storage engine selection, core configuration parameters, permission settings, configuration files, and data placement policies.
[0017] Preferably, the ceph basic packages in step S4 include rpm data packages and python data packages required by the ceph cluster.
[0018] The ceph configuration file includes ceph cluster global global configuration information, monitor, mgr configuration information, osd global configuration information, and recovery node osd configuration information.
[0019] Preferably, the specific process of configuring the ceph osd block soft link is as follows:
[0020] Based on the osd corresponding to the ceph osd disk partition of the current node, establish a soft link between the osdblock file and the corresponding disk partition path, and modify the user and user group of the soft link file to the ceph service running user.
[0021] Preferably, the specific process of restoring the configuration file in the osd path directory in step S5 is as follows:
[0022] Restore the configuration files of bluefs, ceph_fsid, fsid, kv_backend, magic, mkfs_done, ready, require_osd_release, type, and whoami in the osd path directory according to the cluster and current osd information.
[0023] Preferably, the specific process of starting the osd service in step S6 is as follows:
[0024] Start the osd service of the current node through the systemctl start ceph-osd@{osd_id} command.
[0025] Preferably, while performing step S6, the ceph cluster status is monitored in real time during a specified period, and it is determined whether the ceph cluster status has returned to normal during the specified period. If not, start the standby cluster recovery mechanism:
[0026] S61: Execute ceph osd out {osd_id} on the ceph monitor node to set all osds on the faulty node to out, triggering the data recovery mechanism;
[0027] S62: Execute ceph osd crush remove {osd_name} to remove all osds on the faulty node from the crushmap, triggering the uniform distribution of data among nodes;
[0028] S63: Execute ceph osd rm {osd_id} to delete all osds on the faulty node;
[0029] S64: Install the ceph base package on the new node: including the rpm packages and python packages required by the ceph cluster;
[0030] S65: Configure the ceph configuration file on the new node, where the ceph configuration file includes the ceph cluster global configuration information, monitor configuration information, mgr configuration information, and new node osd configuration information;
[0031] S66: Execute the ceph-osd --mkfs command on the new node to initialize the osd of the current node;
[0032] S67: Execute the ceph osd crush command on the new node to add the osd of the current node to the crushmap;
[0033] S68: Execute the command "systemctl start ceph-osd@{osd_id}" on the new node to start the osd service of the current node.
[0034] S69: Wait for the ceph cluster status to be normal.
[0035] The beneficial effects of the present invention include:
[0036] The method for quickly restoring a ceph node provided by the present invention, when a faulty node is detected in the system during the node failure monitoring process, moves the data on the ceph osd disk of the faulty node to a redundant node, and removes the ceph osd disk of the faulty node from the ceph distributed storage system; if the ceph osd disk is normal, then checks the network configuration and disk configuration of the ceph osd disk of the restored node to ensure that the network configuration and disk configuration of the restored node disk are the same as those of the faulty node. Installs the ceph basic package, configures the ceph configuration file, and configures the ceph osd block soft link; restores the configuration file in the osd path directory, and installs the reconfigured ceph osd disk at the faulty node of the ceph distributed storage system; starts the osd service, and waits for the ceph cluster status to return to normal. By means of a method for quickly recovering data without triggering time-consuming data reconstruction, the data recovery efficiency is increased by more than 10 times, the node recovery time is shortened, and the risk of data loss during the recovery process is reduced.
[0037] First, when there is a faulty node in the ceph distributed storage system and the ceph osd disk of the faulty node itself is normal, it indicates that the operating system failure causes the node to be unavailable. Then, by removing the faulty node disk, it is ensured that the network configuration and disk configuration of the restored node are the same as those of the faulty node, installs the ceph basic package, configures the ceph configuration file, and configures the ceph osd block soft link. Restores the configuration file in the osd path directory, starts the osd service, and waits for the ceph cluster status to return to normal for fault recovery. This process uses a method for quickly recovering data without triggering time-consuming data reconstruction, effectively improving the data recovery efficiency, and shortening the node recovery time and reducing the risk of data loss during the recovery process by ensuring that the network configuration and disk configuration of the restored node are the same as those of the faulty node.
[0038] Second, when there is a faulty node in the ceph distributed storage system and the ceph osd disk of the faulty node itself is not normal, then directly replaces the ceph osd disk for fault recovery, avoiding the long time and the risk of data loss when the entire system recovers the cluster through the cluster data recovery mechanism.
[0039] Again, when there are faulty nodes in the Ceph distributed storage system and the disks themselves are normal, but due to various reasons, the method of removing the disks of the faulty nodes for fault recovery still cannot achieve system fault recovery, then execute the data recovery mechanism of the Ceph cluster, and use the data recovery mechanism of the Ceph cluster as the final alternative fault recovery mechanism to ensure that the entire Ceph cluster can achieve fault recovery when a fault occurs. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 It is a schematic diagram of the architecture of the Ceph distributed storage system of the present invention.
[0041] Figure 2 It is a schematic flowchart of the method for quickly recovering Ceph nodes of the present invention.
[0042] Figure 3 It is a schematic flowchart of the standby cluster recovery mechanism of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0043] The following further describes the present invention in detail with reference to the Figures 1 to 3 accompanying drawings:
[0044] Embodiment 1
[0045] Referring to the Figures 1 - 2 accompanying drawings, a method for quickly recovering Ceph nodes includes the following steps:
[0046] S1: Collect the specified data of each node in the Ceph distributed storage system in real time, analyze the specified data of each node collected in real time, and determine whether there are faulty nodes in each node. If so, execute step S2. By collecting and analyzing the specified data of each node in the Ceph distributed storage system in real time, real-time monitoring of the Ceph distributed storage system is achieved.
[0047] S2: Move the OSD data disk of the faulty node to the recovery node. By moving the OSD data disk of the faulty node to the recovery node, it is ensured that the entire Ceph distributed storage system can operate normally, effectively reducing the risk of data loss in the system. Then remove the Ceph OSD disk of the faulty node. Through the disk removal method, fault recovery of disk reconfiguration is performed separately without executing the data recovery mechanism of the Ceph cluster in the system, greatly reducing the time for system fault recovery.
[0048] S3: Check whether the ceph osd disks are normal. If so, check the network configuration and disk configuration of the ceph osd disks of the recovery node to determine whether the network configuration and disk configuration of the recovery node disks are the same as those of the faulty node. If so, execute step S4; if the ceph osd disks are not normal, directly replace the ceph osd disks of the faulty node in the ceph distributed storage system. When the disk itself is damaged, only disk replacement can be used for fault recovery.
[0049] S4: Install the ceph basic package, configure the ceph configuration file, and configure the ceph osd block soft link.
[0050] S5: Restore the configuration file in the osd path directory and install the reconfigured ceph osd disks at the faulty node of the ceph distributed storage system.
[0051] S6: Start the osd service and wait for the ceph cluster status to return to normal.
[0052] When using the above method of removing disks for fault recovery in this embodiment, it is necessary to ensure that the ceph osd disks themselves are normal, otherwise directly replace them. When implementing this solution, it is necessary to ensure that the operating system of the new node has been installed; the ceph osd disks of the faulty node need to be inserted into the new node, and the network configuration of the new node needs to be the same as that of the faulty node. When a ceph node fails, a method that can perform fast data recovery without triggering time-consuming data reconstruction can increase the data recovery efficiency by more than 10 times. The method of restoring the node by ensuring that the network configuration and disk configuration of the recovery node are the same as those of the faulty node shortens the node recovery time and reduces the risk of data loss during the recovery process.
[0053] Preferably, the network configuration of the ceph osd disks of the recovery node in step S3 includes public network configuration, cluster network configuration, messaging type configuration, firewall configuration, and MTU setting;
[0054] The disk configuration includes storage engine selection, core configuration parameters, permission settings, configuration files, and data placement policies.
[0055] Preferably, the ceph basic package in step S4 includes the rpm data packets and python data packets required by the ceph cluster;
[0056] The ceph configuration file includes the ceph cluster global global configuration information, monitor, mgr configuration information, osd global configuration information, and recovery node osd configuration information.
[0057] Embodiment 2
[0058] Based on Embodiment 1, the specific process of configuring the ceph osd block soft link is as follows:
[0059] According to the osd corresponding to the ceph osd disk partition of the current node, establish a soft link between the osdblock file and the corresponding disk partition path, and modify the user and user group to which the soft link file belongs to the ceph service running user.
[0060] In this embodiment, the specific process of restoring the configuration file in the osd path directory in step S5 is as follows:
[0061] According to the cluster and current osd information, restore the bluefs, ceph_fsid, fsid, kv_backend, magic, mkfs_done, ready, require_osd_release, type, and whoami configuration files in the osd path directory. The specific process of starting the osd service in step S6 is as follows:
[0062] Start the osd service of the current node through the systemctl start ceph-osd@{osd_id} command.
[0063] Embodiment 3
[0064] Based on Embodiment 1 or Embodiment 2, refer to Figure 3 ., while in step S6, monitor the ceph cluster status in real time during the specified period, and determine whether the ceph cluster status returns to normal during the specified period. If not, start the standby cluster recovery mechanism:
[0065] S61: Execute ceph osd out {osd_id} on the ceph monitor node to set all osds on the faulty node to out, triggering the data recovery mechanism. It can ensure the integrity and consistency of the data, and at the same time can minimize the risk of data loss. Automatically repair data corruption or loss through the recovery process to ensure data reliability.
[0066] S62: Execute ceph osd crush remove {osd_name} to remove all osds on the faulty node from the crushmap, triggering the uniform distribution of data among the nodes. It mainly ensures the uniform distribution of data among the nodes after osd removal to improve the performance and reliability of the cluster.
[0067] S63: Execute ceph osd rm {osd_id} to delete all osds on the faulty node;
[0068] S64: Install the Ceph base package on the new node: including the RPM packages and Python packages required for the Ceph cluster;
[0069] S65: Configure the Ceph configuration file on the new node, where the Ceph configuration file includes the global configuration information of the Ceph cluster, the monitor configuration information, the mgr configuration information, and the OSD configuration information of the new node;
[0070] S66: Execute the ceph-osd --mkfs command on the new node to initialize the OSD of the current node;
[0071] S67: Execute the ceph osd crush command on the new node to add the OSD of the current node to the crushmap;
[0072] S68: Execute the systemctl start ceph-osd@{osd_id} command on the new node to start the OSD service of the current node;
[0073] S69: Wait for the Ceph cluster status to be normal.
[0074] Use the above cluster recovery mechanism as the backup Ceph distributed storage system fault recovery mechanism. When the fault recovery cannot be achieved by removing the disk and reconfiguring, it can be used as the final fault recovery option to finally achieve the system fault recovery.
[0075] In summary, for the method for quickly recovering Ceph nodes provided by the present invention, when a faulty node is detected in the system during the node fault monitoring process, the data on the Ceph OSD disk of the faulty node is moved to the redundant node, and the Ceph OSD disk on the faulty node is removed from the Ceph distributed storage system; if the Ceph OSD disk is normal, then check the network configuration and disk configuration of the recovered node's Ceph OSD disk to ensure that the network configuration and disk configuration of the recovered node's disk are the same as those of the faulty node. Install the Ceph base package, configure the Ceph configuration file, and configure the Ceph osd block soft link; restore the configuration file in the osd path directory, and install the reconfigured Ceph OSD disk at the faulty node of the Ceph distributed storage system; start the OSD service and wait for the Ceph cluster status to return to normal. By using a method for quickly recovering data without triggering time-consuming data reconstruction, the data recovery efficiency is increased by more than 10 times, the node recovery time is shortened, and the risk of data loss during the recovery process is reduced.
[0076] When there are faulty nodes in the Ceph distributed storage system and the Ceph OSD disks of the faulty nodes are normal, it indicates that the operating system failure causes the nodes to be unavailable. Then, by removing the disks of the faulty nodes, ensure that the network configuration and disk configuration of the recovery nodes are the same as those of the faulty nodes, install the Ceph basic package, configure the Ceph configuration file, and configure the Ceph OSD block soft link. Restore the configuration file in the OSD path directory, start the OSD service, and wait for the Ceph cluster status to return to normal for fault recovery. This process uses a method of quickly recovering data without triggering time-consuming data reconstruction, effectively improving the data recovery efficiency, and shortening the node recovery time and reducing the risk of data loss during the recovery process by ensuring that the network configuration and disk configuration of the recovery nodes are the same as those of the faulty nodes.
[0077] When there are faulty nodes in the Ceph distributed storage system and the Ceph OSD disks of the faulty nodes are not normal, then directly replace the Ceph OSD disks for fault recovery to avoid the long time and risk of data loss when the entire system recovers the cluster through the cluster data recovery mechanism. When there are faulty nodes in the Ceph distributed storage system and the disks are normal, but due to various reasons, the method of removing the disks of the faulty nodes for fault recovery still cannot achieve system fault recovery, then execute the data recovery mechanism of the Ceph cluster, and use the data recovery mechanism of the Ceph cluster as the final alternative fault recovery mechanism to ensure that the entire Ceph cluster can achieve fault recovery when a fault occurs.
Claims
1. A method for quickly recovering a ceph node, characterized in that: The following steps are involved: S1: collect the specified data of each node in the ceph distributed storage system in real time, and analyze the specified data of each node collected in real time to determine whether there is a faulty node in each node. If so, execute step S2; S2: Move the OSD data disk of the failed node to the recovery node; S3: Check whether the ceph osd disk is normal. If so, check the network configuration and disk configuration of the ceph osd disk of the recovery node to determine whether the network configuration and disk configuration of the recovery node disk are consistent with those of the faulty node. If so, execute step S4; if the ceph osd disk is abnormal, directly replace the ceph osd disk of the faulty node in the ceph distributed storage system; S4: Install the ceph basic package, configure the ceph configuration file, and configure the ceph osd block soft link; S5: Restore the configuration files in the osd path directory and install the reconfigured ceph osd disk to the failed node of the ceph distributed storage system; S6: Start the osd service and wait for the ceph cluster status to return to normal; The specific process of configuring the ceph osd block soft link is as follows: According to the osd corresponding to the ceph osd disk partition of the current node, create a soft link between the osd block file and the corresponding disk partition path, and modify the user and user group to which the soft link file belongs to be the ceph service running user.
2. A method for quickly recovering a ceph node according to claim 1, characterized in that: The network configuration of the ceph osd disk of the restored node in step S3 includes public network configuration, cluster network configuration, message transmission type configuration, firewall configuration and MTU setting; The disk configuration includes storage engine selection, core configuration parameters, permission settings, configuration files, and data placement strategies.
3. A method for quickly recovering a ceph node according to claim 1, characterized in that: The ceph basic package in step S4 includes the rpm data package and python data package required by the ceph cluster; The ceph configuration file includes ceph cluster global configuration information, monitor, mgr configuration information, osd global configuration information and recovery node osd configuration information.
4. A method for quickly recovering a ceph node according to claim 1, characterized in that: The specific process of restoring the configuration file in the osd path directory in step S5 is as follows: According to the cluster and current osd information, restore the bluefs, ceph_fsid, fsid, kv_backend, magic, mkfs_done, ready, require_osd_release, type, and whoami configuration files in the osd path directory.
5. A method for quickly recovering a ceph node according to claim 1, characterized in that: The specific process of starting the osd service in step S6 is as follows: Start the osd service of the current node through the systemctl start ceph-osd@{osd_id} command.
6. A method for quickly recovering a ceph node according to claim 1, characterized in that: At the same time as step S6, the ceph cluster status is monitored in real time during the specified period, and it is determined whether the ceph cluster status has returned to normal during the specified period. If not, the backup cluster recovery mechanism is started: S61: Execute ceph osd out {osd_id} on the ceph monitor node to set all osds on the failed node to out, triggering the data recovery mechanism; S62: Execute ceph osd crush remove {osd_name} to remove all osds on the faulty node from the crushmap, triggering the uniform distribution of data among all nodes; S63: Execute ceph osd rm {osd_id} to delete all osds on the faulty node; S64: Install the ceph basic package on the new node: including the rpm package and python package required by the ceph cluster; S65: Configure a ceph configuration file on the new node, wherein the ceph configuration file includes ceph cluster global configuration information, monitor configuration information, mgr configuration information, and new node osd configuration information; S66: Execute the ceph-osd --mkfs command on the new node to initialize the osd of the current node; S67: Execute the ceph osd crush command on the new node to add the osd of the current node to the crushmap; S68: Execute the systemctl start ceph-osd@{osd_id} command on the new node to start the osd service of the current node; S69: Wait for the Ceph cluster status to be normal.
Citation Information
Patent Citations
Ceph-based faulted hard disk processing method and apparatus
CN107832164A
Method and system for storage virtualization
US20190042424A1