A method, apparatus, and medium for restoring a backup

By updating the configuration file and obtaining the status when the disk is unplugged, the RAID1 array is automatically reconstructed, which solves the problem of backup inaccessible recovery caused by the disk being accidentally unplugged, and improves the reliability and business stability of the distributed storage system.

CN114995760BActive Publication Date: 2025-07-22JINAN INSPUR DATA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210609486.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-31
Publication Date
2025-07-22
Estimated Expiration
2042-05-31

AI Technical Summary

Technical Problem

In RAID1 backup, the backup cannot be automatically restored after the disk is accidentally unplugged, resulting in the inability to automatically restore data after the disk failure, affecting the reliability and business stability of the distributed storage system.

Method used

Update the configuration file when the disk is unplugged, obtain the disk status, and reconstruct RAID1 according to the updated configuration file to restore the backup. By sensing the disk insertion and unplugging operations, the RAID1 array is automatically adjusted to achieve data recovery.

Benefits of technology

Automatically restore backup after disk failure recovery, reducing the impact of failure on the overall stability of the cluster, improving the reliability of distributed storage systems and business stability in failure scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114995760B_ABST
    Figure CN114995760B_ABST
Patent Text Reader

Abstract

The present application discloses a method, apparatus and medium for restoring a backup, relating to the technical field of distributed storage. The method includes: updating a configuration file when a first disk is removed; wherein the configuration file is a file recording information of two disks constituting a RAID1, and the disks include a first disk and a second disk; obtaining the status of the second disk when the first disk is inserted; reconstructing the RAID1 according to the status of the second disk and the updated configuration file to restore the backup. The removal of the disk can be regarded as a disk failure, and the insertion of the disk can be regarded as a disk recovery. After the disk failure is recovered, in this method, the RAID1 is reconstructed according to the status of the disk and the updated configuration file, the data backup is restored, the impact of the failure on the overall stability of the cluster is reduced, and the reliability of the distributed storage system and the service stability in the failure scenario are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of distributed storage technology, and particularly to a method, device, and medium for restoring backups. Background Art

[0002] After introducing the new generation storage engine Bluestore, the data index information of the objects recorded on the data storage device (Object Storage Device, OSD) is stored in the rocksdb KV database. To improve storage performance (frequent database access, if the access speed is slow, it will affect the entire storage performance), the database is generally stored on a solid state drive (SSD) partition.

[0003] Generally, an SSD disk is divided into 6 to 12 SSD partitions as the db data partitions of the OSD. When an SSD disk fails, all the data of the db databases of the corresponding OSDs will be lost, causing these OSDs to be unable to continue working online. To enable the OSDs to continue working, currently, the redundant array of independent disks (RAID) 1 is usually used to implement the backup function of the OSD database data. In this way, when one of the SSD disks fails, all kinds of data required for the OSD to run online can be read from another disk. When using RAID1 to back up the OSD database data, the disk may be accidentally unplugged, resulting in the situation that when the disk is inserted again, since RAID1 will not automatically add the disk to the RAID1 array for database data recovery, the backup cannot be restored.

[0004] It can be seen that how to automatically restore the backup after the disk failure is a technical problem that needs to be solved urgently by those skilled in the art. Summary of the Invention

[0005] The purpose of this application is to provide a method, device, and medium for restoring backups, which are used to automatically restore the backup after the disk failure.

[0006] To solve the above technical problem, this application provides a method for restoring backups, including:

[0007] When the first disk is unplugged, update the configuration file; wherein, the configuration file is a file recording the information of the two disks constituting RAID1, and the disks include the first disk and the second disk;

[0008] When the first disk is inserted, obtain the status of the second disk;

[0009] Reconstruct the RAID1 according to the status of the second disk and the updated configuration file to restore the backup.

[0010] Preferably, when the first disk is removed, updating the configuration file includes:

[0011] When the first disk is removed, obtain the status of the second disk;

[0012] When the status of the second disk is the first status, write the information of the first disk to the first preset position in the configuration file, and write the information of the second disk to the second preset position in the configuration file;

[0013] When the status of the second disk is the second status, write the information of the second disk to the first preset position in the configuration file, and write the information of the first disk to the second preset position in the configuration file.

[0014] Preferably, reconstructing the RAID1 according to the status of the second disk and the updated configuration file to restore the backup includes:

[0015] When the status of the second disk is the first status, add the first disk to the RAID1;

[0016] When the status of the second disk is the second status, determine whether the information of the first disk is included at the second preset position of the updated configuration file;

[0017] If so, reconstruct the RAID1 using the first disk.

[0018] Preferably, when the information of the first disk is not included at the second preset position of the updated configuration file, reconstructing the RAID1 according to the status of the second disk and the updated configuration file to restore the backup further includes:

[0019] Insert the second disk;

[0020] Determine whether the information of the second disk is included at the second preset position of the updated configuration file;

[0021] If not, reconstruct the RAID1 using the first disk and the second disk.

[0022] Preferably, constructing the RAID1 with two disks includes:

[0023] Divide the disks into partitions of the same size; where the number of partitions is at least 2;

[0024] Form multiple RAID1s from the partitions.

[0025] Preferably, when it is detected that the first disk and / or the second disk is pulled out, it further includes: outputting a prompt message for characterizing the failure of the RAID1.

[0026] Preferably, when it is detected that the RAID1 is being reconstructed, it further includes:

[0027] Collect the total amount of data being reconstructed, the current reconstruction progress, and the current reconstruction speed;

[0028] Display the total amount of data being reconstructed, the current reconstruction progress, and the current reconstruction speed.

[0029] To solve the above technical problems, the present application further provides a device for restoring a backup, including:

[0030] An update module, configured to update a configuration file when the first disk is pulled out; wherein, the configuration file is a file recording information of two disks constituting the RAID1, and the disks include the first disk and the second disk;

[0031] An acquisition module, configured to acquire the status of the second disk when the first disk is inserted;

[0032] A reconstruction module, configured to reconstruct the RAID1 according to the status of the second disk and the updated configuration file to restore the backup.

[0033] To solve the above technical problems, the present application further provides a device for restoring a backup, including:

[0034] A memory, configured to store a computer program;

[0035] A processor, configured to implement the steps of the above method for restoring a backup when executing the computer program.

[0036] To solve the above technical problems, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above method for restoring a backup are implemented.

[0037] A method for restoring a backup provided by this application includes: updating a configuration file when a first disk is removed; wherein, the configuration file is a file recording information of two disks constituting RAID1, and the disks include the first disk and the second disk; when the first disk is inserted, obtaining the status of the second disk; reconstructing RAID1 according to the status of the second disk and the updated configuration file to restore the backup. Disk removal can be regarded as a disk failure, and disk insertion can be regarded as a disk recovery. After the disk failure is recovered, in this method, RAID1 is reconstructed based on the status of the disk and the updated configuration file, the data backup is restored, the impact of the failure on the overall stability of the cluster is reduced, and the reliability of the distributed storage system and the service stability in the failure scenario are improved.

[0038] In addition, this application also provides a device for restoring a backup and a computer-readable storage medium, which have the same or corresponding technical features as the method for restoring a backup mentioned above, and the effects are the same. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] To more clearly illustrate the embodiments of this application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0040] Figure 1 It is a flowchart of a method for restoring a backup provided by this application;

[0041] Figure 2 It is a structural diagram of a device for restoring a backup provided by an embodiment of this application;

[0042] Figure 3 It is a structural diagram of a device for restoring a backup provided by another embodiment of this application;

[0043] Figure 4 It is an overall architecture diagram of OSD metadata backup of a distributed storage system provided by an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0044] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the drawings in the embodiments of this application. Obviously, the described embodiments are only some embodiments of this application, rather than all embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of this application.

[0045] The core of this application is to provide a method, device and medium for restoring a backup, which are used to automatically restore the backup after the disk failure is recovered.

[0046] For ease of understanding, the following introduces the hardware structure used in the technical solution of this application. After introducing the new generation storage engine Bluestore, the data index information of the objects recorded on the OSD is stored in the rocksdb KV database. To improve storage performance (the database is frequently accessed, and if the access speed is slow, it will affect the entire storage performance), the database is generally stored on the SSD partition. In addition, some metadata information of the cluster on the OSD (osdmap, pglog, pginfo, superblock) is also stored in the rocksdb database in the form of kv to improve access performance.

[0047] Generally, a single SSD disk will be divided into 6 to 12 SSD partitions as the db data partitions of the OSD. When a single SSD disk fails, all the database information of the corresponding OSDs will be lost, causing these OSDs to be unable to continue working online. At this time, after the SSD is replaced, the databases of the OSDs are all lost, and the data on the corresponding OSDs cannot be accessed according to the index information. Therefore, it is necessary to format the corresponding hard disk drive (HDD) (the data storage disk of the OSD), delete the relevant OSDs, and then perform the operations of re-adding. After that, the data on these HDDs is restored through data reconstruction. This will cause a problem. If the water level of the storage pool is already very high and the amount of data used is very large during a failure, since multiple OSDs need to be reconstructed at the same time, the amount of data reconstruction in this storage pool will be very large and the reconstruction time will be very long. In addition, during data reconstruction, the reconstructed data will occupy a part of the performance of the OSD, so the front-end service will experience a certain performance degradation during reconstruction. On the other hand, since a large amount of data is in a scenario where members are missing during data reconstruction, if an unknown failure occurs on other disks at this time, there is a possibility that the data will be lost and cannot be recovered. Therefore, a long data reconstruction will affect the reliability of the cluster. Therefore, RAID1 is used to perform real-time mirror backup of the OSD database (backed up to another SSD disk). In this way, when one of the SSD disks fails, all the data required for the OSD to run online can be read from the other disk. However, there are also some problems with the original RAID1 mechanism. For example, when a disk is manually removed, RAID1 will automatically remove the corresponding disk from the RAID array. But after the disk is inserted, RAID1 will not add it to the RAID1 array for database data recovery; after the overall failure of RAID1, manual intervention is required to complete the reorganization of RAID1. Therefore, this application optimizes the mechanism of RAID1 to achieve automatic recovery of the backup after disk failure.

[0048] To enable those skilled in the art to better understand the solution of this application, the following further detailed description of this application will be given in conjunction with the accompanying drawings and specific embodiments. Figure 1 It is a flowchart of a method for restoring a backup provided by this application. As Figure 1 shown, the method includes:

[0049] S10: When the first disk is removed, update the configuration file; wherein, the configuration file is a file recording information of two disks constituting RAID1, and the disks include the first disk and the second disk.

[0050] When implementing redundant backup of data using RAID1, first deploy disks in the user interface (UI). Divide preset-sized partitions on two SSD disks according to the overall requirements of the cluster. The size and number of partitions are not limited and are determined according to the size of the data to be backed up. In practice, to meet both the need for backing up data and making full use of the space on the disks, preferably, divide partitions of the same size on the two SSD disks and form a RAID1 array with these two partitions (note that the uuid of the disk needs to be used to prevent the disk drive letter from drifting). Taking the backup of OSD data as an example, divide a series of partitions on RAID1 as the database partitions of OSD, and specify block.db as the partition on RAID1 during the initialization of OSD to complete the initialization of OSD. Support multiple pairs of partitions in a single node to form multiple RAID1 arrays. After the deployment is completed, write the disk information constituting RAID1 into the corresponding configuration file (raid.conf) to check the status of the RAID1 disk later. In this embodiment, RAID1 is composed of the first disk and the second disk, so the information of the first disk and the second disk is written into the corresponding configuration file. The configuration file records all the disk lists of the entire RAID1 and the disk lists of online and offline disks.

[0051] After completing the above steps, under the online data writing process, the database backup mechanism of OSD can be implemented. When data is written to the database, it will be automatically synchronously written to the respective partitions of two SSDs. However, in practice, abnormal situations may occur, such as the SSD disk being pulled out and then inserted again after being pulled out. When the disk is pulled out, RAID1 will automatically remove the corresponding disk from the RAID1 array. But after the disk is inserted, RAID1 will not add it to the RAID1 array for database data recovery. Therefore, in order to achieve automatic backup recovery after disk failure, in the embodiments of the present application, according to the status of the other disk (i.e., whether the other disk is online) and according to the configuration file to determine whether the disk data is complete, and then use it as the basis for whether to restart RAID. Therefore, after the first disk is pulled out, by modifying the udev rule, it is necessary to sense the disk pull-out operation and update the configuration file in a timely manner. Through the updated configuration file, the latest offline or online disk information can be seen.

[0052] S11: In the case of the first disk being inserted, obtain the status of the second disk.

[0053] S12: Reconstruct RAID1 according to the status of the second disk and the updated configuration file to recover the backup.

[0054] Since RAID1 will not actively add the disk to the array after the disk is inserted, after the first disk is inserted, by modifying the udev rule, sense that the disk insertion is completed for RAID1 reorganization and data synchronization. Specifically, reconstruct RAID1 according to the status of the second disk and the updated configuration file. The status of the disk includes the online status and the offline status. When the disk is in the inserted state, it means the disk is in the online state, and when the disk is in the pulled-out state, it means the disk is in the offline state. There may be multiple situations that cause the disk to be offline. For example, the disk is pulled out manually or a sudden power failure causes the disk to be in the pulled-out state. When the second disk is in the online state, the interface can be directly called to add the first disk to RAID1 to trigger data synchronization; but when the second disk is in the offline state, it is necessary to reconstruct RAID1 according to the disk information recorded in the configuration file. Since the information recorded in the disk that was most recently online in the configuration file is relatively complete, RAID1 can be reconstructed according to the information of the disk that was most recently online in the configuration file.

[0055] The method for restoring a backup provided by this embodiment includes: updating a configuration file when the first disk is removed; wherein, the configuration file is a file recording information of two disks constituting RAID1, and the disks include the first disk and the second disk; when the first disk is inserted, obtaining the status of the second disk; reconstructing RAID1 according to the status of the second disk and the updated configuration file to restore the backup. The removal of the disk can be considered as a disk failure, and the insertion of the disk can be considered as a disk recovery. The status of the disk can reflect the running status of RAID1. According to the updated configuration file, the latest online disk information can be obtained, and the data on the latest online disk is relatively complete. Therefore, in this method, RAID1 is reconstructed based on the status of the disk and the updated configuration file, the backup of the data is restored, the impact of the failure on the overall stability of the cluster is reduced, and the reliability of the distributed storage system and the service stability in the failure scenario are improved.

[0056] In implementation, in order to accurately record the information of the inserted and removed disks in the configuration file, a preferred implementation is that when the first disk is removed, updating the configuration file includes:

[0057] When the first disk is removed, obtaining the status of the second disk;

[0058] When the status of the second disk is the first status, writing the information of the first disk to the first preset position in the configuration file, and writing the information of the second disk to the second preset position in the configuration file;

[0059] When the status of the second disk is the second status, writing the information of the second disk to the first preset position in the configuration file, and writing the information of the first disk to the second preset position in the configuration file.

[0060] In implementation, the status of the second disk is obtained by inputting a command. Here, the disk status is the first status, i.e., the online status; the second status is the offline status. When the first disk is pulled out, i.e., the first disk is offline, the information of the first disk needs to be written to the first preset position in the configuration file. Since the second disk is always in the online state, the information of the second disk is written to the second preset position in the configuration file. It should be noted that the first preset position is used to record the information of the offline disk, and the second preset position is used to record the information of the online disk. For example, after the first disk is pulled out, if the second disk is in the active state in RAID1, then the raid.conf file is updated, the offline devices line is overwritten and updated, the information of the pulled-out disk, i.e., the first disk, is written, the lastactive devices line is overwritten and updated, and the information of the second disk is written. After the first disk is pulled out, if the second disk is not in the active state in RAID1, it means that the second disk has been pulled out before the first disk is pulled out. When the second disk is pulled out, since the first disk has not been pulled out, the first disk is in the active state. Then the raid.conf file is updated, the offline devices line is overwritten and updated, the information of the pulled-out disk, i.e., the second disk, is written, and the lastactivedevices line is overwritten and updated, and the information of the first disk is written.

[0061] In this embodiment, when updating the configuration file, placing the information of the online disk and the offline disk at the first preset position and the second preset position respectively can distinguish the online disk and the offline disk, facilitating obtaining the latest online disk and understanding the information of the offline disk.

[0062] In the above steps, the updated configuration file is obtained, and the information of the latest online and offline disks is obtained through the updated configuration file. RAID1 reconstruction can be achieved in combination with the status of the second disk. In implementation, the preferred implementation method is to reconstruct RAID1 according to the status of the second disk and the updated configuration file to restore the backup, including:

[0063] When the status of the second disk is the first status, add the first disk to RAID1;

[0064] When the status of the second disk is the second status, determine whether the information of the first disk is included at the second preset position of the updated configuration file;

[0065] If so, reconstruct RAID1 using the first disk.

[0066] After the first disk is inserted, obtain the status of the second disk. When the status of the second disk is the online status, it indicates that the information recorded on the second disk is relatively complete. Therefore, the interface can be directly called to add the disk to RAID1, triggering data synchronization. When the status of the second disk is the offline status, it indicates that the data on the second disk is not the latest, that is, the recorded information may be incomplete. Therefore, it is necessary to determine the disk with relatively complete data recorded in the first disk and the second disk according to the updated configuration file. The disk that was most recently online recorded in the updated configuration file is the disk with relatively complete recorded information. Therefore, when the status of the second disk is the second status, it is determined whether the information of the first disk is included in the second preset position of the updated configuration file. If it is included, it indicates that the first disk was pulled out later than the second disk. Therefore, the information recorded on the first disk is more complete than the information recorded on the second disk. Therefore, the first disk is used to reconstruct RAID1. Specifically, after the first disk is inserted, the second disk is in a non-active state in RAID1. It is judged whether the inserted disk is among them through the information in the "lastactive devices" line in the configuration file. If it is, it indicates that the db data of this disk was the latest before the previous raid1 failure. RAID1 can be restarted with this disk, and the interface is called to forcibly restart RAID1 to start the osd service to provide read and write services.

[0067] First, the status of the second disk is obtained in the embodiment provided. When the status of the second disk is the online status, the first disk is directly added to RAID1; when the status of the second disk is the offline status, RAID1 reconstruction is performed according to the information in the updated configuration file. The online status of the second disk indicates that the operation status of the entire RAID1 is normal; through the information in the updated configuration file, it is possible to ensure that the data in the disk is relatively complete as much as possible, so that the data of the reconstructed RAID1 is relatively complete.

[0068] Based on the above embodiment, when the status of the second disk is the offline status, it is possible that the information of the first disk is not included in the second preset position of the updated configuration file. To implement the reconstruction of RAID1, a preferred implementation is that reconstructing RAID1 according to the status of the second disk and the updated configuration file to restore the backup further includes:

[0069] Insert the second disk;

[0070] Judge whether the information of the second disk is included in the second preset position of the updated configuration file;

[0071] If not, reconstruct RAID1 using the first disk and the second disk.

[0072] When the second disk is in an offline state and the first disk is not included at the second preset position of the updated configuration file, it indicates that the first disk was the earlier removed disk before the RAID1 failure, and the data may not be the most complete. Therefore, no operation is performed on the first disk, and wait for the second disk to be inserted. After the second disk is inserted, determine whether the second disk's information is included at the second preset position of the updated configuration file. When the second disk's information is also not included at the second preset position of the updated configuration file, that is, neither the first disk information nor the second disk information is in the lastactive devices information, it indicates that the two disks in the RAID1 may have been removed simultaneously or an abnormal situation has occurred. However, these two disks retain all the data of this RAID1 group. The RAID1 can be restarted to force the reconstruction of the RAID1 with these two disks and start the OSD service.

[0073] In this embodiment, when the two disks are removed simultaneously or an abnormal situation occurs, after the two disks are inserted again, by forcing the reconstruction of the RAID1 with these two disks and starting the OSD service, the automatic restoration of the backup after the disk failure is realized.

[0074] In the implementation, when the partitions are placed in the same RAID1 array, when the RAID1 array fails, the backup data on the RAID1 array cannot be used. Therefore, preferably, the implementation method of constructing a RAID1 with two disks includes:

[0075] Divide the disks into partitions of the same size; where the number of partitions is at least 2;

[0076] Form multiple RAID1s with the partitions.

[0077] In this embodiment, forming multiple RAID1 arrays with the partitions enables obtaining data from the remaining RAID1 arrays when one RAID1 array fails. In addition, in the scenario of multiple RAID1 arrays, write a "#mdx" before the corresponding disk information, where x represents the mapping serial number of the RAID1, and it is convenient to identify the corresponding disk by this mark.

[0078] In order to facilitate the user to understand the status of the RAID1, in the implementation, preferably, when it is detected that the first disk and / or the second disk is removed, it further includes: outputting a prompt message indicating the failure of the RAID1.

[0079] When detecting the removal of the first disk and the insertion or removal status of the second disk, sensors or software can be used for detection. When it is detected that the first disk and / or the second disk is removed, it indicates that there is a fault in the disks of RAID1. There are no restrictions on the way of the prompt message used to characterize the RAID1 fault and the content of the prompt message, which are determined according to the actual situation. For example, the RAID1 fault light is used for prompting, and the RAID1 fault alarm is used. When the software detects that the entire RAID1 fails, the fault lights of the corresponding two disks are lit, or a fault alarm is sent to the front-end interface for display.

[0080] The output provided in this embodiment is used to characterize the prompt message of the RAID1 fault, enabling the user to timely understand the status of RAID1 through the prompt message.

[0081] To facilitate the user's understanding of the RAID1 reconstruction situation, as a preferred implementation, when it is detected that RAID1 is being reconstructed, it further includes:

[0082] Collect the total amount of data to be reconstructed, the current reconstruction progress, and the current reconstruction speed;

[0083] Display the total amount of data to be reconstructed, the current reconstruction progress, and the current reconstruction speed.

[0084] In implementation, all parameters related to the RAID1 reconstruction can be displayed. To facilitate the user to quickly obtain the relevant information of the reconstruction parameters, only the most representative reconstruction parameters are displayed. For example, in this embodiment, the total amount of data to be reconstructed, the current reconstruction progress, and the current reconstruction speed are displayed. There are no restrictions on the display method and the form of the displayed content, which are determined according to the actual situation. When the software detects that RAID1 is in the data reconstruction (synchronization) state, for example, the current reconstruction progress can be displayed in the form of a progress bar or a percentage. When performing the RAID1 reconstruction, an indicator light can also be used to represent that RAID1 is being reconstructed. It should be noted that the way of using the indicator light here needs to be distinguished from the indicator light during the above-mentioned RAID1 fault prompt, and different colors of the indicator light can be used to represent different situations of the detected RAID1.

[0085] In this embodiment, when it is detected that RAID1 is being reconstructed, the reconstruction parameters are displayed, enabling the user to always understand the RAID1 reconstruction situation; and, the representative reconstruction parameters such as the total amount of data to be reconstructed, the current reconstruction progress, and the current reconstruction speed are displayed, enabling the user to quickly understand the RAID1 reconstruction situation.

[0086] In the above embodiments, the method for restoring a backup is described in detail. The present application also provides corresponding embodiments of the apparatus for restoring a backup. It should be noted that the present application describes the embodiments of the apparatus part from two perspectives, one is from the perspective of functional modules, and the other is from the perspective of hardware.

[0087] Figure 2 The structure diagram of the apparatus for restoring a backup provided by an embodiment of the present application. This embodiment is based on the perspective of functional modules and includes:

[0088] An update module 10, configured to update a configuration file when the first disk is removed; wherein, the configuration file is a file recording information of two disks constituting RAID1, and the disks include the first disk and the second disk;

[0089] An acquisition module 11, configured to acquire the status of the second disk when the first disk is inserted;

[0090] A reconstruction module 12, configured to reconstruct RAID1 according to the status of the second disk and the updated configuration file to restore the backup.

[0091] Since the embodiments of the apparatus part correspond to the embodiments of the method part, for the embodiments of the apparatus part, please refer to the description of the embodiments of the method part, which will not be elaborated here.

[0092] For the apparatus for restoring a backup provided by this embodiment, when the first disk is removed, the update module updates the configuration file; wherein, the configuration file is a file recording information of two disks constituting RAID1, and the disks include the first disk and the second disk; when the first disk is inserted, the acquisition module acquires the status of the second disk; and the reconstruction module reconstructs RAID1 according to the status of the second disk and the updated configuration file to restore the backup. In this apparatus, removing a disk can be considered as a disk failure, and inserting a disk can be considered as a disk recovery. After the disk failure is recovered, RAID1 is reconstructed according to the status of the disk and the updated configuration file, the backup of the data is restored, the impact of the failure on the overall stability of the cluster is reduced, and the reliability of the distributed storage system and the service stability in the failure scenario are improved.

[0093] Figure 3 The structure diagram of the apparatus for restoring a backup provided by another embodiment of the present application. This embodiment is based on the hardware perspective. As Figure 3 shown, the apparatus for restoring a backup includes:

[0094] A memory 20, configured to store a computer program;

[0095] A processor 21, configured to implement the steps of the method for restoring a backup as mentioned in the above embodiments when executing the computer program.

[0096] The device for restoring backup provided in this embodiment may include, but is not limited to, a smart phone, a tablet computer, a notebook computer, a desktop computer, etc.

[0097] Among them, the processor 21 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 21 may be implemented in at least one hardware form of a digital signal processor (DSP), a field-programmable gate array (FPGA), or a programmable logic array (PLA). The processor 21 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as a central processing unit (CPU); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 21 may be integrated with a graphics processing unit (GPU), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 21 may further include an artificial intelligence (AI) processor, and the AI processor is used to process computational operations related to machine learning.

[0098] The memory 20 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 20 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In this embodiment, the memory 20 is at least used to store the following computer program 201. After the computer program is loaded and executed by the processor 21, it can implement the relevant steps of the method for restoring backup disclosed in any of the foregoing embodiments. In addition, the resources stored in the memory 20 may further include an operating system 202 and data 203, etc., and the storage method may be temporary storage or permanent storage. Among them, the operating system 202 may include Windows, Unix, Linux, etc. The data 203 may include, but is not limited to, the data involved in the method for restoring backup mentioned above.

[0099] In some embodiments, the device for restoring backup may further include a display screen 22, an input / output interface 23, a communication interface 24, a power supply 25, and a communication bus 26.

[0100] Those skilled in the art can understand, Figure 3The structure shown does not constitute a limitation on the device for restoring backups, and may include more or fewer components than those shown in the figure.

[0101] The device for restoring backups provided by an embodiment of the present application includes a memory and a processor. When the processor executes the program stored in the memory, the following method can be implemented: the method for restoring backups, with the same effect as above.

[0102] Finally, the present application also provides an embodiment corresponding to a computer-readable storage medium. A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps recorded in the above method embodiment are implemented.

[0103] It can be understood that if the method in the above embodiment is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and executes all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, removable disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0104] The computer-readable storage medium provided by the present application includes the above-mentioned method for restoring backups, with the same effect as above.

[0105] To enable those skilled in the art to better understand the technical solution of the present application, the following is a further detailed description of the above-mentioned present application in conjunction with the attached Figure 4 For a further detailed description of the above-mentioned present application, Figure 4 This is an overall architecture diagram of OSD metadata backup for a distributed storage system provided by an embodiment of the present application. As Figure 4 shown, this architecture diagram includes:

[0106] On the management software, through the UI, deploy to eject disks, load disks, and deploy and expand the disks. Send messages through the Web front end to generate a message queue. Process the messages, such as disk ejection, disk addition, creating RAID1, and OSD initialization. After the initialization is completed, write the information to the disk and the service information database, or send it to the underlying OSD through specific commands. In the OSD, the management of the RAID1 located in the kernel is achieved through udev events and management software commands, that is, the RAID1 is divided into the SSD1 partition and the SSD2 partition, and metadata is stored in the RAID1. The metadata specifically includes RocksDB, BlueFs, and BlockDevice, and these metadata are stored on BlueStore through the initialization module. In addition, the metadata can also be stored on BlueStore after passing through the exception handling module. The BlockDevice in BlueStore can also store data in the HDD disk located in the kernel.

[0107] Thus, it can be seen that when the OSD metadata is written, two copies of data will be written on two SSD disks respectively. In this way, when one SSD disk fails, the db data required for the OSD operation can be read from the backup disk, which can ensure the validity of the existing data on the HDD disk after the SSD failure, and greatly reduce the amount of reconstructed data and the reconstruction time caused by the SSD failure.

[0108] Based on the overall architecture of the OSD metadata backup in the above-mentioned distributed storage system, when the first disk is ejected, update the configuration file; where the configuration file is a file recording the information of the two disks constituting the RAID1, and the disks include the first disk and the second disk; when the first disk is inserted, obtain the status of the second disk; reconstruct the RAID1 according to the status of the second disk and the updated configuration file to restore the backup. Disk ejection can be considered as disk failure, and disk insertion can be considered as disk recovery. After the disk failure is recovered, in this method, the RAID1 is reconstructed according to the status of the disk and the updated configuration file, and the data backup is restored, reducing the impact of the failure on the overall stability of the cluster, and improving the reliability of the distributed storage system and the business stability in the failure scenario.

[0109] The method, apparatus, and medium for restoring backup provided by the present application have been introduced in detail above. The various embodiments in the specification are described in a progressive manner, and the key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and reference can be made to the description in the method part for related parts. It should be noted that for those of ordinary skill in the art in the technical field of the present application, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

[0110] It should also be noted that in this specification, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or device including the said element.

Claims

1. A method for restoring a backup, characterized in that, including: updating a configuration file when the first disk is removed; wherein, the configuration file is a file recording information of two disks constituting RAID1, and the disks include the first disk and the second disk; acquiring the status of the second disk when the first disk is inserted; reconstructing the RAID1 according to the status of the second disk and the updated configuration file to restore the backup; the updating the configuration file when the first disk is removed includes: acquiring the status of the second disk when the first disk is removed; when the status of the second disk is the first status, writing the information of the first disk to a first preset position in the configuration file, and writing the information of the second disk to a second preset position in the configuration file; wherein, the first status is the online status; when the status of the second disk is the second status, writing the information of the second disk to the first preset position in the configuration file, and writing the information of the first disk to the second preset position in the configuration file; wherein, the second status is the offline status; the reconstructing the RAID1 according to the status of the second disk and the updated configuration file to restore the backup includes: adding the first disk to the RAID1 when the status of the second disk is the first status; when the status of the second disk is the second status, determining whether the information of the first disk is included at the second preset position of the updated configuration file; if so, reconstructing the RAID1 by using the first disk.

2. The method for restoring a backup according to claim 1, characterized in that when the information of the first disk is not included at the second preset position of the updated configuration file, the reconstructing the RAID1 according to the status of the second disk and the updated configuration file to restore the backup further includes: inserting the second disk; determining whether the information of the second disk is included at the second preset position of the updated configuration file; if not, reconstructing the RAID1 by using the first disk and the second disk.

3. The method for restoring a backup according to claim 1, wherein constructing the RAID1 by using two disks includes: dividing the disks into partitions of the same size; wherein, the number of the partitions is at least 2; forming multiple RAID1s with the partitions.

4. The method for restoring a backup according to claim 3, wherein when it is detected that the first disk and / or the second disk is removed, further including: outputting a prompt message for indicating that the RAID1 fails.

5. The method for restoring a backup according to claim 3, wherein when it is detected that the RAID1 is reconstructed, further including: collecting the total amount of data reconstructed, the current reconstruction progress, and the current reconstruction speed; displaying the total amount of data reconstructed, the current reconstruction progress, and the current reconstruction speed.

6. A device for restoring a backup, characterized in that, including: an updating module, configured to update a configuration file when the first disk is removed; wherein, the configuration file is a file recording information of two disks constituting RAID1, and the disks include the first disk and the second disk; an acquiring module, configured to acquire the status of the second disk when the first disk is inserted; A reconstruction module, configured to reconstruct the RAID1 according to the status of the second disk and the updated configuration file to restore the backup; The update module is specifically configured to: acquire the status of the second disk when the first disk is removed; when the status of the second disk is the first status, write the information of the first disk to a first preset position in the configuration file, and write the information of the second disk to a second preset position in the configuration file; wherein, the first status is the online status; when the status of the second disk is the second status, write the information of the second disk to the first preset position in the configuration file, and write the information of the first disk to the second preset position in the configuration file; wherein, the second status is the offline status; The reconstruction module is specifically configured to: add the first disk to the RAID1 when the status of the second disk is the first status; when the status of the second disk is the second status, determine whether the information of the first disk is included at the second preset position of the updated configuration file; if so, reconstruct the RAID1 by using the first disk.

7. A device for restoring a backup, characterized in that, including: a memory, configured to store a computer program; a processor, configured to implement the steps of the method for restoring backup according to any one of claims 1 to 5 when executing the computer program.

8. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the method for restoring backup according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Multi-fault disk data recovery method for RAID (Redundant Arrays of Independent Disks) and system thereof

    CN106371947A

  • Method and system for recovering data of data storage equipment

    CN106933707A