Disk repair method and device, electronic equipment, storage medium and program product
By directly copying data from redundant data storage disks to replacement disks, the problem of low disk failure repair efficiency in the prior art is solved, and an efficient single data migration process is realized, reducing time and resource consumption.
Patent Information
- Application Number
- CN202510820657.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, the disk failure repair process is inefficient, requires two data copy operations, which is time-consuming and resource consumption.
By receiving the virtual group identification information of the failed disk sent by the management device and the identification information of the second disk stored in redundant data, the metadata information is directly obtained from the second disk, and the redundant data is parsed and copied to the replaced first disk, avoiding two data copying steps from the failed disk to the spare disk and then to the new disk.
It significantly reduces disk repair time and resource consumption, improves repair efficiency, and realizes efficient data recovery during a single data migration.
Smart Images

Figure CN120336090A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of distributed storage, and in particular, to a disk repair method, apparatus, electronic device, storage medium, and program product. Background Art
[0002] During the operation of a distributed storage system, when a disk storing data fails, the conventional processing procedure is to first perform an offline replacement, replacing the faulty disk with a new disk. After the disk replacement is completed, it is also necessary to completely restore the data originally stored on the disk of the object storage device (OSD) to ensure the normal operation of the distributed storage system and data integrity. In related technologies, to cope with disk failures, standby nodes and standby disks need to be prepared in advance. When a disk fails, the data of the faulty disk needs to be copied to the standby disk first; after the new disk goes online, the data needs to be copied from the standby disk to the new disk again. In this way, the entire data recovery process requires two data copy operations, which not only makes the process cumbersome but also consumes a large amount of time and resources, resulting in low disk repair efficiency. Summary of the Invention
[0003] Embodiments of the present application provide a disk repair method, apparatus, electronic device, storage medium, and program product to solve the problem of low efficiency in existing disk repair methods.
[0004] To solve the above technical problems, the present application is implemented as follows: In a first aspect, an embodiment of the present application provides a disk repair method, which is applied to a first disk. The method includes: Receiving identification information of a first virtual group on a faulty disk and identification information of a second disk corresponding to the identification information of the first virtual group sent by a management device, where the first disk is a disk used to replace the faulty disk, the faulty disk stores first data of the first virtual group, the second disk stores second data of the first virtual group, and the second data is obtained by performing data redundancy on the first data; Sending a metadata information query request to the second disk, where the metadata information query request carries the identification information of the first virtual group; Receiving the metadata information corresponding to the first virtual group sent by the second disk; Based on the metadata information corresponding to the first virtual group, parsing out data information of the second data stored on the second disk; According to the data information of the second data, copying the second data from the second disk to the first disk.
[0005] Optionally, copying the second data from the second disk to the first disk according to the data information of the second data includes: Sending a Transmission Control Protocol (TCP) connection request to the second disk and establishing a TCP connection channel with the second disk based on the TCP connection request; Copying the second data from the second disk to the first disk based on the TCP connection channel according to the data information of the second data.
[0006] Optionally, before receiving the identification information of the first virtual group on the failed disk and the identification information of the second disk corresponding to the identification information of the first virtual group sent by the management device, the method further includes: Sending the first storage node service (OSD) status information of the first disk to the management device, where the first OSD status information is used to indicate the status of replacing the failed disk with the first disk.
[0007] Optionally, after copying the second data from the second disk to the first disk according to the data information of the second data, the method further includes: When the copying of the second data from the second disk to the first disk is completed, sending a notification message for indicating the completion of the copying of the second data to the management device.
[0008] In a second aspect, an embodiment of the present application provides a disk repair method applied to a management device. The method includes: Obtaining the identification information of the first virtual group on the failed disk; Querying, from a pre-established metadata table, the identification information of the second disk corresponding to the identification information of the first virtual group. The first data of the first virtual group is stored on the failed disk, and the second data of the first virtual group is stored on the second disk, and the second data is obtained by performing data redundancy on the first data; Sending the identification information of the first virtual group and the identification information of the second disk to the first disk, where the first disk is a disk used to replace the failed disk.
[0009] Optionally, obtaining the identification information of the first virtual group on the failed disk includes: Receiving the first OSD status information of the first disk sent by the first disk, where the first OSD status information is used to indicate the status of replacing the failed disk with the first disk; Based on the first OSD status information, obtaining the identification information of the first virtual group on the failed disk.
[0010] Optionally, after sending the identification information of the first virtual group and the identification information of the second disk to the first disk, the method further includes: Receiving a notification message sent by the first disk for indicating that the second data copy is completed; Based on the notification message, modifying the first OSD status information to second OSD status information, where the first OSD status information is used to indicate the status that the first disk has completed data repair of the first virtual group.
[0011] Optionally, before obtaining the identification information of the first virtual group on the faulty disk, the method further includes: Obtaining stored data; Dividing the stored data into multiple data slices according to a preset data size; Based on the identification information of the stored data and the identification information of each data slice, determining the identification information of the virtual group corresponding to each data slice; Performing data redundancy on each data slice to obtain a data copy corresponding to each data slice; Querying, from a pre-established metadata table, the identification information of the target disk corresponding to the identification information of the target virtual group, and storing the target data copy to the target disk, where the target virtual group is the virtual group whose identification information corresponds to the target data slice, the target data slice is any one of the multiple data slices, and the target data copy is the data copy corresponding to the target data slice.
[0012] In a third aspect, an embodiment of the present application provides a disk repair method, which is applied to a second disk. The method includes: Receiving a metadata information query request sent by a first disk, where the metadata information query request carries the identification information of a first virtual group; Based on the metadata information query request, sending metadata information corresponding to the first virtual group to the first disk, where the metadata information includes data information of second data of the first virtual group; Sending the second data to the first disk according to the data information of the second data.
[0013] Optionally, the sending the second data to the first disk according to the data information of the second data includes: Receiving a TCP connection request sent by the first disk, and establishing a TCP connection channel with the first disk based on the TCP connection request; Sending the second data to the first disk based on the TCP connection channel according to the data information of the second data.
[0014] Fourthly, an embodiment of the present application further provides a disk repair device, which is applied to a first disk. The device includes: A first receiving module, configured to receive identification information of a first virtual group on a faulty disk and identification information of a second disk corresponding to the identification information of the first virtual group sent by a management device, where the first disk is a disk used to replace the faulty disk, the faulty disk stores first data of the first virtual group, the second disk stores second data of the first virtual group, and the second data is obtained by performing data redundancy on the first data; A first sending module, configured to send a metadata information query request to the second disk, where the metadata information query request carries the identification information of the first virtual group; A second receiving module, configured to receive the metadata information corresponding to the first virtual group sent by the second disk; A first parsing module, configured to parse data information of the second data stored on the second disk based on the metadata information corresponding to the first virtual group; A first copying module, configured to copy the second data from the second disk to the first disk according to the data information of the second data.
[0015] Fifthly, an embodiment of the present application further provides a disk repair device, which is applied to a management device. The device includes: A first obtaining module, configured to obtain identification information of a first virtual group on a faulty disk; A first querying module, configured to query, from a pre-established metadata table, identification information of a second disk corresponding to the identification information of the first virtual group, where the faulty disk stores first data belonging to the first virtual group, the second disk stores second data belonging to the first virtual group, and the second data is obtained by performing data redundancy on the first data; A fourth sending module, configured to send the identification information of the first virtual group and the identification information of the second disk to a first disk, where the first disk is a disk used to replace the faulty disk.
[0016] Sixthly, an embodiment of the present application further provides a disk repair device, which is applied to a second disk. The device includes: A fourth receiving module, configured to receive a metadata information query request sent by a first disk, where the metadata information query request carries identification information of a first virtual group; A fifth sending module, configured to send, based on the metadata information query request, metadata information corresponding to the first virtual group to the first disk, where the metadata information includes data information of the second data of the first virtual group; The sixth sending module is configured to send the second data to the first disk according to the data information of the second data.
[0017] In a seventh aspect, an embodiment of the present application further provides a first disk, which includes a transceiver and a processor. The transceiver is configured to: Receive the identification information of the first virtual group on the failed disk and the identification information of the second disk corresponding to the identification information of the first virtual group sent by the management device. The first disk is a disk for replacing the failed disk. The first data of the first virtual group is stored on the failed disk, and the second data of the first virtual group is stored on the second disk. The second data is obtained by performing data redundancy on the first data; Send a metadata information query request to the second disk, where the metadata information query request carries the identification information of the first virtual group; Receive the metadata information corresponding to the first virtual group sent by the second disk; The processor is configured to: Based on the metadata information corresponding to the first virtual group, parse out the data information of the second data stored on the second disk; The transceiver is configured to: According to the data information of the second data, copy the second data from the second disk to the first disk.
[0018] In an eighth aspect, an embodiment of the present application further provides a management device, which includes a transceiver and a processor. The transceiver is configured to: Obtain the identification information of the first virtual group on the failed disk; Query, from a pre-established metadata table, the identification information of the second disk corresponding to the identification information of the first virtual group. The first data belonging to the first virtual group is stored on the failed disk, and the second data belonging to the first virtual group is stored on the second disk. The second data is obtained by performing data redundancy on the first data; The transceiver is configured to: Send the identification information of the first virtual group and the identification information of the second disk to the first disk. The first disk is a disk for replacing the failed disk.
[0019] In a ninth aspect, an embodiment of the present application further provides a second disk, which includes a transceiver and a processor. The transceiver is configured to: Receive the metadata information query request sent by the first disk. The metadata information query request carries the identification information of the first virtual group; Based on the metadata information query request, metadata information corresponding to the first virtual group is sent to the first disk, wherein the metadata information includes data information of second data of the first virtual group; The second data is sent to the first disk according to data information of the second data.
[0020] In the tenth aspect, an embodiment of the present application also provides an electronic device, including a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program implements the steps of the above-mentioned disk repair method when executed by the processor.
[0021] In the eleventh aspect, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned disk repair method are implemented.
[0022] In a twelfth aspect, a computer program product is provided, comprising computer instructions, which, when executed by a processor, implement the steps of the disk repair method as described above.
[0023] In the disk repair method of the embodiment of the present application, a first disk receives identification information of a first virtual group on a faulty disk and identification information of a second disk corresponding to the identification information of the first virtual group sent by a management device, wherein the first disk is a disk used to replace the faulty disk, the faulty disk stores first data of the first virtual group, and the second disk stores second data of the first virtual group, wherein the second data is obtained by performing data redundancy on the first data; a metadata information query request is sent to the second disk, wherein the metadata information query request carries the identification information of the first virtual group; metadata information corresponding to the first virtual group is received from the second disk; based on the metadata information corresponding to the first virtual group, data information of the second data stored on the second disk is parsed; and according to the data information of the second data, the second data is copied from the second disk to the first disk.
[0024] In this embodiment, when the failed disk is replaced by the first disk, the first disk directly obtains metadata information from the second disk storing the redundant data of the failed disk, parses the second data to be restored, and then directly copies the second data. This eliminates the two data copying steps from the failed disk to the backup disk and then to the new disk in the related art, and the disk data is repaired only through one data migration process, which can significantly reduce time and resource consumption and improve the efficiency of disk repair. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments of the present application. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0026] Figure 1 is one of the flowcharts of the disk repair method provided by the embodiments of the present application; Figure 2 is the second flowchart of the disk repair method provided by the embodiments of the present application; Figure 3 is the third flowchart of the disk repair method provided by the embodiments of the present application; Figure 4 is the fourth flowchart of the disk repair method provided by the embodiments of the present application; Figure 5 is the fifth flowchart of the disk repair method provided by the embodiments of the present application; Figure 6 is the sixth flowchart of the disk repair method provided by the embodiments of the present application; Figure 7 is one of the structural diagrams of the disk repair device provided by another embodiment of the present application; Figure 8 is the second structural diagram of the disk repair device provided by another embodiment of the present application; Figure 9 is the third structural diagram of the disk repair device provided by another embodiment of the present application; Figure 10 is the structural diagram of the first disk provided by another embodiment of the present application; Figure 11 is the structural diagram of the management device provided by another embodiment of the present application; Figure 12 is the structural diagram of the second disk provided by another embodiment of the present application. Detailed implementation manners
[0027] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some, rather than all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts belong to the scope of protection of the present application.
[0028] The embodiments of the present application provide a disk repair method, which is applied to the first disk. Refer to Figure 1 , Figure 1 is the flowchart of the disk repair method provided by the embodiments of the present application, asFigure 1 As shown in the figure, it includes the following steps: Step 101: Receive the identification information of the first virtual group on the faulty disk sent by the management device and the identification information of the second disk corresponding to the identification information of the first virtual group, where the first disk is the disk used to replace the faulty disk, the first data of the first virtual group is stored on the faulty disk, the second data of the first virtual group is stored on the second disk, and the second data is obtained by data redundancy of the first data; In the distributed storage system of the present application, multiple virtual groups (Vritual Group, VG) can be defined, each virtual group has an identification information (Identity, ID), and under the multi-copy redundancy mechanism, the data of each virtual group can be redundantly obtained into multiple data copies, and then the multiple data copies corresponding to each virtual group are respectively stored on different disks. Exemplarily, under the three-copy redundancy mechanism, the data of each virtual group can be redundantly formed into three copies, and then the three data copies are respectively stored on three different disks, and each data copy corresponds to one disk. Then the management device will also store the mapping relationship between the identification information of each virtual group and the identification information of the corresponding three disks in the metadata table (VG_MAP), so that when the identification information of the virtual group is known, the disk storing its data copy can be found according to the metadata table, and when the identification information of the disk is known, all virtual groups on the disk can be found according to the metadata table.
[0029] After the faulty disk is damaged, the faulty disk can be replaced with the first disk offline. The first disk is a new disk with normal functions, and then the data repair task for the faulty disk can be carried out. First, the management device can query the identification information of all virtual groups on the faulty disk according to the metadata table, such as the identification information of the first virtual group on the faulty disk mentioned above. Then, according to the identification information of the first virtual group, using the metadata table, query the identification information of all disks storing the data copies of the first virtual group, such as the identification information of the second disk mentioned above. It should be noted that the first virtual group and the second disk are only for illustrative purposes. In addition to storing the data of the first virtual group, the faulty disk may also store the data of other virtual groups. In addition to being stored on the second disk, the data copies of the first virtual group may also be stored on other disks. The specific execution process is the same and will not be elaborated herein.
[0030] When the management device queries and obtains the identification information of the first virtual group of the faulty disk and the identification information of the second disk corresponding to the identification information of the first virtual group according to the metadata table, it sends this information to the first disk. The first data of the first virtual group was originally stored on the faulty disk. Therefore, the first disk that replaces the faulty disk needs to recover the first data. The first disk can use other disks (i.e., the aforementioned second disks) that store data copies (i.e., the aforementioned second data) of the first data to perform the data recovery process.
[0031] Step 102: Send a metadata information query request to the second disk. The metadata information query request carries the identification information of the first virtual group. In this step, metadata (BlockMeta) information records key information such as the location, size, and check value of the corresponding data (BlockData), and is the "index" for system data management. The first disk can first send a metadata information query request to the second disk. The metadata information query request carries the identification information of the first virtual group to obtain the metadata information of the first virtual group from the second disk. Then, the first disk can index the specific second data (BlockData) based on the metadata information.
[0032] Step 103: Receive the metadata information corresponding to the first virtual group sent by the second disk. In this step, after receiving the metadata information query request from the first disk, the second disk returns the metadata information of the first virtual group to the first disk according to the identification information of the first virtual group in the metadata information query request. The metadata information of the first virtual group sent by the second disk records key information such as the storage location, size, and check value of the second data, which is the basis for the first disk to copy the second data later.
[0033] Step 104: Based on the metadata information corresponding to the first virtual group, parse out the data information of the second data stored on the second disk. In this step, although the data content of the second data stored on the second disk is the same as the data content of the first data, the metadata information corresponding to the second data stored on the second disk is different from the metadata information corresponding to the first data stored on the faulty disk. For example, one piece of metadata information may record the location of the data on disk A, while another piece of metadata information may record the location of the data on disk B. If the metadata information is directly copied, it may be incompatible with the storage environment of the new disk, such as disk paths and storage formats. Therefore, the first disk needs to parse out the data information of the second data from the metadata information corresponding to the first virtual group and then copy the second data.
[0034] Step 105: Copy the second data from the second disk to the first disk according to the data information of the second data.
[0035] In this step, based on the data information of the second data parsed out, the first disk can copy the second data from the second disk through the Transmission Control Protocol (TCP) to repair the first data on the faulty disk. During the copying process, the checksum value of the second data may be verified to ensure data integrity.
[0036] In one implementation, the first disk receives the identification information of the first virtual group on the faulty disk and the identification information of the second disk corresponding to the identification information of the first virtual group sent by the management device. Here, the first disk is the disk used to replace the faulty disk. The first data of the first virtual group is stored on the faulty disk, and the second data of the first virtual group is stored on the second disk. The second data is obtained by data redundancy of the first data. The first disk sends a metadata information query request to the second disk, and the metadata information query request carries the identification information of the first virtual group. The first disk receives the metadata information corresponding to the first virtual group sent by the second disk. Based on the metadata information corresponding to the first virtual group, the data information of the second data stored on the second disk is parsed out. Then, according to the data information of the second data, the second data is copied from the second disk to the first disk.
[0037] In this implementation, after the faulty disk is replaced by the first disk, without relying on a spare disk, the first disk directly obtains the metadata information from the second disk storing the redundant data of the faulty disk, parses out the second data to be restored, and then directly copies the second data. This omits the two data copy steps from the faulty disk to the spare disk and then to the new disk in the related art, and realizes the repair of disk data only through one data migration process, which can significantly reduce time and resource consumption and improve the efficiency of disk repair.
[0038] Optionally, the step of copying the second data from the second disk to the first disk according to the data information of the second data includes: Send a Transmission Control Protocol (TCP) connection request to the second disk and establish a TCP connection channel with the second disk based on the TCP connection request; Based on the data information of the second data, copy the second data from the second disk to the first disk through the TCP connection channel.
[0039] In one embodiment, the first disk sends a TCP connection request to the second disk, and the two parties can establish a TCP connection channel through three-way handshake. The first disk parses the data information of the second data stored on the second disk according to the metadata information obtained previously. Then, through the established TCP connection channel, the first disk gradually copies the second data to the second disk.
[0040] In this embodiment, the first disk directly uses redundant data (the second disk) for single-copying, without relying on a spare disk, reducing redundant steps and improving efficiency. The TCP protocol guarantees transmission reliability, and combines with metadata to accurately locate data shards, ensuring an efficient and secure repair process.
[0041] Optionally, before sending the identification information of the first virtual group on the failed disk sent by the receiving management device and the identification information of the second disk corresponding to the identification information of the first virtual group, the method further includes: Sending the first storage node service OSD status information of the first disk to the management device, where the first OSD status information is used to indicate the status of replacing the failed disk with the first disk.
[0042] In one embodiment, after the operation and maintenance personnel replace the failed disk with the first disk, the storage node service (Object Storage Daemon, OSD) status can be updated, the OSD status is modified to the "new disk replaced status", and the OSD service corresponding to the first disk is pulled up. The OSD service of the first disk reports the first OSD status information to the management device, and the first OSD status information is used to indicate to the management device that the failed disk has been replaced with the first disk, so that the management device determines to start the data repair process of the failed disk based on the first OSD status information.
[0043] In this embodiment, by reporting the first OSD status information to the management device, the management device can perceive the completion of disk replacement in real time, avoid manual intervention, and reduce the fault recovery time.
[0044] Optionally, after copying the second data from the second disk to the first disk according to the data information of the second data, the method further includes: When the copying of the second data from the second disk to the first disk is completed, sending a notification message to the management device for indicating the completion of the copying of the second data.
[0045] In one implementation, when the first disk finishes copying all the second data under the first virtual group to the first disk, the first disk can send a notification message indicating the completion of the second data copy to the management device. At this time, the management device can mark the first virtual group on the faulty disk as the status of having completed replacement and repair.
[0046] In this implementation, by the first disk actively notifying the management device of the completion of the repair, the system can update the virtual group status in a timely manner, avoiding delays or misoperations caused by manual intervention.
[0047] An embodiment of the present application provides a disk repair method, which is applied to a management device. Refer to Figure 2 , Figure 2 is the flowchart of the disk repair method provided by the embodiment of the present application. As Figure 2 shown, it includes the following steps: Step 201, obtain the identification information of the first virtual group on the faulty disk; In the distributed storage system of the present application, multiple virtual groups (Vritual Group, VG) can be defined. Each virtual group has an identification information (Identity, ID). Under the multi-copy redundancy mechanism, the data of each virtual group can be redundantly obtained into multiple data copies, and then the multiple data copies corresponding to each virtual group are respectively stored on different disks. Exemplarily, under the three-copy redundancy mechanism, the data of each virtual group can be redundantly formed into three copies, and then the three data copies are respectively stored on three different disks, and each data copy corresponds to one disk. Then the cluster status management service stores the mapping relationship between the identification information of each virtual group and the identification information of the corresponding three disks in the metadata table (VG_MAP). Thus, when the identification information of the virtual group is known, the disk storing its data copy can be found according to the metadata table, and when the identification information of the disk is known, all the virtual groups on the disk can be found according to the metadata table.
[0048] Refer to Figure 3 , a cluster status management service is deployed on the management device, which can realize the management of the status of each disk in the cluster. During the operation of the cluster, a large number of data writing exceptions occur. The cluster status management service detects the status of each disk. After determining the faulty disk, the operation and maintenance personnel can offline replace the faulty disk with the first disk. The first disk is a new disk with normal functions. Then the cluster status management service performs the data repair task on the faulty disk. First, the cluster status management service can query the identification information of all the virtual groups on the faulty disk according to the metadata table, such as the identification information of the first virtual group on the faulty disk mentioned above.
[0049] Step 202: Query the identification information of the second disk corresponding to the identification information of the first virtual group from the pre-established metadata table. The first data of the first virtual group is stored on the faulty disk, and the second data of the first virtual group is stored on the second disk. The second data is obtained by performing data redundancy on the first data. In this step, the cluster status management service queries, based on the identification information of the first virtual group and using the metadata table, the identification information of all disks storing data replicas of the first virtual group, detects whether the OSD status of the disks where other data replicas are located is normal, and determines a disk as the data repair source for the faulty disk. For example, as mentioned above, the second disk stores the data replica (second data) corresponding to the first data, and the first data on the faulty disk can be repaired using the second data on the second disk.
[0050] Step 203: Send the identification information of the first virtual group and the identification information of the second disk to the first disk, where the first disk is the disk used to replace the faulty disk.
[0051] In this step, after the cluster status management service in the management device determines the first virtual group that needs to be repaired on the faulty disk and the data repair source of the faulty disk - the second disk, it sends the identification message of the first virtual group and the identification message of the second disk to the first disk, so that the first disk can execute the subsequent disk repair process based on this information.
[0052] In an implementation, the management device obtains the identification information of the first virtual group on the faulty disk; queries the identification information of the second disk corresponding to the identification information of the first virtual group from the pre-established metadata table. The first data of the first virtual group is stored on the faulty disk, and the second data of the first virtual group is stored on the second disk. The second data is obtained by performing data redundancy on the first data; sends the identification information of the first virtual group and the identification information of the second disk to the first disk, where the first disk is the disk used to replace the faulty disk.
[0053] In this implementation, by obtaining the identification information of the virtual group of the faulty disk and querying the metadata table, the management device can accurately locate the redundant data source (the second disk), achieving the automation and efficiency of data repair. This mechanism avoids the cumbersome processes of manual intervention and manual search for redundant data, reducing the fault recovery time.
[0054] Optionally, the obtaining the identification information of the first virtual group on the faulty disk includes: Receive the first OSD status information of the first disk sent by the first disk, where the first OSD status information is used to indicate the status of replacing the faulty disk with the first disk; Based on the first OSD status information, obtain the identification information of the first virtual group on the faulty disk.
[0055] In one implementation, after the operation and maintenance personnel replace the faulty disk with the first disk, the storage node service (Object Storage Daemon, OSD) status can be updated, the OSD status is modified to the "new disk replaced status", and the OSD service corresponding to the first disk is pulled up. The OSD service of the first disk will report the first OSD status information to the management device, and the first OSD status information is used to indicate to the management device that the faulty disk has been replaced with the first disk, so that the management device determines to start the data repair process of the faulty disk based on the first OSD status information.
[0056] In this implementation, by receiving the first OSD status information, the management device can perceive the completion of disk replacement in real time, avoid manual intervention, and reduce the fault recovery time.
[0057] Optionally, after sending the identification information of the first virtual group and the identification information of the second disk to the first disk, the method further includes: Receive a notification message sent by the first disk indicating the completion of the second data copy; Based on the notification message, modify the first OSD status information to the second OSD status information, where the first OSD status information is used to indicate the status that the first disk has completed the data repair of the first virtual group.
[0058] In one implementation, when the first disk finishes copying all the second data in the first virtual group to the first disk, the first disk can send a notification message of the completion of the second data copy to the management device. At this time, the management device can mark the first virtual group on the faulty disk as the status of having completed replacement and repair. The management device can also detect all the replacement and repair records on the first disk. When all the replacement and repair records on the first disk are cleared, it indicates that the replacement and repair task of the first disk is completed. At this time, the management device can set the OSD status of the first disk to normal, indicating that all the data on the first disk has been completely repaired.
[0059] In this embodiment, after the first disk finishes data copying, it actively notifies the management device, enabling the management device to confirm the completion of the repair in real time. The management device accurately determines the end of the repair process by marking the virtual group status and clearing the repair record, avoiding redundant operations. At the same time, the OSD status is set to normal to ensure that the cluster status is consistent with the actual data and prevent subsequent failures caused by inconsistent status.
[0060] Optionally, before obtaining the identification information of the first virtual group on the faulty disk, the method further includes: Obtain stored data; Divide the stored data into multiple data slices according to a preset data size; Based on the identification information of the stored data and the identification information of each data slice, determine the identification information of the virtual group corresponding to each data slice; Perform data redundancy on each data slice to obtain a data copy corresponding to each data slice; Query the identification information of the target disk corresponding to the identification information of the target virtual group from a pre-established metadata table, and store the target data copy to the target disk, where the target virtual group is the virtual group whose identification information corresponds to the target data slice, the target data slice is any one of the multiple data slices, and the target data copy is the data copy corresponding to the target data slice.
[0061] In one embodiment, when the business system uploads stored data, the management device divides the stored data into multiple data slices according to a preset data size. Then, according to the hash algorithm, a hash value calculation is performed on the identification information (ID) of the stored data and the identification information (ID) of each data slice, so as to obtain the identification information (ID) of the virtual group corresponding to each data slice.
[0062] Performing data redundancy on each data slice can obtain multiple data copies corresponding to the data slice. Then, according to the corresponding relationship between the virtual group and the disk determined in the previous metadata table, the multiple data copies corresponding to the data slice are respectively stored to multiple disks under the virtual group corresponding to the data slice.
[0063] In this embodiment, efficient and reliable distributed storage is achieved through data sharding and redundant storage.
[0064] An embodiment of the present application provides a disk repair method, which is applied to the second disk. Refer to Figure 4 , Figure 4 is the flowchart of the disk repair method provided by the embodiment of the present application. As Figure 4 shown, it includes the following steps: Step 401: Receive a metadata information query request sent by a first disk, where the metadata information query request carries identification information of a first virtual group; Step 402: Based on the metadata information query request, send metadata information corresponding to the first virtual group to the first disk, where the metadata information includes data information of second data of the first virtual group; Step 403: According to the data information of the second data, send the second data to the first disk.
[0065] Optionally, the step of sending the second data to the first disk according to the data information of the second data includes: Receive a TCP connection request sent by the first disk, and establish a TCP connection channel with the first disk based on the TCP connection request; According to the data information of the second data, send the second data to the first disk based on the TCP connection channel.
[0066] It should be noted that this embodiment, as the implementation manner of the second disk corresponding to the embodiment shown in Figure 1 can refer to the relevant description of the embodiment shown in Figure 1 For the sake of avoiding repeated description, this embodiment will not be elaborated herein, and the same beneficial effects can still be achieved.
[0067] Refer to Figure 5 and Figure 6 Next, the process of the disk repair method according to the embodiments of the present application will be described in detail: 1. A cluster status management service is deployed on a management device, which can manage the status of each disk in the cluster. During the operation of the cluster, a large number of data writing anomalies occur. The cluster status management service detects the status of each disk. After determining a faulty disk, an operation and maintenance personnel can offline replace the faulty disk with a first disk, and the first disk is a new disk with normal functions. After the operation and maintenance personnel replace the faulty disk with the first disk, the OSD status can be updated, the OSD status is modified to "new disk replaced status", and the OSD service corresponding to the first disk is pulled up.
[0068] 2. The OSD service of the first disk reports first OSD status information to the cluster status management service of the management device. The first OSD status information is used to indicate to the cluster status management service of the management device that the faulty disk has been replaced with the first disk, so that the cluster status management service of the management device can determine to start the data repair process of the faulty disk based on the first OSD status information.
[0069] 3. The cluster status management service can query the identification information of all virtual groups on the faulty disk according to the metadata table, such as the identification information of the first virtual group on the faulty disk mentioned above. Then, based on the identification information of the first virtual group, the cluster status management service uses the metadata table to query the identification information of all disks storing data replicas of the first virtual group, detects whether the OSD status of the disks where other data replicas are located is normal, and determines a disk as the data repair source for the faulty disk, such as the second disk mentioned above. The second disk stores the data replica (second data) corresponding to the first data, and the first data on the faulty disk can be repaired using the second data on the second disk.
[0070] 4. After the cluster status management service determines the first virtual group to be repaired on the faulty disk and the data repair source of the faulty disk - the second disk, it notifies the OSD service of the first disk of the identification message of the first virtual group and the identification message of the second disk. The first disk can use other disks (i.e., the second disk mentioned above) storing the data replica (i.e., the second data) of the first data to perform the data recovery process.
[0071] 5. The OSD service of the first disk can first send a metadata information query request to the OSD service of the second disk. The metadata information query request carries the identification information of the first virtual group to obtain the metadata information under the first virtual group from the second disk. Then, the first disk indexes the specific second data according to the metadata information and copies the second data from the second disk through the TCP connection channel.
[0072] 6. When the first disk finishes copying all the second data under the first virtual group to the first disk, the OSD service of the first disk can send a notification message indicating that the copying of the second data is completed to the cluster status management service of the management device. At this time, the management device can mark the first virtual group on the faulty disk as the state of having completed replacement and repair. The cluster status management service of the management device can also detect all the replacement and repair records on the first disk. When all the replacement and repair records on the first disk are cleared, it indicates that the replacement and repair task of the first disk is completed. At this time, the cluster status management service of the management device can set the OSD status of the first disk to normal, indicating that all the data on the first disk has been completely repaired and the first disk can be used.
[0073] Figure 6 The Metadata Server (MDS) in it is the core component responsible for managing metadata. In the disk replacement and repair process, MDS dynamically updates the status of virtual groups and disks through the metadata table and triggers the data repair process. After the repair is completed, MDS marks the corresponding virtual group as "completed replacement and repair", and finally sets the OSD status of the new disk to "normal".
[0074] See Figure 7 , Figure 7 which is a structural diagram of a disk repair device provided in an embodiment of the present application. The disk repair device is applied to a first disk. As Figure 7 shown, the disk repair device 700 includes: A first receiving module 701, configured to receive the identification information of the first virtual group on the faulty disk and the identification information of the second disk corresponding to the identification information of the first virtual group sent by the management device. Wherein, the first disk is a disk used to replace the faulty disk, the first data of the first virtual group is stored on the faulty disk, and the second data of the first virtual group is stored on the second disk, and the second data is obtained by performing data redundancy on the first data; A first sending module 702, configured to send a metadata information query request to the second disk, where the metadata information query request carries the identification information of the first virtual group; A second receiving module 703, configured to receive the metadata information corresponding to the first virtual group sent by the second disk; A first parsing module 704, configured to parse the data information of the second data stored on the second disk based on the metadata information corresponding to the first virtual group; A first copying module 705, configured to copy the second data from the second disk to the first disk according to the data information of the second data.
[0075] Optionally, the first copying module includes: A first establishing unit, configured to send a Transmission Control Protocol (TCP) connection request to the second disk and establish a TCP connection channel with the second disk based on the TCP connection request; A first copying unit, configured to copy the second data from the second disk to the first disk based on the TCP connection channel according to the data information of the second data.
[0076] Optionally, the device further includes: A second sending module, configured to send the first Object-based Storage Device (OSD) status information of the first disk to the management device, where the first OSD status information is used to indicate the status of replacing the faulty disk with the first disk.
[0077] Optionally, the device further includes: A third sending module, configured to send a notification message indicating the completion of copying the second data to the management device when the copying of the second data from the second disk to the first disk is completed.
[0078] See Figure 8 , Figure 8 which is a structural diagram of a disk repair device provided by an embodiment of the present application. The disk repair device is applied to a management device. As Figure 8 shown, the disk repair device 800 includes: A first acquisition module 801, configured to acquire identification information of a first virtual group on a faulty disk; A first query module 802, configured to query, from a pre-established metadata table, identification information of a second disk corresponding to the identification information of the first virtual group. The first data belonging to the first virtual group is stored on the faulty disk, and the second data belonging to the first virtual group is stored on the second disk. The second data is obtained by performing data redundancy on the first data; A fourth sending module 803, configured to send the identification information of the first virtual group and the identification information of the second disk to a first disk, where the first disk is a disk used to replace the faulty disk.
[0079] Optionally, the first acquisition module includes: A first receiving unit, configured to receive first OSD status information of the first disk sent by the first disk, where the first OSD status information is used to indicate a status of replacing the faulty disk with the first disk; A first acquisition unit, configured to acquire, based on the first OSD status information, identification information of a first virtual group on the faulty disk.
[0080] Optionally, the device further includes: A third receiving module, configured to receive a notification message sent by the first disk for indicating that the copying of the second data is completed; A first modification module, configured to modify, based on the notification message, the first OSD status information to second OSD status information, where the first OSD status information is used to indicate a status that the first disk has completed data repair of the first virtual group.
[0081] Optionally, the device further includes: A second acquisition module, configured to acquire stored data; A first slicing module, configured to divide the stored data into multiple data slices according to a preset data size; A first determination module, configured to determine, based on the identification information of the stored data and the identification information of each data slice, identification information of a virtual group corresponding to each data slice; A first redundancy module, configured to perform data redundancy on each data slice to obtain a data copy corresponding to each data slice; The first query module is configured to query, from a pre-established metadata table, the identification information of a target disk corresponding to the identification information of a target virtual group, and store a target data copy into the target disk. The target virtual group is a virtual group whose identification information corresponds to a target data slice, the target data slice is any one of the multiple data slices, and the target data copy is a data copy corresponding to the target data slice.
[0082] See Figure 9 , Figure 9 is a structural diagram of a disk repair device provided in an embodiment of the present application. The disk repair device is applied to a second disk. As Figure 9 shown, the disk repair device 900 includes: A fourth receiving module 901, configured to receive a metadata information query request sent by a first disk, where the metadata information query request carries the identification information of a first virtual group; A fifth sending module 902, configured to send, based on the metadata information query request, metadata information corresponding to the first virtual group to the first disk, where the metadata information includes data information of second data of the first virtual group; A sixth sending module 903, configured to send the second data to the first disk according to the data information of the second data.
[0083] Optionally, the sixth sending module includes: A second receiving unit, configured to receive a TCP connection request sent by the first disk, and establish a TCP connection channel with the first disk based on the TCP connection request; A first sending unit, configured to send the second data to the first disk based on the TCP connection channel according to the data information of the second data.
[0084] An embodiment of the present application further provides a first disk. Since the principle of the first disk for solving problems is similar to that of the disk repair method in the embodiment of the present application, the implementation of the first disk can refer to the implementation of the method, and the repeated parts will not be described again. As Figure 10 shown, the first disk in the embodiment of the present application includes: a processor 1000, configured to read a program in a memory 1020 and execute the following processes: through a transceiver 1010: Receive the identification information of a first virtual group on a failed disk and the identification information of a second disk corresponding to the identification information of the first virtual group sent by a management device, where the first disk is a disk used to replace the failed disk, the first data of the first virtual group is stored on the failed disk, and the second data of the first virtual group is stored on the second disk, and the second data is obtained by performing data redundancy on the first data; Send a metadata information query request to the second disk, where the metadata information query request carries the identification information of the first virtual group; Receive the metadata information corresponding to the first virtual group sent by the second disk; The processor 1000 is configured to read the program in the memory 1020 and execute the following processes: Based on the metadata information corresponding to the first virtual group, parse the data information of the second data stored on the second disk; According to the data information of the second data, copy the second data from the second disk to the first disk.
[0085] Among them, in Figure 10 The bus architecture may include any number of interconnected buses and bridges, specifically, various circuits represented by one or more processors represented by the processor 1000 and the memory represented by the memory 1020 are linked together. The bus architecture can also link together various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art, and therefore, they will not be further described herein. The bus interface provides an interface. The transceiver 1010 can be multiple components, that is, including a transmitter and a transceiver, and provides a unit for communicating with various other devices on the transmission medium. The processor 1000 is responsible for managing the bus architecture and general processing, and the memory 1020 can store the data used by the processor 1000 when performing operations.
[0086] Optionally, the processor 1000 is configured to read the program in the memory 1020 and execute the following processes: Send a Transmission Control Protocol (TCP) connection request to the second disk and establish a TCP connection channel with the second disk based on the TCP connection request; According to the data information of the second data, copy the second data from the second disk to the first disk based on the TCP connection channel.
[0087] Optionally, before receiving the identification information of the first virtual group on the failed disk and the identification information of the second disk corresponding to the identification information of the first virtual group sent by the management device, the method further includes: Send the first Open Storage Device (OSD) status information of the first disk to the management device, where the first OSD status information is used to indicate the status of replacing the failed disk with the first disk.
[0088] Optionally, the processor 1000 is configured to read the program in the memory 1020 and execute the following processes: through the transceiver 1010: When the copying of the second data from the second disk to the first disk is completed, a notification message for indicating the completion of the copying of the second data is sent to the management device.
[0089] An embodiment of the present application further provides a management device. Since the principle for the management device to solve problems is similar to the disk repair method in the embodiment of the present application, the implementation of the management device can refer to the implementation of the method, and the repeated parts will not be described again. As Figure 11 shown, the management device in the embodiment of the present application includes: a processor 1100, configured to read a program in a memory 1120 and execute the following processes: Obtain identification information of a first virtual group on a faulty disk; Query, from a pre-established metadata table, identification information of a second disk corresponding to the identification information of the first virtual group, where the first data of the first virtual group is stored on the faulty disk, and the second data of the first virtual group is stored on the second disk, and the second data is obtained by performing data redundancy on the first data; The processor 1100 is configured to read a program in the memory 1120 and execute the following processes: through a transceiver 1110: Send the identification information of the first virtual group and the identification information of the second disk to a first disk, where the first disk is a disk used to replace the faulty disk.
[0090] Among them, in Figure 11 , the bus architecture may include any number of interconnected buses and bridges. Specifically, various circuits represented by one or more processors represented by the processor 1100 and a memory represented by the memory 1120 are linked together. The bus architecture may also link together various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art. Therefore, they will not be further described herein. The bus interface provides an interface. The transceiver 1110 may be multiple components, that is, including a transmitter and a transceiver, and provides a unit for communicating with various other devices on a transmission medium. The processor 1100 is responsible for managing the bus architecture and general processing, and the memory 1120 may store data used by the processor 1100 when executing operations.
[0091] Optionally, the processor 1100 is configured to read a program in the memory 1120 and execute the following processes: Receive first OSD status information of the first disk sent by the first disk, where the first OSD status information is used to indicate the status of replacing the faulty disk with the first disk; Based on the first OSD status information, obtain identification information of a first virtual group on the faulty disk.
[0092] Optionally, the processor 1100 is configured to read a program in the memory 1120 and execute the following process: through the transceiver 1110: Receive a notification message sent by the first disk indicating that the second data copy is completed; The processor 1100 is configured to read a program in the memory 1120 and execute the following process: Based on the notification message, modify the first OSD status information to second OSD status information, where the first OSD status information is used to indicate the status that the first disk has completed data repair of the first virtual group.
[0093] Optionally, the processor 1100 is configured to read a program in the memory 1120 and execute the following process: Obtain stored data; Divide the stored data into multiple data slices according to a preset data size; Based on the identification information of the stored data and the identification information of each data slice, determine the identification information of the virtual group corresponding to each data slice; Perform data redundancy on each data slice to obtain a data copy corresponding to each data slice; Query, from a pre-established metadata table, the identification information of the target disk corresponding to the identification information of the target virtual group, and store the target data copy to the target disk, where the target virtual group is the virtual group whose identification information corresponds to the target data slice, the target data slice is any one of the multiple data slices, and the target data copy is the data copy corresponding to the target data slice.
[0094] An embodiment of the present application further provides a second disk. Since the principle of the second disk to solve problems is similar to the disk repair method in the embodiment of the present application, the implementation of the second disk can refer to the implementation of the method, and the repeated parts will not be described again. As Figure 12 shown, in the second disk of the embodiment of the present application, it includes: a processor 1200, configured to read a program in the memory 1220 and execute the following process: through the transceiver 1210: Receive a metadata information query request sent by the first disk, where the metadata information query request carries the identification information of the first virtual group; Based on the metadata information query request, send the metadata information corresponding to the first virtual group to the first disk, where the metadata information includes the data information of the second data of the first virtual group; Send the second data to the first disk according to the data information of the second data.
[0095] Among them, inFigure 12 In this case, the bus architecture may include any number of interconnected buses and bridges, and various circuits of one or more processors represented by the processor 1200 and the memory represented by the memory 1220 are specifically linked together. The bus architecture may also link together various other circuits such as peripheral devices, voltage regulators, and power management circuits, etc., which are well known in the art, and thus will not be further described herein. The bus interface provides an interface. The transceiver 1210 may be multiple components, that is, including a transmitter and a transceiver, and provides a unit for communicating with various other devices on the transmission medium. The processor 1200 is responsible for managing the bus architecture and general processing, and the memory 1220 may store data used by the processor 1200 when executing operations.
[0096] Optionally, the processor 1200 is configured to read a program in the memory 1220 and execute the following processes: through the transceiver 1210: Receive a TCP connection request sent by the first disk, and establish a TCP connection channel with the first disk based on the TCP connection request; Send the second data to the first disk based on the TCP connection channel according to the data information of the second data.
[0097] The embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements each process of the above-mentioned disk repair method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here. Among them, the computer-readable storage medium is, for example, a Read-Only Memory (ROM), a Random Access Memory (RAM), a magnetic disk, or an optical disc, etc.
[0098] The embodiment of the present application further provides a computer program product, including computer instructions, which implement each process of the above-mentioned Figure 1 、 Figure 2 or Figure 4 shown method embodiments when executed by a processor, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0099] It should be noted that in this text, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without more limitations, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including that element.
[0100] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of this application.
[0101] The embodiments of this application have been described above in conjunction with the accompanying drawings. However, this application is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative and not restrictive. Those of ordinary skill in the art, under the inspiration of this application and without departing from the purpose of this application and the scope protected by the claims, can also make many forms, all of which fall within the protection scope of this application.
Claims
1. A disk repair method, characterized in that, Applied to the first disk, the method includes: Receiving identification information of a first virtual group on a failed disk and identification information of a second disk corresponding to the identification information of the first virtual group sent by a management device, where the first disk is a disk used to replace the failed disk, the first data of the first virtual group is stored on the failed disk, the second data of the first virtual group is stored on the second disk, and the second data is obtained by performing data redundancy on the first data; Sending a metadata information query request to the second disk, where the metadata information query request carries the identification information of the first virtual group; Receiving the metadata information corresponding to the first virtual group sent by the second disk; Based on the metadata information corresponding to the first virtual group, parsing out the data information of the second data stored on the second disk; According to the data information of the second data, copying the second data from the second disk to the first disk.
2. The disk repair method according to claim 1, wherein The copying the second data from the second disk to the first disk according to the data information of the second data includes: Sending a Transmission Control Protocol (TCP) connection request to the second disk and establishing a TCP connection channel with the second disk based on the TCP connection request; According to the data information of the second data, copying the second data from the second disk to the first disk based on the TCP connection channel.
3. The disk repair method according to claim 1, characterized in that Before receiving the identification information of the first virtual group on the failed disk and the identification information of the second disk corresponding to the identification information of the first virtual group sent by the management device, the method further includes: Sending the first Object-based Storage Device (OSD) status information of the first disk to the management device, where the first OSD status information is used to indicate the status of replacing the failed disk with the first disk.
4. The disk repair method according to claim 1, characterized in that After copying the second data from the second disk to the first disk according to the data information of the second data, the method further includes: When the copying of the second data from the second disk to the first disk is completed, sending a notification message indicating the completion of the copying of the second data to the management device.
5. A disk repair method, characterized in that, Applied to the management device, the method includes: Obtaining the identification information of the first virtual group on the failed disk; Querying, from a pre-established metadata table, the identification information of the second disk corresponding to the identification information of the first virtual group, where the first data of the first virtual group is stored on the failed disk, the second data of the first virtual group is stored on the second disk, and the second data is obtained by performing data redundancy on the first data; Sending the identification information of the first virtual group and the identification information of the second disk to the first disk, where the first disk is a disk used to replace the failed disk.
6. The disk repair method according to claim 5, wherein, The obtaining the identification information of the first virtual group on the failed disk includes: Receiving the first OSD status information of the first disk sent by the first disk, where the first OSD status information is used to indicate the status of replacing the failed disk with the first disk; Based on the first OSD status information, obtain the identification information of the first virtual group on the faulty disk.
7. The disk repair method according to claim 6, characterized in that After sending the identification information of the first virtual group and the identification information of the second disk to the first disk, the method further includes: Receiving a notification message sent by the first disk for indicating the completion of the second data copy; Based on the notification message, modifying the first OSD status information to second OSD status information, where the first OSD status information is used to indicate the status that the first disk has completed the data repair of the first virtual group.
8. The disk repair method according to any one of claims 5 to 7, characterized in that, Before obtaining the identification information of the first virtual group on the faulty disk, the method further includes: Obtaining stored data; Dividing the stored data into multiple data slices according to a preset data size; Based on the identification information of the stored data and the identification information of each data slice, determining the identification information of the virtual group corresponding to each data slice; Performing data redundancy on each data slice to obtain a data copy corresponding to each data slice; Querying, from a pre-established metadata table, the identification information of the target disk corresponding to the identification information of the target virtual group, and storing the target data copy to the target disk, where the target virtual group is the virtual group corresponding to the identification information of the target data slice, the target data slice is any one of the multiple data slices, and the target data copy is the data copy corresponding to the target data slice.
9. A disk repair method, characterized in that, Applied to the second disk, the method includes: Receiving a metadata information query request sent by the first disk, where the metadata information query request carries the identification information of the first virtual group; Based on the metadata information query request, sending metadata information corresponding to the first virtual group to the first disk, where the metadata information includes the data information of the second data of the first virtual group; Sending the second data to the first disk according to the data information of the second data.
10. The disk repair method according to claim 9, wherein The sending the second data to the first disk according to the data information of the second data includes: Receiving a TCP connection request sent by the first disk, and establishing a TCP connection channel with the first disk based on the TCP connection request; Based on the data information of the second data, sending the second data to the first disk through the TCP connection channel.
11. A disk repair device, characterized in that, Applied to the first disk, the apparatus includes: A first receiving module, configured to receive the identification information of the first virtual group on the faulty disk sent by the management device and the identification information of the second disk corresponding to the identification information of the first virtual group, where the first disk is a disk for replacing the faulty disk, the first data of the first virtual group is stored on the faulty disk, the second data of the first virtual group is stored on the second disk, and the second data is obtained by performing data redundancy on the first data; A first sending module, configured to send a metadata information query request to the second disk, where the metadata information query request carries the identification information of the first virtual group; A second receiving module, configured to receive metadata information corresponding to the first virtual group sent by the second disk; A first parsing module, configured to parse data information of the second data stored on the second disk based on the metadata information corresponding to the first virtual group; A first copying module, configured to copy the second data from the second disk to the first disk according to the data information of the second data; 12. A disk repair device, characterized in that, Applied to a management device, the apparatus includes: A first obtaining module, configured to obtain identification information of a first virtual group on a faulty disk; A first querying module, configured to query, from a pre-established metadata table, identification information of a second disk corresponding to the identification information of the first virtual group, where the faulty disk stores first data belonging to the first virtual group, the second disk stores second data belonging to the first virtual group, and the second data is obtained by performing data redundancy on the first data; A fourth sending module, configured to send the identification information of the first virtual group and the identification information of the second disk to a first disk, where the first disk is a disk for replacing the faulty disk; 13. A disk repair device, characterized in that, Applied to a second disk, the apparatus includes: A fourth receiving module, configured to receive a metadata information query request sent by the first disk, where the metadata information query request carries the identification information of the first virtual group; A fifth sending module, configured to send, based on the metadata information query request, metadata information corresponding to the first virtual group to the first disk, where the metadata information includes data information of the second data of the first virtual group; A sixth sending module, configured to send the second data to the first disk according to the data information of the second data; 14. A first disk, characterized in that, The first disk includes a transceiver and a processor, and the transceiver is configured to: Receive the identification information of the first virtual group on the faulty disk and the identification information of the second disk corresponding to the identification information of the first virtual group, where the first disk is a disk for replacing the faulty disk, the faulty disk stores the first data of the first virtual group, the second disk stores the second data of the first virtual group, and the second data is obtained by performing data redundancy on the first data; Send a metadata information query request to the second disk, where the metadata information query request carries the identification information of the first virtual group; Receive the metadata information corresponding to the first virtual group sent by the second disk; The processor is configured to: Parse the data information of the second data stored on the second disk based on the metadata information corresponding to the first virtual group; The transceiver is configured to: Copy the second data from the second disk to the first disk according to the data information of the second data; 15. A management device, characterized in that, The management device includes a transceiver and a processor, and the processor is configured to: Obtain the identification information of the first virtual group on the faulty disk; Query the identification information of the second disk corresponding to the identification information of the first virtual group from a pre-established metadata table. The first data belonging to the first virtual group is stored on the faulty disk, and the second data belonging to the first virtual group is stored on the second disk. The second data is obtained by data redundancy of the first data. The transceiver is configured to: Send the identification information of the first virtual group and the identification information of the second disk to the first disk, where the first disk is a disk used to replace the faulty disk.
16. A second disk, characterized in that, The second disk includes a transceiver and a processor. The transceiver is configured to: Receive a metadata information query request sent by the first disk, where the metadata information query request carries the identification information of the first virtual group. Based on the metadata information query request, send the metadata information corresponding to the first virtual group to the first disk, where the metadata information includes the data information of the second data of the first virtual group. Send the second data to the first disk according to the data information of the second data.
17. An electronic device, characterized in that, Comprising a processor, a memory, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps of the disk repair method according to any one of claims 1 to 4, or when the computer program is executed by the processor, it implements the steps of the disk repair method according to any one of claims 5 to 8, or when the computer program is executed by the processor, it implements the steps of the disk repair method according to claim 9 or 10.
18. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is executed by the processor, it implements the steps of the disk repair method according to any one of claims 1 to 4, or when the computer program is executed by the processor, it implements the steps of the disk repair method according to any one of claims 5 to 8, or when the computer program is executed by the processor, it implements the steps of the disk repair method according to claim 9 or 10.
19. A computer program product, characterized in that, Comprising computer instructions. When the computer instructions are executed by the processor, it implements the steps of the disk repair method according to any one of claims 1 to 4, or when the computer instructions are executed by the processor, it implements the steps of the disk repair method according to any one of claims 5 to 8, or when the computer instructions are executed by the processor, it implements the steps of the disk repair method according to claim 9 or 10.
Citation Information
Patent Citations
Disk array fault recovery method and device
CN114518976A