Data recovery method and related device
By saving the recovered information of the failed data block written to the target storage node in the distributed storage system and reporting it to the management node, the problem of data recovery failure caused by management node failure or network fluctuations is solved, and the complete recovery of data blocks and data consistency is achieved.
Patent Information
- Application Number
- CN202210343229.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-31
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-03-31
AI Technical Summary
In a distributed storage system, management node failure or network fluctuations lead to the inability to accurately and completely record the relevant information of failed data blocks and verification blocks, which in turn cannot ensure the consistency of data before and after storage.
When the management node fails to successfully record the data to be recovered corresponding to the failed data block, the client saves it to the target storage node among multiple storage nodes, and then reports it to the management node by the target storage node to realize the recovery of the data block.
Through this method, data block recovery failure caused by management node exceptions or network exceptions is avoided, and each data block that failed to write can be restored, ensuring the consistency of data before and after storage.
Smart Images

Figure CN114880165B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of distributed storage, and in particular, to a data recovery method and related device. Background Art
[0002] Distributed storage means that the client splits the data to be stored into multiple data blocks, encodes the multiple data blocks through an erasure code algorithm to obtain redundant check blocks, and then stores each data block and check block on different storage nodes respectively to implement the storage of the data to be stored. When there are data with storage failures among the data blocks and check blocks, the data with storage failures can be recovered through the successfully stored data.
[0003] Suppose the client splits the data to be stored into n data blocks, and m check blocks are obtained by encoding the n data blocks. During the process of storing the n data blocks and m check blocks to the corresponding storage nodes, if the sum of the successfully stored data blocks and check blocks is greater than n and less than n + m, the client needs to report the relevant information of all data blocks and / or check blocks with storage failures to the management node for recording, so that the data blocks and / or check blocks with storage failures can be recovered according to the successfully stored data blocks and / or check blocks later.
[0004] If, when the client sends the relevant information of the data blocks and / or check blocks with storage failures, the management node fails and undergoes primary and standby switching, or the network fluctuates, at this time, the management node cannot accurately and completely record the relevant information of the data blocks and / or check blocks with storage failures, and the data cannot be recovered later, thus the consistency of the data before and after storage cannot be guaranteed. Summary of the Invention
[0005] To overcome the deficiencies of the prior art and ensure the consistency of the data before and after storage, embodiments of the present invention provide a data recovery method and related device.
[0006] The technical solutions provided by the embodiments of the present invention are as follows:
[0007] In a first aspect, an embodiment of the present invention provides a data recovery method for a client in a distributed storage system. The distributed storage system further includes a management node and multiple storage nodes. The client is communicatively connected to the management node, the client is communicatively connected to the multiple storage nodes, and the management node is communicatively connected to the multiple storage nodes. The client sends a write data request to the multiple storage nodes, where each storage node needs to write the data block in the write data request it receives. The method includes:
[0008] Receive response messages returned by each of the storage nodes for their respective write data requests, where each response message is used to indicate whether each storage node has successfully written the data block in its respective write data request;
[0009] If there are response messages indicating write failures, and the number of the response messages indicating write failures is less than a preset value, send the data recovery information corresponding to each write-failed data block to the management node, so that the management node records each piece of the data recovery information;
[0010] If the management node fails to record, save each piece of the data recovery information to a target storage node among the multiple storage nodes, so that the target storage node reports each piece of the data recovery information to the management node, and after the reporting is successful, the management node restores the corresponding write-failed data block according to each piece of the data recovery information.
[0011] Optionally, the data block includes a first data block and a second data block, and the second data block is parity data obtained by performing erasure coding on the first data block. Before the step of saving each piece of the data recovery information to a target storage node among the multiple storage nodes, the method further includes:
[0012] Calculate the number of candidate target storage nodes according to the number of storage nodes, the number of first data blocks, and the number of second data blocks;
[0013] Use the candidate target storage nodes with the smallest storage space utilization rate among the calculated number of candidates as the target storage nodes.
[0014] In a second aspect, an embodiment of the present invention provides a data recovery method for a target storage node among multiple storage nodes of a distributed storage system. The distributed storage system further includes a client and a management node. The client is communicatively connected to the management node, the client is communicatively connected to the multiple storage nodes, the management node is communicatively connected to the multiple storage nodes, and the client sends a write data request to the multiple storage nodes. Each storage node needs to write the data block in the write data request it receives. The method includes:
[0015] Receive the data recovery information corresponding to each write-failed data block sent by the client. The data recovery information is sent by the client to the target storage node when the management node fails to record the data recovery information. The number of write-failed data blocks is the same as the number of response messages indicating write failures. Each response message is generated by each storage node for its respective write data request and returned to the client;
[0016] Save each piece of the data information to be restored, and report it to the management node. After the reporting is successful, the management node restores the corresponding data block with a write failure according to each piece of the data information to be restored.
[0017] Optionally, the management node is in an online state, and the target storage node records the save time of each piece of the data information to be restored. The method further includes:
[0018] Receive the reported quantity sent by the management node;
[0019] Send the earliest-reported quantity of the data information to be restored to the management node.
[0020] Optionally, the management node is in an offline state. The method further includes:
[0021] When the management node recovers from the offline state to the online state, send all the data information to be restored to the management node.
[0022] In a third aspect, an embodiment of the present invention provides a data recovery method for a management node applied to a distributed storage system. The distributed storage system further includes a client and multiple storage nodes. The client is communicatively connected to the management node, the client is communicatively connected to the multiple storage nodes, the management node is communicatively connected to the multiple storage nodes, and the client sends a write data request to the multiple storage nodes. Each storage node needs to write the data block in the write data request it receives. The method includes:
[0023] Receive the data information to be restored corresponding to each data block with a write failure sent by a target storage node among the multiple storage nodes. The data information to be restored sent by the target storage node is saved to the target storage node by the client when the management node fails to record the data information to be restored. The number of the data blocks with a write failure is the same as the number of response messages indicating a write failure. Each response message is generated by each storage node for its respective write data request and returned to the client;
[0024] Restore the corresponding data block with a write failure according to each piece of the data information to be restored.
[0025] Optionally, the method further includes:
[0026] Obtain the CPU occupancy rate, CPU load, and the number of the storage nodes;
[0027] Calculate the reporting quantity according to the total quantity of the data information to be restored, the CPU occupancy rate, the CPU load, and the number of the storage nodes, where the reporting quantity is used to represent the maximum quantity of the data information to be restored that the target storage node can report;
[0028] Send the reporting quantity to each of the storage nodes, so that when each of the storage nodes serves as a target storage node, it sends the data information to be restored locally saved to the management node according to the reporting quantity.
[0029] In a fourth aspect, an embodiment of the present invention provides a data recovery device for a client in a distributed storage system. The distributed storage system further includes a management node and a plurality of storage nodes. The client is communicatively connected to the management node, the client is communicatively connected to the plurality of storage nodes, and the management node is communicatively connected to the plurality of storage nodes. The client sends a write data request to the plurality of storage nodes. Each of the storage nodes needs to write the data block in the received write data request. The device includes:
[0030] A first receiving module, configured to receive a response message returned by each of the storage nodes for its respective write data request, where each response message is used to represent whether each storage node successfully writes the data block in its respective write data request;
[0031] A first sending module, configured to, if there is a response message indicating a write failure and the number of the response messages indicating a write failure is less than a preset value, send the data information to be restored corresponding to each write-failed data block to the management node, so that the management node records each piece of the data information to be restored;
[0032] The first sending module is further configured to, if the management node records a failure, save each piece of the data information to be restored to a target storage node among the plurality of storage nodes, so that the target storage node reports each piece of the data information to be restored to the management node, and after the reporting is successful, the management node restores the corresponding write-failed data block according to each piece of the data information to be restored.
[0033] In a fifth aspect, an embodiment of the present invention provides a target storage node among a plurality of storage nodes in a distributed storage system. The distributed storage system further includes a client and a management node. The client is communicatively connected to the management node, the client is communicatively connected to the plurality of storage nodes, and the management node is communicatively connected to the plurality of storage nodes. The client sends a write data request to the plurality of storage nodes. Each of the storage nodes needs to write the data block in the received write data request. The device includes:
[0034] A second receiving module, configured to receive the data information to be recovered corresponding to each data block with a write failure sent by the client, where the data information to be recovered is sent by the client to the target storage node when the management node fails to record the data information to be recovered. The number of data blocks with write failures is the same as the number of response messages indicating write failures, and each of the response messages is generated by each storage node for its respective write data request and returned to the client;
[0035] A storage module, configured to save each piece of the data information to be recovered for reporting to the management node, and after successful reporting, enable the management node to recover the corresponding data block with a write failure according to each piece of the data information to be recovered.
[0036] In a sixth aspect, an embodiment of the present invention provides a management node applied to a distributed storage system. The distributed storage system further includes a client and multiple storage nodes. The client is communicatively connected to the management node, the client is communicatively connected to the multiple storage nodes, and the management node is communicatively connected to the multiple storage nodes. The client sends a write data request to the multiple storage nodes. Each of the storage nodes needs to write the data block in the write data request it receives. The apparatus includes:
[0037] A third receiving module, configured to receive the data information to be recovered corresponding to each data block with a write failure sent by a target storage node among the multiple storage nodes. The data information to be recovered sent by the target storage node is saved to the target storage point when the management node fails to record the data information to be recovered. The number of data blocks with write failures is the same as the number of response messages indicating write failures, and each of the response messages is generated by each storage node for its respective write data request and returned to the client;
[0038] A recovery module, configured to recover the corresponding data block with a write failure according to each piece of the data information to be recovered.
[0039] In a seventh aspect, an embodiment of the present invention provides a distributed storage system. The distributed storage system includes a client, a management node, and multiple storage nodes. The client is communicatively connected to the management node, the client is communicatively connected to the multiple storage nodes, and the management node is communicatively connected to the multiple storage nodes;
[0040] The client is configured to send a write data request to the multiple storage nodes;
[0041] Each of the storage nodes is configured to receive a write data request sent by the client and write the data block in its respective write data request, where the write data requests received by each of the storage nodes are different;
[0042] The client is further configured to receive a response message returned by each of the storage nodes for its respective write data request, where each response message is used to indicate whether each storage node has successfully written the data block in its respective write data request;
[0043] The client is further configured to, if there is a response message indicating a write failure and the number of the response messages indicating a write failure is less than a preset value, send the data information to be recovered corresponding to each data block with a write failure to the management node, so that the management node records each piece of the data information to be recovered;
[0044] The client is further configured to, if the management node records a failure, save each piece of the data information to be recovered to a target storage node among the multiple storage nodes;
[0045] The target storage node is configured to report each piece of the data information to be recovered to the management node;
[0046] The management node is configured to, after the target storage node reports successfully, recover the corresponding data block with a write failure according to each piece of the data information to be recovered.
[0047] In a eighth aspect, an embodiment of the present invention provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, it implements the data recovery method described in the first aspect, or implements the data recovery method described in the second aspect, or implements the data recovery method described in the third aspect.
[0048] In a ninth aspect, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the data recovery method described in the first aspect, or implements the data recovery method described in the second aspect, or implements the data recovery method described in the third aspect.
[0049] Compared with the prior art, in the technical solution provided by the embodiment of the present invention, when the management node fails to successfully record the data recovery information corresponding to the data block with a write failure, the client saves the data recovery information corresponding to the data block with a write failure to a target storage node among multiple storage nodes, so that the target storage node subsequently reports the data recovery information corresponding to the data block with a write failure to the management node, and after the successful reporting, the management node restores each data block with a write failure corresponding to each piece of data recovery information, avoiding the inability to accurately and completely record the data recovery information corresponding to the data block with a write failure due to an exception of the management node or a network exception, enabling each data block with a write failure to be restored, and thus ensuring the consistency of the data before and after storage. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0051] Figure 1 FIG. 1 is a schematic structural diagram of a distributed storage system provided by an embodiment of the present invention;
[0052] Figure 2 FIG. 2 is a schematic block diagram of the structure of a computer device provided by an embodiment of the present invention;
[0053] Figure 3 FIG. 3 is a schematic flowchart of a data recovery method provided by an embodiment of the present invention;
[0054] Figure 4 FIG. 4 is a schematic flowchart of a method for determining a target storage node provided by an embodiment of the present invention;
[0055] Figure 5 FIG. 5 is another schematic flowchart of a data recovery method provided by an embodiment of the present invention;
[0056] Figure 6 FIG. 6 is a schematic flowchart of a method for reporting data recovery information provided by an embodiment of the present invention;
[0057] Figure 7 FIG. 7 is still another schematic flowchart of a data recovery method provided by an embodiment of the present invention;
[0058] Figure 8 FIG. 8 is a schematic flowchart of a method for determining the number of reported data recovery information provided by an embodiment of the present invention;
[0059] Figure 9A block diagram of a data recovery device applied to a client provided by an embodiment of the present invention;
[0060] Figure 10 A block diagram of a data recovery device applied to a target storage node provided by an embodiment of the present invention;
[0061] Figure 11 A block diagram of a data recovery device applied to a management node provided by an embodiment of the present invention.
[0062] Icons: 100 - computer device; 110 - memory; 120 - processor; 200 - data recovery device applied to the client; 201 - first receiving module; 202 - first sending module; 203 - first determining module; 300 - data recovery device applied to the target storage node; 301 - second receiving module; 302 - storage module; 303 - second sending module; 400 - data recovery device applied to the management node; 401 - third receiving module; 402 - recovery module; 403 - second determining module; 404 - third sending module. Detailed implementation manners
[0063] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.
[0064] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0065] It should be noted that: Similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0066] In addition, terms such as "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance.
[0067] It should be noted that, without conflict, the features in the embodiments of the present invention can be combined with each other.
[0068] In order to overcome the problem that during the process of master-slave switchover due to a failure of the management node or during a network fluctuation period, the relevant information of the data blocks and / or parity blocks with storage failures sent by the client cannot be accurately and completely recorded by the management node, resulting in the subsequent inability to recover these data blocks and / or parity blocks with storage failures, the common approach is to periodically scan the data blocks and / or parity blocks stored in each storage node in the distributed storage system, which consumes a large amount of resources.
[0069] In view of this, embodiments of the present invention provide a data recovery method and related device to reduce the consumed resources on the premise of ensuring data consistency before and after storage. The following will be described in detail.
[0070] Please refer to Figure 1 , Figure 1 which is a schematic structural diagram of a distributed storage system provided by an embodiment of the present invention. The distributed storage system includes a client, multiple storage nodes, and a management node. The client is communicatively connected to the management node, the client is communicatively connected to multiple storage nodes, and the management node is communicatively connected to multiple storage nodes.
[0071] The client (Client) can interact with the upper-layer application or an external host, can receive data from the upper-layer application or an external host, encode the received data through an erasure code algorithm to obtain redundant parity blocks, distribute each data block and parity block to different storage nodes for storage by sending a write data request to the storage nodes, and read the stored data from the storage nodes by sending a read data request to the storage nodes. The client can also be used to send the relevant information of the data blocks and / or parity blocks with storage failures to the management node for recording. In particular, when the management node fails to successfully record the relevant information of the data blocks and / or parity blocks with storage failures, the client is also used to save the relevant information of the data blocks and / or parity blocks with storage failures to the target storage node among the multiple storage nodes. The client can be a server, a personal computer (Personal Computer, hereinafter referred to as PC), a laptop computer, etc. The client can also be one or more program modules on a device, or a virtual machine or container running on a device. The client can also be a cluster composed of multiple devices, for example, it can be a collective term for multiple program modules distributed on multiple devices.
[0072] The storage node receives read data requests and / or write data requests to read the data stored by the client, and / or store data blocks and / or parity blocks from the client. When any storage node is determined as the target storage node, it is also used to save the relevant information of the data blocks and / or parity blocks with storage failures sent by the client, and report the relevant information of the data blocks and / or parity blocks with storage failures to the management node. The storage node can be a server, a PC, a laptop, etc. The storage node can be a physical storage node or a logical storage node divided from a physical storage node.
[0073] The management node can be used to receive the relevant information of the data blocks and / or parity blocks with storage failures sent by the client and / or the target storage node, and recover each data block and / or parity block with storage failures according to the relevant information. The management node can be a server, a PC, a laptop, etc. The management node can also be one or more program modules on a device, or a virtual machine or container running on a device. The management node can also be a cluster composed of multiple devices, for example, it can be a general term for multiple program modules distributed on multiple devices.
[0074] Figure 2 Fig. shows a schematic block diagram of a structure of a computer device 100 provided by an embodiment of the present invention. The computer device 100 can be Figure 1 the client in, or the storage node, or the management node. The computer device 100 may include a memory 110 and a processor 120.
[0075] Among them, the processor 120 can be a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or an integrated circuit for controlling the execution of one or more programs for the data recovery method applied to the client, or the data recovery method applied to the target storage node, or the data recovery method applied to the management node provided by the following method embodiments.
[0076] The memory 110 can be a ROM or other type of static storage device that can store static information and instructions, a RAM or other type of dynamic storage device that can store information and instructions, or an Electrically Erasable Programmable Read-Only Memory (EEPROM), a Compact Disc Read-Only Memory (CD-ROM), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 110 can exist independently and be connected to the processor 120 through a communication bus. The memory 110 can also be integrated with the processor 120. Among them, the memory 110 is used to store machine-executable instructions for implementing the solution of this application. The processor 120 is used to execute the machine-executable instructions stored in the memory 110 to implement the embodiments of the data recovery method applied to the client, or the data recovery method applied to the target storage node, or the data recovery method applied to the management node.
[0077] Since the computer device 100 provided by the embodiments of the present invention is another implementation form of the data recovery method applied to the client, or the data recovery method applied to the target storage node, or the data recovery method applied to the management node provided by the following method embodiments, the technical effects that can be obtained thereby can refer to the following method embodiments and will not be elaborated herein.
[0078] Please refer to Figure 3 , Figure 3 which is a schematic flowchart of a data recovery method provided by an embodiment of the present invention, and its execution subject is the Figure 1 client in, and the method includes steps S101, S102, and S103.
[0079] S101, receiving response messages returned by each storage node for their respective write data requests.
[0080] Among them, each response message is used to indicate whether each storage node has successfully written the data block in its respective write data request. The write data request is sent by the client to the storage node, and the write data requests sent by the client to each storage node are all different.
[0081] Each storage node needs to write the data blocks in the write data requests it receives respectively. And regardless of whether the data blocks are successfully written or not, the client can receive corresponding response messages. In one case, the storage node returns a response message indicating write failure or write success to the client. In another case, the storage node fails or the network fails, resulting in the storage node not returning a response message within a preset duration. At this time, the client will receive a timeout response message, and the client determines that the write of this storage node fails based on the timeout response message. Thus, the client can determine the number of successfully written data blocks and the number of failed written data blocks according to the number of response messages indicating write success and the number of response messages indicating write failure.
[0082] S102, if there is a response message indicating write failure and the number of response messages indicating write failure is less than a preset value, send the data information to be recovered corresponding to each failed written data block to the management node, so that the management node records each piece of data information to be recovered.
[0083] Among them, the number of failed written data blocks is the same as the number of response messages indicating write failure. The preset value refers to the upper limit value of the number of failed written data blocks. When the number of failed written data blocks is greater than the preset value, it means that the failed written data blocks cannot be recovered. The reason is that the distributed storage system usually uses the erasure code (hereinafter referred to as: EC) technology to implement data storage, that is, the data to be stored is divided into n data blocks, and then the n data blocks are subjected to erasure coding to obtain m check blocks for data recovery processing. The size of the preset value is generally the same as the number of check blocks m. When the total number of failed written data blocks does not exceed m, the failed written data can be recovered by performing EC inverse coding on the successfully written data blocks and the check blocks. On the contrary, when the total number of failed written data blocks is greater than m, the failed written data cannot be recovered by the successfully written data blocks and the check blocks.
[0084] The data information to be recovered corresponding to each failed written data block includes the data block index and the write location. The write location includes the address of the storage node where the write is performed and the number of the corresponding disk in the storage node. If the client determines that the total number of failed written data blocks is less than the preset value according to each received response message, it is necessary to send the data information to be recovered corresponding to each failed written data block to the management node for recording, so that the management node can recover each failed written data block according to the recorded data information to be recovered.
[0085] S103. If the management node fails to record, each piece of data information to be restored is saved to a target storage node among multiple storage nodes, so that the target storage node reports each piece of data information to be restored to the management node, and after the reporting is successful, the management node restores the data block with a write failure corresponding to each piece of data information to be restored.
[0086] Among them, the target storage node is determined from multiple storage nodes by the client when it determines that the management node fails to successfully record the data information to be restored, according to the erasure code configuration of the distributed storage system and the storage space usage of each storage node currently in the online state. The client can set a response time interval and determine whether the management node successfully records the data information to be restored based on whether it receives a message from the management node regarding the recording status of the data information to be restored within the response time interval. If there are network fluctuations or the management node itself fails and needs to perform primary and standby switching, the management node may not be able to successfully record the data information to be restored.
[0087] When the client determines that the management node fails to record, it sends each piece of data information to be restored to the target storage node for storage. As an implementation method, the target storage node can store each piece of data information to be restored in the local leveldb database, and the format can be as follows in the table:
[0088] Structure key KSBLK@Block name Key value status-disk_id-dn_ip
[0089] In the above table, KSBLK@block name is the index of the data block with a write failure, status indicates whether this piece of data information to be restored has been reported, dn_ip is the address of the storage node with a write failure, and disk_id is the number of the disk with a write failure. As another implementation method, the target storage node can also store each piece of data information to be restored in a specific data block server.
[0090] When the network returns to normal or the primary and standby switching of the management node is completed, the target storage node then sends each piece of data information to be restored stored locally to the management node for recording, and then the management node restores each corresponding data block with a write failure according to each piece of recorded data information to be restored.
[0091] The method provided in the above embodiments has the beneficial effect that when the management node fails to successfully record the data information to be recovered corresponding to the data block with a write failure, the client saves the data information to be recovered corresponding to the data block with a write failure to the target storage node among multiple storage nodes, so that the target storage node subsequently reports the data information to be recovered corresponding to the data block with a write failure to the management node, enabling the management node to recover each data block with a write failure corresponding to each piece of data information to be recovered, avoiding the inability to accurately and completely record the data information to be recovered corresponding to the data block with a write failure due to an exception of the management node or a network exception, enabling each data block with a write failure to be recovered, and thus ensuring the consistency of the data before and after storage.
[0092] Based on this, an embodiment of the present invention further provides a possible implementation manner for determining the target storage node. Please refer to Figure 3 ..., Figure 4 which Figure 4 is a schematic flowchart of a method for determining a target storage node provided by an embodiment of the present invention. The method includes steps S110 and S111.
[0093] S110, calculate the candidate number of the target storage node according to the number of storage nodes, the number of first data blocks, and the number of second data blocks;
[0094] Among them, the data block includes a first data block and a second data block. The second data block is parity data obtained by performing erasure coding on the first data block, that is, a parity block, and the first data block is obtained by the client dividing the data to be stored.
[0095] The number of storage nodes, the number of first data blocks, the number of second data blocks, and the candidate number of the target storage node satisfy the following formula:
[0096]
[0097] In the formula, m is the number of first data blocks, n is the number of second data blocks, N is the number of storage nodes, and M is the candidate number of the target storage node.
[0098] For example, when N is 5, m is 2, and n is 4, the value of the candidate number M of the target storage node is calculated as 3 according to the above formula.
[0099] S111, use the candidate number of storage nodes with the smallest storage space utilization rate as the target storage node.
[0100] Optionally, sort all storage nodes in ascending order of storage space utilization rate, and starting from the beginning of the sequence, use several storage nodes that meet the candidate number as the target storage node.
[0101] For example, the storage space utilization rates of storage node 1, storage node 2, storage node 3, storage node 4, and storage node 5 are 30%, 20%, 32%, 35%, and 15% respectively. When sorted in ascending order of storage space utilization rate, the result is storage node 5, storage node 2, storage node 3, storage node 1, and storage node 4. If the number of candidates to be selected is 3, then the 3 storage nodes with the lowest storage space utilization rates, namely storage node 5, storage node 2, and storage node 3, are used as the target storage nodes.
[0102] It should be noted that the storage nodes in step S201 and step S202 are in an online state, that is, they can communicate normally with the client and the management node.
[0103] Please refer to Figure 5 , Figure 5 Another process schematic diagram of the data recovery method provided by the embodiment of the present invention, the execution subject of which is Figure 1 the target storage nodes among multiple storage nodes in
[0104] S201, receiving the data information to be recovered corresponding to each write-failed data block sent by the client.
[0105] Among them, the data information to be recovered is sent by the client to the target storage node when the management node fails to record the data information to be recovered. The number of write-failed data blocks is the same as the number of response messages indicating write failure. Each response message is generated by each storage node for its own write data request and returned to the client.
[0106] S202, saving each piece of data information to be recovered for reporting to the management node, and after successful reporting, enabling the management node to recover the write-failed data blocks corresponding to each piece of data information to be recovered.
[0107] Among them, after the target storage node completes the local storage of the data information to be recovered, it will select a reporting method according to whether the management node is in an online state, so that the management node can recover each write-failed data block corresponding to each piece of recorded data information to be recovered.
[0108] When the management node is in an online state, in order to avoid excessive job pressure on the management node due to a large number of reported data information to be recovered, which affects the performance of its own business. For this reason, the embodiment of the present invention provides a possible implementation method for reporting the data information to be recovered. Please refer to Figure 6 , Figure 6 A process schematic diagram of a method for reporting data information to be recovered provided by the embodiment of the present invention. This method includes step S210 and step S211.
[0109] S210, receive the reporting quantity sent by the management node.
[0110] Among them, the target storage node receives the reporting quantity in the response message returned by the management node, that is, the maximum quantity of data information to be restored that the target storage node can send to the management node at the current moment.
[0111] For example, if the reporting quantity is 40, it means that the target storage node can send at most 40 pieces of data information to be restored to the management node at the current moment.
[0112] S211, send the earliest-reported quantity of data information to be restored to the management node.
[0113] Among them, the target storage node records the time when each piece of data to be restored is saved to the local leveldb database, that is, the save time.
[0114] Optionally, the target storage node sorts each piece of data information to be restored in chronological order of the recorded time, and starts from the beginning of the sequence to send several pieces of data information to be restored that meet the reporting quantity to the management node.
[0115] For example, if the reporting quantity is 40, if the length of the sequence is greater than 40, the first 40 pieces of data information to be restored in the sequence will be sent to the management node; if the length of the sequence is not greater than 40, all the data information to be restored in the sequence will be sent to the management node.
[0116] After the target storage node completes the reporting at the current moment, it sends a request message to the management node again to determine the reporting quantity of data information to be restored at the next moment.
[0117] It should be noted that, in order to further reduce the job pressure on the management node, when the target storage node reports data information to be restored, it will send the data information to be restored with the reporting status of "not reported" to the management node according to the reporting status of each piece of data information to be restored, and the data information to be restored with the reporting status of "reported" will not be reported repeatedly.
[0118] When the management node is in the offline state, it means that the target storage node cannot communicate with the management node. At this time, the target storage node needs to wait to communicate with the management node again. And when the management node drops the line, it may be in the process of restoring some data blocks with write failures corresponding to the data information to be restored, and the "offline event" will abnormally terminate the recovery procedure of these data blocks with write failures. In this regard, the embodiments of the present invention provide the following another possible implementation manner for reporting data information to be restored:
[0119] When the management node resumes from the offline state to the online state, all data information to be restored will be sent to the management node.
[0120] Among them, when the management node returns to the online state again, the target storage node reports all the reported and unreported data information to be restored stored locally to the management node at the current moment, so that the management node can find out the data blocks whose restoration process is abnormally terminated, thus ensuring data consistency.
[0121] Please refer to Figure 7 , Figure 7 which is another schematic flowchart of the data restoration method provided by the embodiment of the present invention, and its execution subject is the Figure 1 management node in , and this method includes steps S301 and S302.
[0122] S301: Receive the data information to be restored corresponding to each data block with a write failure sent by the target storage node among multiple storage nodes.
[0123] Among them, the data information to be restored sent by the target storage node is saved to the target storage node when the management node fails to record the data information to be restored, and the number of data blocks with a write failure is the same as the number of response messages indicating a write failure. Each response message is generated by each storage node for its own write data request and returned to the client.
[0124] S302: Restore the corresponding data blocks with a write failure according to each piece of data information to be restored.
[0125] Among them, the management node generates a restoration data list for each piece of data information to be restored in the order of recording time. According to the restoration data list, for each piece of data information to be restored, the management node restores each data block with a write failure in turn according to the index of the data block with a write failure, the address of the storage node with a write failure, and the number of the disk with a write failure included. Moreover, after the management node restores each data block with a write failure, it sends a bad block deletion message to the storage node with a write failure to make it delete the incomplete data block written before.
[0126] To avoid excessive job pressure on the management node during data restoration, which may cause too much impact on the performance of its own services, the embodiment of the present invention limits the number of data information to be restored sent by the target storage node at one time. Please refer to Figure 8 , Figure 8 which is a schematic flowchart of a method for determining the number of data information to be restored reported by the embodiment of the present invention. This method includes steps S310, S311, and step S312.
[0127] S310: Obtain the CPU occupancy rate, CPU load, and the number of storage nodes.
[0128] Among them, the CPU occupancy rate and load of the management node can reflect its job pressure and resource usage. The number of storage nodes reflects the maximum number of target storage nodes that can send data information to be restored to the management node at the current moment.
[0129] For example, if the management node obtains that the number of storage nodes communicating with it at the current moment is 5, it means that at most 5 storage nodes are used as target storage nodes to send data information to be restored to the management node at the current moment.
[0130] S311. Calculate the reporting quantity according to the total quantity of data information to be restored, the CPU occupancy rate, the CPU load, and the number of storage nodes.
[0131] Among them, the reporting quantity is used to represent the maximum quantity of data information to be restored that the target storage node can report. The total quantity of data information to be restored refers to the quantity of data information to be restored that has been recorded by the management node and added to the recovery data list at the current moment.
[0132] At the current moment, the total quantity of data information to be restored, the CPU occupancy rate, the number of storage nodes, and the reporting quantity satisfy the following formula:
[0133]
[0134] In the formula, q is the total quantity of data information to be restored at the current moment, q max is the maximum value of the total quantity of data information to be restored preset, u is the actual occupancy rate of the CPU at the current moment, u max is the maximum occupancy rate of the CPU preset, N is the number of storage nodes in the online state at the current moment, K is the reporting quantity, and p is the preset weight coefficient.
[0135] Since the number of CPUs in the management node may be multiple, the above formula is corrected by using the average load of the CPU and the number of CPUs at the current moment, and the following formula for calculating the reporting quantity is obtained:
[0136]
[0137] In the formula, a is the average load of the CPU, and n is the number of CPUs.
[0138] For example, when p is 0.75, u max is 50, q max is 1024, n is 12, a is 12, u is 30, q is 1, and N is 5, the reporting quantity K obtained according to the above formula is 40.
[0139] S312, sending the reported quantity to each storage node, so that each storage node, when serving as a target storage node, sends the locally stored information of the data to be recovered to the management node according to the reported quantity.
[0140] Among them, the management node can periodically calculate the reporting quantity and send it to each storage node, so that when the storage node acts as a target storage node, it reports and processes the data information to be recovered stored in the local leveldb database according to the reporting quantity.
[0141] Optionally, after receiving the request message sent by the target storage node, the management node may also write the reported quantity calculated in the current period into a response message and send it to the target storage node.
[0142] In order to execute the corresponding steps in the above embodiment and each possible implementation method, the following provides a data recovery device 200 applied to a client, a data recovery device 300 applied to a target storage node, and a data recovery device 400 applied to a management node. Figure 9 , Figure 10 and Figure 11 , Figure 9 A block diagram of a data recovery device 200 applied to a client provided by an embodiment of the present invention is shown. Figure 10 A block diagram of a data recovery device 300 applied to a target storage node provided by an embodiment of the present invention is shown. Figure 11 A block diagram of a data recovery device 400 applied to a management node provided in an embodiment of the present invention is shown. It should be noted that the basic principles and technical effects of the data recovery device 200 applied to a client, the data recovery device 300 applied to a target storage node, and the data recovery device 400 applied to a management node provided in an embodiment of the present invention are the same as those in the above embodiments, and are not mentioned in the embodiment of the present invention for the sake of brief description.
[0143] The data recovery device 200 applied to the client includes a first receiving module 201 , a first sending module 202 and a first determining module 203 .
[0144] The first receiving module 201 is used to receive a response message returned by each storage node in response to its respective write data request, wherein each response message is used to indicate whether each storage node has successfully written a data block in its respective write data request.
[0145] The first sending module 202 is used to send the data information to be recovered corresponding to each data block that failed to be written to the management node if there is a response message indicating a write failure and the number of response messages indicating a write failure is less than a preset value, so that the management node records each piece of data information to be recovered.
[0146] The first sending module 202 is further configured to, if the management node fails to record, save each piece of data information to be restored to a target storage node among multiple storage nodes, so that the target storage node reports each piece of data information to be restored to the management node, and after the reporting is successful, the management node restores the data block corresponding to each piece of data information to be restored that has failed to be written.
[0147] The first determining module 203 is configured to calculate the candidate quantity of the target storage node according to the quantity of storage nodes, the quantity of first data blocks, and the quantity of second data blocks, where the data blocks include first data blocks and second data blocks, and the second data blocks are parity data obtained by performing erasure coding on the first data blocks; and use the candidate quantity of storage nodes with the lowest storage space utilization rate as the target storage node.
[0148] The data recovery device 300 applied to the target storage node includes a second receiving module 301, a storage module 302, and a second sending module 303.
[0149] The second receiving module 301 is configured to receive the data information to be restored corresponding to each data block that has failed to be written sent by the client. The data information to be restored is sent by the client to the target storage node when the management node fails to record the data information to be restored. The quantity of data blocks that have failed to be written is the same as the quantity of response messages indicating the failure to write. Each response message is generated by each storage node for its own write data request and returned to the client.
[0150] The storage module 302 is configured to save each piece of data information to be restored for reporting to the management node, and after the reporting is successful, the management node restores the data block corresponding to each piece of data information to be restored that has failed to be written.
[0151] When the management node is in the online state, the second receiving module 301 is further configured to receive the reporting quantity sent by the management node. The second sending module 303 is configured to send the earliest saved reporting quantity of data information to be restored to the management node, where the target storage node records the saving time of each piece of data information to be restored.
[0152] When the management node is in the offline state, the second sending module 303 is configured to send all the data information to be restored to the management node when the management node resumes from the offline state to the online state.
[0153] The data recovery device 400 applied to the management node includes a third receiving module 401, a recovery module 402, a second determining module 403, and a third sending module 404.
[0154] A third receiving module 401 is configured to receive the to-be-restored data information corresponding to each data block with a write failure sent by a target storage node among multiple storage nodes. The to-be-restored data information sent by the target storage node is saved to the target storage node when the client fails to record the to-be-restored data information at the management node. The number of data blocks with write failures is the same as the number of response messages indicating write failures. Each response message is generated by each storage node for its own write data request and returned to the client.
[0155] A recovery module 402 is configured to recover the corresponding data blocks with write failures according to each piece of to-be-restored data information.
[0156] A second determination module 403 is configured to obtain the CPU occupancy rate, the CPU load, and the number of storage nodes; calculate a reporting quantity according to the total quantity of the to-be-restored data information, the CPU occupancy rate, the CPU load, and the number of storage nodes. The reporting quantity is used to represent the maximum quantity of the to-be-restored data information that the target storage node can report.
[0157] A third sending module 404 is configured to send the reporting quantity to each storage node, so that when each storage node serves as the target storage node, it sends the locally saved to-be-restored data information to the management node according to the reporting quantity.
[0158] An embodiment of the present invention further provides a readable storage medium including computer-executable instructions. When the computer-executable instructions are executed, they can be used to perform the relevant operations in the data recovery method applied to the client, or the data recovery method applied to the target storage node, or the data recovery method applied to the management node provided by the foregoing method embodiment.
[0159] In summary, in an embodiment of the present invention, a data recovery method and related device, a client receives response messages returned by each storage node for their respective write data requests, where each response message is used to indicate whether each storage node has successfully written the data block in its respective write data request; if there is a response message indicating a write failure and the number of response messages indicating a write failure is less than a preset value, the client sends the data information to be recovered corresponding to each data block with a write failure to the management node, so that the management node records each data information to be recovered; if the management node fails to record, the client saves each data information to be recovered to a target storage node among multiple storage nodes, so that the target storage node reports each data information to be recovered to the management node, and after the reporting is successful, the management node recovers the data block corresponding to each data information to be recovered with a write failure. Since the client saves the data information to be recovered corresponding to the data block with a write failure to a target storage node among multiple storage nodes when the management node fails to successfully record the data information to be recovered corresponding to the data block with a write failure, so that the target storage node subsequently reports the data information to be recovered corresponding to the data block with a write failure to the management node, so that the management node recovers each data block with a write failure corresponding to each data information to be recovered, it avoids the inability to accurately and completely record the data information to be recovered corresponding to the data block with a write failure due to an exception of the management node or a network exception, enables each data block with a write failure to be recovered, and thus ensures the consistency of the data before and after storage.
[0160] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A data recovery method, characterized in that, it is applied to a client in a distributed storage system, the distributed storage system further includes a management node and a plurality of storage nodes, the client is communicatively connected to the management node, the client is communicatively connected to the plurality of storage nodes, the management node is communicatively connected to the plurality of storage nodes, the client sends a write data request to the plurality of storage nodes, wherein each of the storage nodes needs to write the data blocks in the received write data request, and the method includes: Receiving a response message returned by each of the storage nodes for its respective write data request, wherein each of the response messages is used to indicate whether each of the storage nodes has successfully written the data blocks in its respective write data request; If there is a response message indicating a write failure, and the number of the response messages indicating a write failure is less than a preset value, sending, to the management node, the data information to be recovered corresponding to each write-failed data block, so that the management node records each piece of the data information to be recovered; the data information to be recovered includes a data block index and a write location, and the write location includes the address of the storage node where the data is written and the number of the corresponding disk in the storage node; If the management node fails to record, saving each piece of the data information to be recovered to a target storage node among the plurality of storage nodes, so that the target storage node reports each piece of the data information to be recovered to the management node, and after the reporting is successful, the management node recovers the corresponding write-failed data block according to each piece of the data information to be recovered.
2. The method according to claim 1, characterized in that, the data blocks include a first data block and a second data block, the second data block is parity data obtained by erasure coding of the first data block, and before the step of saving each piece of the data information to be recovered to a target storage node among the plurality of storage nodes, the method further includes: Calculating the number of candidate target storage nodes according to the number of the storage nodes, the number of the first data blocks, and the number of the second data blocks; Taking the candidate target storage nodes with the smallest storage space utilization rate among the calculated number of candidate target storage nodes as the target storage nodes.
3. A data recovery method, characterized in that, it is applied to a target storage node among a plurality of storage nodes in a distributed storage system, the distributed storage system further includes a client and a management node, the client is communicatively connected to the management node, the client is communicatively connected to the plurality of storage nodes, the management node is communicatively connected to the plurality of storage nodes, the client sends a write data request to the plurality of storage nodes, wherein each of the storage nodes needs to write the data blocks in the received write data request, and the method includes: Receiving the to-be-recovered data information corresponding to each data block with a write failure sent by the client, where the to-be-recovered data information is sent by the client to the target storage node when the management node fails to record the to-be-recovered data information. The number of data blocks with write failures is the same as the number of response messages indicating write failures. Each of the response messages is generated by each storage node for its respective write data request and returned to the client; the to-be-recovered data information includes a data block index and a write location, and the write location includes the address of the storage node where the write is performed and the number of the corresponding disk in the storage node. Saving each piece of the to-be-recovered data information for reporting to the management node, and after successful reporting, enabling the management node to recover the corresponding data block with a write failure according to each piece of the to-be-recovered data information.
4. The method according to claim 3, characterized in that the management node is in an online state, and the target storage node records the save time of each piece of the to-be-recovered data information. The method further includes: Receiving the reporting quantity sent by the management node; Sending the earliest-reported quantity of pieces of the to-be-recovered data information to the management node.
5. The method according to claim 3, characterized in that the management node is in an offline state. The method further includes: When the management node recovers from the offline state to the online state, sending all the to-be-recovered data information to the management node.
6. A data recovery method, characterized in that applied to a management node in a distributed storage system. The distributed storage system further includes a client and multiple storage nodes. The client is communicatively connected to the management node, the client is communicatively connected to the multiple storage nodes, the management node is communicatively connected to the multiple storage nodes. The client sends a write data request to the multiple storage nodes. Among them, each storage node needs to write the data block in the write data request it receives. The method includes: Receiving the to-be-recovered data information corresponding to each data block with a write failure sent by the target storage node among the multiple storage nodes. The to-be-recovered data information sent by the target storage node is saved by the client to the target storage node when the management node fails to record the to-be-recovered data information. The number of data blocks with write failures is the same as the number of response messages indicating write failures. Each of the response messages is generated by each storage node for its respective write data request and returned to the client; the to-be-recovered data information includes a data block index and a write location, and the write location includes the address of the storage node where the write is performed and the number of the corresponding disk in the storage node. Recovering the corresponding data block with a write failure according to each piece of the to-be-recovered data information.
7. The method according to claim 6, characterized in that the method further includes: Obtaining the CPU occupancy rate, the CPU load, and the number of the storage nodes; Calculate the reporting quantity according to the total quantity of the data information to be restored, the CPU occupancy rate, the CPU load, and the number of the storage nodes, where the reporting quantity is used to represent the maximum quantity of the data information to be restored that the target storage node can report; Send the reporting quantity to each of the storage nodes, so that when each of the storage nodes serves as a target storage node, it sends the data information to be restored locally saved to the management node according to the reporting quantity.
8. A data recovery device Characterized in that It is applied to a client in a distributed storage system, the distributed storage system further includes a management node and a plurality of storage nodes, the client is communicatively connected to the management node, the client is communicatively connected to the plurality of storage nodes, the management node is communicatively connected to the plurality of storage nodes, and the client sends a write data request to the plurality of storage nodes, wherein each of the storage nodes needs to write the data blocks in the write data request it receives, and the device includes: A first receiving module, configured to receive a response message returned by each of the storage nodes for its respective write data request, where each of the response messages is used to represent whether each of the storage nodes has successfully written the data blocks in its respective write data request; A first sending module, configured to, if there is a response message indicating a write failure and the number of the response messages indicating a write failure is less than a preset value, send the data information to be restored corresponding to each write-failed data block to the management node, so that the management node records each of the data information to be restored; the data information to be restored includes a data block index and a write location, and the write location includes the address of the storage node where the data is written and the number of the corresponding disk in the storage node; The first sending module is further configured to, if the management node records a failure, save each of the data information to be restored to a target storage node among the plurality of storage nodes, so that the target storage node reports each of the data information to be restored to the management node, and after the reporting is successful, the management node restores the write-failed data blocks corresponding to each of the data information to be restored.
9. A data recovery device Characterized in that It is applied to a target storage node among a plurality of storage nodes in a distributed storage system, the distributed storage system further includes a client and a management node, the client is communicatively connected to the management node, the client is communicatively connected to the plurality of storage nodes, the management node is communicatively connected to the plurality of storage nodes, and the client sends a write data request to the plurality of storage nodes, wherein each of the storage nodes needs to write the data blocks in the write data request it receives, and the device includes: A second receiving module, configured to receive the to-be-restored data information corresponding to each data block with a write failure sent by the client. The to-be-restored data information is sent by the client to the target storage node when the management node fails to record the to-be-restored data information. The number of data blocks with write failures is the same as the number of response messages indicating write failures. Each of the response messages is generated by each storage node for its respective write data request and returned to the client. The to-be-restored data information includes a data block index and a write location. The write location includes the address of the storage node where the data is written and the number of the corresponding disk in the storage node. A storage module, configured to save each piece of the to-be-restored data information for reporting to the management node, and after successful reporting, cause the management node to restore the data blocks with write failures corresponding to each piece of the to-be-restored data information.
10. A data recovery device Characterized in that It is applied to a management node in a distributed storage system. The distributed storage system further includes a client and multiple storage nodes. The client is communicatively connected to the management node. The client is communicatively connected to the multiple storage nodes. The management node is communicatively connected to the multiple storage nodes. The client sends a write data request to the multiple storage nodes. Among them, each storage node needs to write the data block in its respective received write data request. The device includes: A third receiving module, configured to receive the to-be-restored data information corresponding to each data block with a write failure sent by a target storage node among the multiple storage nodes. The to-be-restored data information sent by the target storage node is saved to the target storage node when the management node fails to record the to-be-restored data information. The number of data blocks with write failures is the same as the number of response messages indicating write failures. Each of the response messages is generated by each storage node for its respective write data request and returned to the client. The to-be-restored data information includes a data block index and a write location. The write location includes the address of the storage node where the data is written and the number of the corresponding disk in the storage node. A recovery module, configured to restore the data blocks with corresponding write failures according to each piece of the to-be-restored data information.
11. A distributed storage system Characterized in that The distributed storage system includes a client, a management node, and multiple storage nodes. The client is communicatively connected to the management node. The client is communicatively connected to the multiple storage nodes. The management node is communicatively connected to the multiple storage nodes. The client is configured to send a write data request to the multiple storage nodes. Each storage node is configured to receive the write data request sent by the client and write the data block in its respective write data request. Among them, the write data requests received by each storage node are different. The client is further configured to receive response messages returned by each of the storage nodes for their respective write data requests, where each response message is used to indicate whether each storage node has successfully written the data block in its respective write data request; The client is further configured to, if there is a response message indicating a write failure and the number of response messages indicating a write failure is less than a preset value, send the data information to be recovered corresponding to each data block with a write failure to the management node, so that the management node records each piece of data information to be recovered; the data information to be recovered includes a data block index and a write location, and the write location includes the address of the storage node where the write is performed and the number of the corresponding disk in the storage node; The client is further configured to, if the management node fails to record, save each piece of data information to be recovered to a target storage node among the multiple storage nodes; The target storage node is configured to report each piece of data information to be recovered to the management node; The management node is configured to, after the successful reporting by the target storage node, recover the corresponding data block with a write failure according to each piece of data information for recovery.
12. A computer device, comprising a memory and a processor, wherein, the memory stores a computer program, and when the processor executes the computer program, it implements the data recovery method according to any one of claims 1-2, or implements the data recovery method according to any one of claims 3-5, or implements the data recovery method according to any one of claims 6-7.
13. A computer-readable storage medium, on which a computer program is stored, wherein, when the computer program is executed by a processor, it implements the data recovery method according to any one of claims 1-2, or implements the data recovery method according to any one of claims 3-5, or implements the data recovery method according to any one of claims 6-7.
Citation Information
Patent Citations
Method, apparatus and system for data reconstruction in distributed storage system
CN106662983A