Strip repairing method, system and equipment
By sensing bandwidth differences between racks, optimizing repair plans, and aggregating helper blocks within the same rack, the system solves the problems of excessive network traffic and long cross-rack transmission time in multi-band repair, improves repair performance and efficiency, and ensures data reliability and fault tolerance.
Patent Information
- Application Number
- CN202510753923.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-09-19
AI Technical Summary
Existing technologies ignore cross-rack bandwidth differences and multi-rack failure scenarios during multi-strip repair, resulting in excessive network traffic and long cross-rack transmission time, resulting in poor repair performance and efficiency.
By sensing the bandwidth differences between racks, an iterative algorithm is used to optimize the repair solution, aggregate helper blocks within the same rack, reduce cross-rack transmission time, reasonably distribute the data transmission burden, and use high-bandwidth links to reduce low-bandwidth bottlenecks.
It effectively reduces network traffic and cross-rack transmission time, improves multi-strip repair performance and efficiency, and enhances the durability and fault tolerance of repair results.
Smart Images

Figure CN120675575A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of computer storage technology, and more specifically, relates to a stripe repair method, system, and device. Background Art
[0002] When data failures occur, erasure codes require rapid data repair to ensure data availability, a process that relies on network transmission of large amounts of data. Existing repair methods primarily design scheduling algorithms for individual stripes to improve repair performance. However, due to full node failures and lazy repair strategies, storage systems must handle scenarios where multiple stripes fail simultaneously. Furthermore, current storage systems often utilize cross-rack deployments, resulting in varying network bandwidth between racks.
[0003] Therefore, ignoring cross-rack bandwidth differences and multi-strip failure scenarios will lead to uneven transmission loads between racks, with the rack with the heaviest transmission load becoming a performance bottleneck during the repair process. Furthermore, ignoring transmission scheduling for cross-rack repair links will further prolong the repair process, resulting in reduced repair performance and efficiency.
[0004] In summary, how to reduce the network traffic and cross-rack transmission time during stripe data repair, and thereby improve the multi-stripe repair performance and efficiency, is a technical problem that urgently needs to be solved. Summary of the Invention
[0005] In response to the defects of the existing technology, the purpose of this application is to provide a stripe repair method, system and equipment, aiming to solve the problems in the existing technology of multi-strip repair caused by excessive network traffic and long cross-rack transmission time, resulting in low multi-strip repair performance and repair efficiency.
[0006] To achieve the above objectives, in a first aspect, the present application provides a stripe repair method, comprising: Randomly initialize each failed stripe to be repaired to obtain the initial repair plan for each failed stripe; Determining bandwidth differences between different racks, and optimizing an initial repair solution through an iterative algorithm based on the bandwidth differences and an optimization goal to obtain an optimized repair solution; the optimization goal is to reduce cross-domain transmission time; According to the optimized repair solution, helper blocks in the same rack are aggregated and encoded to repair failed blocks in each stripe.
[0007] Optionally, the method for generating the initial repair solution includes: Determine the remaining blocks according to the data blocks, check blocks and failed blocks of the failed stripe; K blocks are randomly selected from the remaining blocks of the failed stripe as helper blocks to generate multiple different initial repair schemes, and storage nodes are selected to store the repaired request blocks; wherein k is a positive integer.
[0008] Optionally, the storage node does not store any other blocks in the failed stripe except the failed block; the total number of blocks storing the failed stripe in the area to which the requested block belongs does not exceed m to achieve regional fault tolerance; where m is a positive integer.
[0009] Optionally, the method for obtaining the optimized repair solution includes: Randomly selecting a benchmark repair scheme from multiple initial repair schemes for the current failed stripe, and calculating the time required for cross-rack data transmission in the benchmark repair scheme as the benchmark cross-domain transmission time; Traversing all valid initial repair solutions for the currently failed stripe, and determining a target cross-domain transmission time for each initial repair solution; The target cross-domain transmission time is compared with the benchmark cross-domain transmission time, and any initial repair solution is selected to replace the benchmark repair solution according to the comparison result to obtain an optimized repair solution.
[0010] Optionally, it also includes: Get the current network bandwidth of each rack, and get the bandwidth difference based on the network bandwidth of each rack; According to the bandwidth difference, the transmission task volume of the high-bandwidth rack is increased and the transmission task volume of the low-bandwidth rack is reduced, so as to reduce the cross-domain transmission time.
[0011] Optionally, comparing the target inter-domain transmission time with the benchmark inter-domain transmission time, and selecting any initial repair solution to replace the benchmark repair solution according to the comparison result to obtain an optimized repair solution, includes: Determining whether the target inter-domain transmission time is less than the benchmark inter-domain transmission time; If the target cross-domain transmission time is less than the benchmark cross-domain transmission time, replacing the initial repair solution with the repair solution corresponding to the target cross-domain transmission time as the optimized repair solution; If the target cross-domain transmission time is not less than the benchmark cross-domain transmission time, all valid initial repair solutions and the transmission time comparison process are repeated until all alternative initial repair solutions are traversed or the target cross-domain transmission time is less than the benchmark cross-domain transmission time, and the optimized repair solution is obtained.
[0012] Optionally, the aggregating and encoding helper blocks in the same rack according to the optimized repair solution to repair failed blocks in each stripe includes: Determine the rack where the helper block of each failed stripe is located, and aggregate all the helper blocks in the rack; Using the linear encoding rule of erasure codes, the helper blocks aggregated in the rack are encoded to generate temporary blocks. The temporary blocks contain partial information required to reconstruct the failed blocks. Sending the temporary block generated by each rack to the target rack where the requested block is located; The request block generates a new block from the received temporary block using the linear coding rule of the erasure code, and stores the new block in the storage node where the corresponding request block is located to repair the failed block in the failed stripe.
[0013] In a second aspect, the present application further provides a strip repair system, comprising: An initialization module is used to randomly initialize each failed stripe to be repaired and obtain an initial repair plan for each failed stripe; An optimization module is configured to determine bandwidth differences between different racks and optimize an initial repair solution using an iterative algorithm based on the bandwidth differences and an optimization goal to obtain an optimized repair solution; the optimization goal is to reduce cross-domain transmission time; The repair module is used for aggregating and encoding the helper blocks in the same rack according to the optimized repair solution to repair the failed blocks in each stripe.
[0014] In a third aspect, the present application provides an electronic device comprising: at least one memory for storing programs; and at least one processor for executing the programs stored in the memory. When the programs stored in the memory are executed, the processor is used to execute the method described in the first aspect or any possible implementation of the first aspect.
[0015] In a fourth aspect, the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a processor, the processor executes the method described in the first aspect or any possible implementation of the first aspect.
[0016] In a fifth aspect, the present application provides a computer program product, which, when executed on a processor, enables the processor to execute the method described in the first aspect or any possible implementation of the first aspect.
[0017] It can be understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here.
[0018] In general, the above technical solutions conceived by this application have the following beneficial effects compared with the existing technologies: (1) This application can sense the bandwidth differences between racks and iteratively generate repair schemes with lower cross-rack transmission time, thereby reducing the transmission time of the most loaded rack and thus reducing the repair time. Each scheme determines the helper blocks used to assist in repair and the target blocks for storing failed data in the failed stripe. By aggregating and encoding the helper blocks within the same rack, cross-rack network traffic is reduced, thus solving the problems of excessive network traffic and long cross-rack transmission time caused by multi-strip repair, thereby improving the performance and efficiency of multi-strip repair.
[0019] (2) The storage nodes of this application cannot store any other data blocks or check blocks in the original failed stripe except the failed block. This restriction effectively avoids the risk of multiple related data being lost due to a single node failure and enhances the durability of the repair results.
[0020] (3) This application tends to increase the transmission task volume of nodes located on high-bandwidth racks, while correspondingly reducing the transmission task volume of low-bandwidth racks. The load adjustment strategy based on bandwidth differences can more reasonably distribute the data transmission burden across racks and give priority to using links with better network resources. In this way, during the iterative optimization process, the overall cross-domain transmission time can be more effectively reduced, and the repair efficiency can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 This is one of the flow charts of the strip repair method provided in the embodiment of the present application; Figure 2 This is the second flow chart of the strip repair method provided in the embodiment of the present application; Figure 2 (a) shows the schematic diagram of block layout and time cost; Figure 2 (b) in the figure shows the repair scheme before optimization; Figure 2 (c) in the figure shows the schematic diagram of the optimized repair scheme; Figure 3 It is a structural schematic diagram of a strip repair method and device provided in an embodiment of the present application; Figure 4 It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0022] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0023] The term "and / or" as used herein describes an association between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. The symbol " / " as used herein indicates that the related objects are in an "or" relationship, for example, A / B means either A or B.
[0024] The terms "first" and "second" in this specification and claims are used to distinguish different objects rather than to describe a specific order of objects. For example, "first response message" and "second response message" are used to distinguish different response messages rather than to describe a specific order of response messages.
[0025] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0026] In the description of the embodiments of the present application, unless otherwise specified, "multiple" means two or more, for example, multiple processing units means two or more processing units, etc.; multiple elements means two or more elements, etc.
[0027] First, the technical terms involved in the embodiments of this application are introduced.
[0028] (1) Erasure Code Stripe: Through erasure coding technology, the original data can be divided into several data blocks, and then several check blocks are generated by performing specific matrix operations on these data blocks. These data blocks and the check blocks form an erasure code stripe, which can also be described as an erasure code stripe.
[0029] For example, in an erasure code based In the storage system, a file is divided into The original data blocks are encoded into Total data blocks, of which, ,in, is the number of check blocks, The collection of total data blocks is called an "erasure code stripe". An erasure code stripe is the smallest coding unit in the erasure code. Usually, each data block in an erasure code stripe is stored in a different storage node. The system can tolerate the failure of any m nodes. If no more than If a node fails, the remaining Repair the invalid data in each node.
[0030] (2) Frame: A rack is the physical support structure for storage devices (such as hard drives and servers) and is typically used in data centers or large-scale storage systems. In a distributed storage system based on erasure coding, multiple racks can form a cluster, with each rack containing multiple storage nodes (such as servers).
[0031] The embodiments of the present application are described below in conjunction with the drawings in the embodiments of the present application.
[0032] Reference Figure 1 , the present application provides a stripe repair method, comprising: S101. Randomly initialize each failed stripe to be repaired to obtain an initial repair plan for each failed stripe; S102. Determine the bandwidth differences between different racks, and optimize the initial repair solution through an iterative algorithm based on the bandwidth differences and the optimization goal to obtain an optimized repair solution; the optimization goal is to reduce cross-domain transmission time; S103. According to the optimized repair solution, aggregate and encode the helper blocks in the same rack to repair the failed blocks in each stripe.
[0033] Specifically, the repair process begins when one or more data blocks and parity blocks (collectively referred to as "failed blocks") in a stripe are detected to have failed. For each such failed stripe, a preliminary repair plan is generated that describes how to reconstruct the failed blocks using the data blocks and parity blocks (remaining blocks) that are still available in storage.
[0034] Random initialization is the core method for generating these initial solutions: First, all remaining blocks that the failed stripe depends on are identified. Then, a certain number of blocks (for example, k blocks may be required, depending on the encoding rules of the erasure code) are randomly selected from these remaining blocks as "helper blocks." Each different random selection combination, combined with the target storage node (called the "requesting node") selected to repair the failed block, forms a different initial repair solution. This process may be repeated multiple times to generate multiple different initial solutions, providing a diverse basis for subsequent optimization steps and avoiding the limitations that may arise from starting from a single fixed starting point.
[0035] After generating multiple initial repair plans, the optimal one needs to be selected. The key to optimization is reducing the cross-rack (cross-domain) data transmission time required during the repair process. First, bandwidth information for network connections between different rack pairs in the current storage cluster needs to be collected and analyzed to calculate the relative bandwidth differences between them.
[0036] An iterative algorithm (such as a greedy algorithm) is used to evaluate and improve the initial solutions. Based on known bandwidth differences, the algorithm calculates the total time required to transmit all helper block data to the requesting node for each initial repair solution, specifically focusing on the time required to transmit data across different racks. The algorithm compares the total cross-domain transmission time of different solutions, favoring those that utilize high-bandwidth links and avoid low-bandwidth bottlenecks. Through iterative comparison and evaluation, the algorithm eventually converges to an optimal repair solution with the lowest estimated cross-domain transmission time among all evaluated solutions.
[0037] After the final optimized repair plan is determined, the actual repair execution phase begins. The plan clearly specifies which helper blocks will be used to rebuild the failed block and the storage locations of these helper blocks and the failed block.
[0038] First, these helper blocks within the same rack are aggregated. Then, within that rack, the linear encoding rules of erasure codes are used to partially compute these aggregated helper blocks, generating several "temporary blocks." These temporary blocks contain partial information needed to reconstruct the failed block, but they themselves may not yet be the complete block ultimately required. Next, each rack sends its generated temporary blocks over the network to the target rack (i.e., the rack where the requesting node is located) storing the failed block (which now needs to be repaired). Finally, at the target rack, temporary blocks from all relevant racks are collected, and the linear encoding rules of erasure codes are again applied to combine these temporary blocks to compute a complete, usable new block.
[0039] These new blocks are then written to the requesting nodes specified in the optimization plan, successfully repairing the failed blocks in the failed stripe. This step significantly reduces the amount of data that needs to be transferred across racks by performing some computation and data aggregation within the rack, further reducing network latency and bandwidth consumption.
[0040] Optionally, the method for generating the initial repair plan includes: Determine the remaining blocks based on the data block k, the check block m, and the failed block of the failed stripe; K blocks are randomly selected from the remaining blocks of the failed stripe as helper blocks to generate multiple different initial repair schemes, and storage nodes are selected to store the repaired request blocks; where k is a positive integer.
[0041] The storage node does not store any other blocks in the failed stripe except the failed block; the total number of blocks storing the failed stripe in the region to which the requested block belongs does not exceed m to achieve region-level fault tolerance; where m is a positive integer.
[0042] Specifically, in the embodiment of the present application, if the failed block is 1, the remaining blocks can be recorded as k+m-1. K blocks are selected from the k+m-1 remaining blocks and combined to obtain multiple different initial repair solutions.
[0043] When generating the initial repair plan, you also need to select a storage node to store the repaired blocks. The selection of the storage node must meet the following conditions: The storage node does not store any other blocks in the failed stripe except the failed block, which means that the storage node cannot store other blocks in the failed stripe to avoid data redundancy and potential conflicts.
[0044] The total number of blocks storing failed stripes in the region to which the requested block belongs does not exceed m: This condition is to achieve region-level fault tolerance. By limiting the number of blocks storing failed stripes in each region, we can ensure that a failure in one region will not result in the loss of data for the entire stripe.
[0045] Through the above steps, multiple initial repair plans can be generated and appropriate storage nodes can be selected for each plan. The initial repair plans will then be used in the optimization step to find the optimal repair plan, thereby reducing cross-domain transmission time while ensuring data reliability and fault tolerance.
[0046] Optionally, the method for obtaining the optimized repair solution includes: Randomly select a baseline repair solution from multiple initial repair solutions for the current failed stripe, and calculate the time required for cross-rack data transmission in the baseline repair solution as the baseline cross-domain transmission time; Traverse all valid initial repair plans for the current failed stripe and determine the target cross-domain transmission time for each initial repair plan; The target cross-domain transmission time is compared with the benchmark cross-domain transmission time, and any initial repair solution is selected to replace the benchmark repair solution according to the comparison result to obtain an optimized repair solution.
[0047] Specifically, from the multiple initial repair plans generated for the current failed stripe, one is randomly selected as the baseline repair plan. This baseline plan is then analyzed in detail to determine the amount of data required to be transferred across different racks and the bandwidth limitations between these racks. Based on this information, the total cross-rack data transfer time required to execute the baseline repair plan is calculated and defined as the baseline cross-domain transfer time. This baseline time represents the performance baseline for the currently randomly selected plan.
[0048] Next, we traverse all remaining valid initial repair solutions for the currently failed stripe, i.e., solutions other than the selected baseline solution. For each traversed initial repair solution, we repeat the analysis process similar to the baseline solution, and accurately calculate the total cross-rack data transfer time required to execute the solution as the target cross-domain transfer time.
[0049] Furthermore, the target cross-domain transmission time is compared with the benchmark cross-domain transmission time to replace the initial repair solution and obtain an optimized repair solution, including: Determine whether the target cross-domain transmission time is less than the benchmark cross-domain transmission time; If the target cross-domain transmission time is less than the benchmark cross-domain transmission time, the initial repair plan is replaced by the repair plan corresponding to the target cross-domain transmission time as the optimized repair plan; If the target cross-domain transmission time is not less than the benchmark cross-domain transmission time, all valid initial repair solutions and the transmission time comparison process are repeated until all alternative initial repair solutions are traversed or the target cross-domain transmission time is less than the benchmark cross-domain transmission time, and an optimized repair solution is obtained.
[0050] After calculating the target cross-domain transmission time for an initial repair solution, it is compared with the previously determined baseline cross-domain transmission time. The core goal of the optimization process is to reduce the cross-domain transmission time. Therefore, if the target cross-domain transmission time of the currently traversed solution is less than the baseline cross-domain transmission time, then the solution is considered to be superior. At this point, this superior solution replaces the current baseline repair solution, and the baseline cross-domain transmission time is updated with the target cross-domain transmission time of this solution. The system continues to traverse the next initial repair solution, repeating the above comparison and replacement process. Finally, when all valid initial repair solutions have been traversed, the baseline repair solution that remains is the optimized repair solution found through this iterative method based on benchmark comparison. This optimized solution has the lowest cross-domain transmission time among all the initial solutions examined.
[0051] Optionally, it also includes: Get the current network bandwidth of each rack, and get the bandwidth difference based on the network bandwidth of each rack; Based on the bandwidth difference, the transmission task volume of the high-bandwidth rack is increased and the transmission task volume of the low-bandwidth rack is reduced to reduce the cross-domain transmission time.
[0052] This application tends to increase the transmission task volume of nodes located on high-bandwidth racks, while correspondingly reducing the transmission task volume of low-bandwidth racks. The load adjustment strategy based on bandwidth differences can more reasonably distribute the data transmission burden across racks and give priority to using links with better network resources. Therefore, during the iterative optimization process, the overall cross-domain transmission time can be more effectively reduced, and the repair efficiency can be improved.
[0053] Optionally, according to an optimized repair scheme, helper blocks within the same rack are aggregated and encoded to repair failed blocks in each stripe, including: Determine the rack where the helper block of each failed stripe is located, and aggregate all the helper blocks in the rack; Using the linear encoding rules of erasure codes, the helper blocks aggregated in the rack are encoded to generate temporary blocks. The temporary blocks contain part of the information needed to reconstruct the failed blocks. Send the temporary blocks generated by each rack to the target rack where the requested blocks are located; The received temporary block is generated into a new block by using the linear encoding rule of the erasure code through the request block, and the new block is stored in the storage node where the corresponding request block is located to repair the failed block in the failed stripe.
[0054] Specifically, in the embodiment of the present application, first, for each failed stripe, a set of helper blocks is determined according to the optimized repair solution. Within each such rack, all blocks selected as helper blocks in the rack are aggregated for the next encoding operation.
[0055] Within each rack that aggregates helper blocks, a specific encoding calculation is performed on these aggregated helper blocks using the linear encoding rules of erasure codes. The goal of this encoding process is not to fully reconstruct the failed block, but rather to generate a new set of data blocks, called temporary blocks. These temporary blocks contain the partial information needed to reconstruct the failed block.
[0056] After each rack completes generating its temporary blocks, it sends them over the network to the target rack where the failed blocks are stored. The target rack is where the initial failure was detected and where new blocks are ultimately written. At the target rack, temporary blocks from all relevant racks are collected.
[0057] Then, using the linear encoding rules of the erasure code, combined with these received temporary blocks, a complete and usable new block is accurately calculated. The new block is used to replace the original failed block and is a data block or check block with completely equivalent functionality.
[0058] Finally, these calculated new blocks are written to the requesting node specified in the optimized repair solution, which is the original node storing the failed block. Once the write is completed, the failed block in the failed stripe is successfully repaired, and storage redundancy and data availability are restored.
[0059] The embodiment of the present application significantly reduces the amount of data that needs to be transmitted across racks by performing partial calculations and aggregation within the rack first, thereby effectively reducing network latency and bandwidth consumption and improving repair efficiency.
[0060] Figure 2This is the second flow chart of the method for generating a cross-rack multi-strip repair solution provided by the embodiment of the present application, such as Figure 2 As shown in Figure 1, the cluster using RS(2,2) encoding contains 3 racks and 6 nodes. Figure 2 (a) shows the layout of three stripes in the cluster and the time cost required to transfer a block between racks. Assume that blocks b1,6, b2,4 and b3,1 fail, and letters h and r represent helper blocks and request blocks respectively. Figure 2 As shown in (b), R1's download link has the longest transmission time (Td1=MT=7 seconds), and the repair schemes for all three stripes select R1 as the request rack. This indicates that the bottleneck of the repair process lies in R1's downlink. To reduce R1's download time, R3 is selected as the request rack for solu'1, and solu'1 replaces solu1 in Solu. Therefore, in Figure 2 In (c), after replacement, the maximum transmission time MT of the new multi-repair solution is reduced to 5 seconds.
[0061] Reference Figure 3 , the present application also provides a strip repair system, comprising: Initialization module 310, configured to randomly initialize each failed stripe to be repaired to obtain an initial repair plan for each failed stripe; An optimization module 320 is configured to determine bandwidth differences between different racks and optimize the initial repair solution using an iterative algorithm based on the bandwidth differences and an optimization goal to obtain an optimized repair solution; the optimization goal is to reduce cross-domain transmission time; The repair module 330 is configured to aggregate and encode helper blocks in the same rack according to the optimized repair solution to repair failed blocks in each stripe.
[0062] Optionally, the initialization module is specifically used to: Determine the remaining blocks based on the data blocks, check blocks, and failed blocks of the failed stripe; K blocks are randomly selected from the remaining blocks of the failed stripe as helper blocks to generate multiple different initial repair schemes, and storage nodes are selected to store the repaired request blocks; where k is a positive integer.
[0063] Optionally, the storage node does not store any other blocks in the failed stripe except the failed block; the total number of blocks storing the failed stripe in the region to which the requested block belongs does not exceed m, so as to achieve regional fault tolerance; where m is a positive integer.
[0064] Optionally, the optimization module is specifically used to: Randomly select a baseline repair solution from multiple initial repair solutions for the current failed stripe, and calculate the time required for cross-rack data transmission in the baseline repair solution as the baseline cross-domain transmission time; Traverse all valid initial repair plans for the current failed stripe and determine the target cross-domain transmission time for each initial repair plan; The target cross-domain transmission time is compared with the benchmark cross-domain transmission time to replace the benchmark repair solution and obtain an optimized repair solution.
[0065] Optionally, the optimization module is specifically used to: Get the current network bandwidth of each rack, and get the bandwidth difference based on the network bandwidth of each rack; Based on the bandwidth difference, the transmission task volume of the high-bandwidth rack is increased and the transmission task volume of the low-bandwidth rack is reduced to reduce the cross-domain transmission time.
[0066] Optionally, the optimization module is specifically used to: Determine whether the target cross-domain transmission time is less than the benchmark cross-domain transmission time; If the target cross-domain transmission time is less than the benchmark cross-domain transmission time, the initial repair plan is replaced by the repair plan corresponding to the target cross-domain transmission time as the optimized repair plan; If the target cross-domain transmission time is not less than the benchmark cross-domain transmission time, all valid initial repair solutions and the transmission time comparison process are repeated until all alternative initial repair solutions are traversed or the target cross-domain transmission time is less than the benchmark cross-domain transmission time, and an optimized repair solution is obtained.
[0067] Optionally, the repair module is specifically configured to: Determine the rack where the helper block of each failed stripe is located, and aggregate all the helper blocks in the rack; Using the linear encoding rules of erasure codes, the helper blocks aggregated in the rack are encoded to generate temporary blocks. The temporary blocks contain part of the information needed to reconstruct the failed blocks. Send the temporary blocks generated by each rack to the target rack where the requested blocks are located; The received temporary block is generated into a new block by using the linear encoding rule of the erasure code through the request block, and the new block is stored in the storage node where the corresponding request block is located to repair the failed block in the failed stripe.
[0068] It is understandable that the detailed functional implementation of each of the above units / modules can be found in the introduction of the aforementioned method embodiment, and will not be repeated here.
[0069] It should be understood that the above-mentioned device is used to execute the method in the above-mentioned embodiment. The implementation principle and technical effect of the corresponding program module in the device are similar to those described in the above-mentioned method. The working process of the device can refer to the corresponding process in the above-mentioned method and will not be repeated here.
[0070] Reference Figure 4Based on the methods in the above embodiments, an embodiment of the present application provides an electronic device, which may include: a processor (Processor) 410, a communication interface (Communications Interface) 420, a memory (Memory) 430, and a communication bus 440. The processor 410, the communication interface 420, and the memory 430 communicate with each other via the communication bus 440. The processor 410 may call the logic instructions in the memory 430 to execute the methods in the above embodiments.
[0071] In addition, the logic instructions in the aforementioned memory 430 can be implemented in the form of a software functional unit and, when sold or used as an independent product, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the portion that contributes to the prior art, or the portion of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application.
[0072] Based on the method in the above embodiment, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a processor, the processor executes the method in the above embodiment.
[0073] Based on the method in the above embodiment, an embodiment of the present application provides a computer program product. When the computer program product runs on a processor, the processor executes the method in the above embodiment.
[0074] It is understood that the processor in the embodiments of the present application may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA), other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.
[0075] The method steps in the embodiments of the present application can be implemented by hardware or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, mobile hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be an integral part of the processor. The processor and storage medium can be located in an ASIC.
[0076] The above embodiments can be implemented in whole or in part using software, hardware, firmware, or any combination thereof. When implemented using software, they can be implemented in whole or in part in the form of a computer program product. The computer program product comprises one or more computer instructions. When loaded and executed on a computer, the computer program instructions fully or partially produce the processes or functions described in the embodiments of this application. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted via the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible by a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be magnetic media (e.g., floppy disk, hard disk, tape), optical media (e.g., DVD), or semiconductor media (e.g., solid-state drive (SSD)).
[0077] It will be understood that the various numerical numbers involved in the embodiments of the present application are merely distinctions for the convenience of description and are not intended to limit the scope of the embodiments of the present application.
[0078] It is easy for those skilled in the art to understand that the above is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application should be included in the scope of protection of the present application.
Claims
1. A strip repair method, characterized in that: include: Randomly initialize each failed stripe to be repaired to obtain the initial repair plan for each failed stripe; Determining bandwidth differences between different racks, and optimizing an initial repair solution through an iterative algorithm based on the bandwidth differences and an optimization goal to obtain an optimized repair solution; the optimization goal is to reduce cross-domain transmission time; According to the optimized repair solution, helper blocks in the same rack are aggregated and encoded to repair failed blocks in each stripe.
2. The strip repair method according to claim 1, characterized in that: The method for generating the initial repair solution includes: Determine the remaining blocks according to the data blocks, check blocks and failed blocks of the failed stripe; K blocks are randomly selected from the remaining blocks of the failed stripe as helper blocks to generate multiple different initial repair schemes, and storage nodes are selected to store the repaired request blocks; wherein k is a positive integer.
3. The strip repair method according to claim 1, characterized in that: The storage node does not store any other blocks in the failed stripe except the failed block; the total number of blocks storing the failed stripe in the area to which the requested block belongs does not exceed m to achieve regional fault tolerance; where m is a positive integer.
4. The strip repair method according to claim 1, characterized in that: The method for obtaining the optimized repair solution includes: Randomly selecting a benchmark repair scheme from multiple initial repair schemes for the current failed stripe, and calculating the time required for cross-rack data transmission in the benchmark repair scheme as the benchmark cross-domain transmission time; Traversing all valid initial repair solutions for the currently failed stripe, and determining a target cross-domain transmission time for each initial repair solution; The target cross-domain transmission time is compared with the benchmark cross-domain transmission time, and any initial repair solution is selected to replace the benchmark repair solution according to the comparison result to obtain an optimized repair solution.
5. The strip repair method according to claim 4, characterized in that: Also includes: Get the current network bandwidth of each rack, and get the bandwidth difference based on the network bandwidth of each rack; According to the bandwidth difference, the transmission task volume of the high-bandwidth rack is increased and the transmission task volume of the low-bandwidth rack is reduced, so as to reduce the cross-domain transmission time.
6. The strip repair method according to claim 4, characterized in that: The step of comparing the target inter-domain transmission time with the benchmark inter-domain transmission time and selecting any initial repair solution to replace the benchmark repair solution according to the comparison result to obtain an optimized repair solution includes: Determining whether the target inter-domain transmission time is less than the benchmark inter-domain transmission time; If the target cross-domain transmission time is less than the benchmark cross-domain transmission time, replacing the initial repair solution with the repair solution corresponding to the target cross-domain transmission time as the optimized repair solution; If the target cross-domain transmission time is not less than the benchmark cross-domain transmission time, all valid initial repair solutions and the transmission time comparison process are repeated until all alternative initial repair solutions are traversed or the target cross-domain transmission time is less than the benchmark cross-domain transmission time, and the optimized repair solution is obtained.
7. The strip repair method according to claim 1, characterized in that: The step of aggregating and encoding helper blocks in the same rack according to the optimized repair solution to repair failed blocks in each stripe includes: Determine the rack where the helper block of each failed stripe is located, and aggregate all the helper blocks in the rack; Using the linear encoding rule of erasure codes, the helper blocks aggregated in the rack are encoded to generate temporary blocks. The temporary blocks contain partial information required to reconstruct the failed blocks. Sending the temporary block generated by each rack to the target rack where the requested block is located; The request block generates a new block from the received temporary block using the linear coding rule of the erasure code, and stores the new block in the storage node where the corresponding request block is located to repair the failed block in the failed stripe.
8. A strip repair system, characterized in that: include: An initialization module is used to randomly initialize each failed stripe to be repaired and obtain an initial repair plan for each failed stripe; An optimization module is configured to determine bandwidth differences between different racks and optimize an initial repair solution using an iterative algorithm based on the bandwidth differences and an optimization goal to obtain an optimized repair solution; the optimization goal is to reduce cross-domain transmission time; The repair module is used for aggregating and encoding the helper blocks in the same rack according to the optimized repair solution to repair the failed blocks in each stripe.
9. An electronic device, characterized in that: include: at least one memory for storing a computer program; At least one processor is used to execute the program stored in the memory. When the program stored in the memory is executed, the processor is used to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed on a processor, the processor is caused to execute the method according to any one of claims 1 to 7.