A hybrid acceleration method and device for repairing erasure code failed blocks

By integrating in-network repair and server repair modes in the erasure code storage system, combined with the aggregator resources of the programmable switch, the failure block repair strategy is optimized, the problem of low repair throughput in the existing technology is solved, and efficient failure block repair is achieved.

CN119402453BActive Publication Date: 2025-05-06NAT UNIV OF DEFENSE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510008756.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-06
Estimated Expiration
2045-01-03

AI Technical Summary

Technical Problem

The prior art is constrained by the up and downlink bandwidth of the node in the erasure code storage system during the repair of failure blocks in the erasure code storage system, resulting in low network latency and repair throughput.

Method used

A hybrid acceleration method for repairing the erasure code failure block is proposed. By integrating two modes of in-network repair and server repair in a storage cluster equipped with programmable switches, the optimal repair strategy is calculated based on the current network bandwidth and the number of aggregators of the programmable switches, including repair mode, routing path, aggregators and transmission rate.

Benefits of technology

It realizes the repair throughput of failed data blocks under limited resources, avoids repair degradation to the most primitive incast transmission mode, and significantly reduces data block transmission traffic and repair delay.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119402453B_ABST
    Figure CN119402453B_ABST
Patent Text Reader

Abstract

The present invention discloses a hybrid acceleration method and device for repairing erasure code failure blocks, the method steps include: step S01. monitoring the data damage state in the controlled network, and starting the repair control when receiving the failure repair request sent by the failure repair node; step S02. querying the metadata information of the stored data and determining the Helper node information involved in the repair from the relevant storage nodes according to the current network bandwidth and the number of available aggregators of the programmable switch, and calculating the optimal repair strategy for the failure block, the repair mode is a combination mode, and the routing flow table is issued; according to the determined repair strategy, each Helper node is controlled to send data, and the aggregation operation is performed at the programmable switch, and the failure data block is repaired at the failure repair node. The present invention can efficiently integrate the two modes of in-network repair and server repair, and improve the repair throughput of the failure data block under limited resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data storage systems, and in particular to a hybrid acceleration method and device for repairing erasure code failed blocks. Background Art

[0002] As the scale of the system continues to expand, the training process of large models is more susceptible to system failures (such as infrastructure failures or software errors), resulting in the loss of trained model states, which in turn results in a large waste of time and training resources. Regularly writing large model states (i.e., model parameters and optimizer states) to persistent storage is a key technology to ensure the fault tolerance of the storage system and the smooth progress of model training. Currently, the training state of large models is usually backed up from GPU memory to disk regularly by setting checkpoints for model training. When a failure occurs, the model state can be restored to the most recent checkpoint, thus avoiding a large waste of training resources and time. However, checkpoint technology still has difficulty in achieving model state recovery under hardware problems (such as overheating and power failure), which affects the normal progress of model training. In addition, since the state of large models usually ranges from hundreds of GB to TB, multi-copy fault tolerance of model states will occupy a large amount of storage resources, resulting in high fault tolerance costs.

[0003] Erasure code storage systems can achieve frequent backup of large model states with low storage overhead. For example, Reed-Solomon code (RS) Data unit encoding generation A redundant "check unit". When any data unit fails, Any of the units The failed unit can be decoded and repaired. Since backup technology requires multiple copies of a large number of model parameters, erasure coding technology only needs to generate a few additional check units to ensure fault tolerance for a small number of node failures. Therefore, erasure coding technology can achieve fault tolerance for large model states at a low storage cost.

[0004] However, the complex encoding and decoding process of the erasure code storage system may reduce the efficiency of repairing failed data blocks, thereby affecting the training recovery process of large models after a node failure. In order to accelerate the repair of failed data blocks, the prior art usually adopts an efficient encoding method to improve the encoding efficiency, and further accelerates the repair process through GPU hardware resources, or by building a balanced data layout to improve the repair throughput or reducing network bandwidth consumption through aggregation operations, or designing a pipeline for failed block repair for Repair Pipelining (RP) to serialize the repair operations of failed data across storage nodes, further improving the repair throughput. However, since the server is also involved in the data aggregation process, the above-mentioned repair acceleration method in the prior art will be constrained by the upstream and downstream bandwidth of the node at the same time, and the bottleneck link will introduce significant network delay, thereby affecting the failed block repair throughput. Summary of the invention

[0005] The technical problem to be solved by the present invention is: in response to the technical problems existing in the prior art, the present invention provides a hybrid acceleration method and device for repairing failed blocks of erasure codes with simple implementation method, low cost and high repair throughput, which can efficiently integrate the two modes of in-network repair and server repair, and improve the repair throughput of failed data blocks under limited resources.

[0006] In order to solve the above technical problems, the technical solution proposed by the present invention is:

[0007] A hybrid acceleration method for repairing failed blocks of erasure codes is applicable to repairing failed blocks in a storage cluster equipped with a programmable switch. The method comprises the following steps:

[0008] Step S01: Monitor the data damage status in the controlled network, and when receiving the failed repair node When a failure repair request is sent, the process goes to step S02 to start the repair control;

[0009] Step S02. Repair strategy determination: query the metadata information of the stored data and determine the number of available aggregators from the current network bandwidth and the programmable switch. The storage nodes that are determined to participate in the repair Helper node information, k and m represent the number of original data segmentation blocks and the number of redundant data blocks respectively, and take maximizing the repair throughput of failed blocks as the goal, calculate the optimal repair strategy for failed blocks according to the currently available network bandwidth resources and the number of programmable switch aggregators, the repair strategy includes the repair mode, routing path, the number of occupied switch aggregators and the transmission rate of the corresponding storage node, the repair mode is a combination mode in which some Helper nodes use the in-network aggregation mode and some Helper nodes use the server aggregation mode, and issue the routing flow table according to the determined repair strategy;

[0010] Step S03. Repair strategy execution: Control each Helper node to send data according to the determined repair strategy, and perform aggregation operations at the programmable switch. Repair the invalid data blocks according to the received data packets.

[0011] Furthermore, after step S02, the following steps are further included:

[0012] When each Helper node involved in the repair receives the repair instruction, each Helper node sends data according to the specified sending rules;

[0013] Each Helper node is based on the local block and coding coefficient Calculate the encoding block , and assign them to corresponding sending queues, each sending queue sends data on the corresponding routing path with a determined throughput and identifies the sent data according to the sending order;

[0014] performing an aggregation operation on the data packet payloads having the same identifier at the programmable switch;

[0015] Repairing a failed node Repair the invalid data blocks according to the received data packets.

[0016] Furthermore, step S02 also includes modeling the problem of accelerating the repair of failed blocks. During the modeling process, only the data transmission between the upper network and each rack is abstracted as the uplink and downlink bandwidth from the upper network to the ToR (Top of Rack Switch) switch. During the data aggregation process, the programmable switch processes data with the maximum data packet aggregation throughput PAT, and the remaining data exceeding the maximum data packet aggregation throughput PAT is directly forwarded to the destination repair node. , the same rack uses the intra-network aggregation or server aggregation mode.

[0017] Furthermore, the expression for modeling the problem of accelerating the repair of failed blocks is:

[0018] , st,

[0019] (1)

[0020] (2)

[0021] (3)

[0022] (4)

[0023] (5)

[0024] (6)

[0025] (7)

[0026] (8)

[0027] (9)

[0028] in, Indicates the data sending rate of each Helper node. Indicates that the slave storage node Uplink bandwidth to the ToR switch, Indicates the distance from the ToR switch to the storage node The downlink bandwidth is Indicates the rack The aggregation mode within, when it is the first value, it is the network aggregation, when it is the second value, it is the server aggregation; Indicates the rack Nodes within Whether it is selected as a Helper node, where the value is the first value if it is selected as a Helper node, and the value is the second value if it is not selected as a Helper node; They represent the aggregated traffic within the network, the unaggregated traffic within the network, and the aggregated traffic of the server. Indicates the rack The aggregate server throughput of Indicates ToR switches for each rack Uplink bandwidth, represents the aggregate throughput of all racks, Indicates a failed repair node The downlink bandwidth of the link. Repairing failed nodes The uplink bandwidth of the link. is the number of Helper nodes, and x represents the data sending rate of each Helper node.

[0029] Furthermore, in step S02, the optimal repair strategy for the failed block is determined by solving the failed block repair acceleration problem, including:

[0030] Enter each storage node Uplink and downlink bandwidth and And switches Maximum aggregate packet throughput And each ToR switch Uplink and downlink bandwidth and ;

[0031] According to the uplink bandwidth between the node and the ToR switch, the storage nodes of the failed block are Select the target Helper node;

[0032] The servers in the same rack uniformly use the intra-network aggregation or server aggregation mode. If the maximum throughput of intra-network repair is the minimum value of the uplink bandwidth between the Helper node and the ToR switch, the number of switch aggregators, and the uplink bandwidth between the switch and the upper network topology, the corresponding Helper node selects the intra-network aggregation mode; if the maximum throughput of server aggregation is the minimum value of the uplink and downlink bandwidths between the Helper node and the ToR switch, the corresponding Helper node selects the server aggregation mode;

[0033] Each Helper node sends data at the minimum throughput to coordinate the maximum aggregate throughput of each rack and select the ToR switch with the most aggregators. Repolymerization is performed;

[0034] Determine the size of the aggregated traffic after two rounds of aggregation. If the size of the aggregated traffic meets the requirements, repair the node. The bandwidth of the two downlinks in the rack and The maximum sending throughput of the Helper node and the aggregation mode mode_r in each rack are directly determined based on the constraints; otherwise, if any link bandwidth does not meet the constraints, the aggregation mode of all racks is directly set to server aggregation, and the data sending rate is set to , The minimum value in Racks in server aggregation mode The maximum aggregate throughput of

[0035] Aggregate the data between each rack one by one and save it on the switch The failed block is repaired, and the minimum throughput of the aggregated data is sent to the repair node. .

[0036] Furthermore, the racks in the network aggregation mode The maximum aggregate throughput limit is defined as confine_INA_r=min( ), rack in server aggregation mode The maximum aggregate throughput limit is defined as confine_RP_r=min( , ).

[0037] A hybrid acceleration device for repairing erasure code failed blocks, comprising:

[0038] The repair control startup module is used to monitor the data damage status in the controlled network and when it receives the failed repair node When a failure repair request is sent, the repair control is started;

[0039] The repair control module is used to query the metadata information of the storage data and select the aggregator from the programmable switch according to the current network bandwidth and the number of available aggregators of the programmable switch when the repair control is started. Determine the storage nodes involved in the repair It obtains information of Helper nodes and calculates the repair mode, routing path, number of occupied switch aggregators and transmission rate of all related storage nodes of the failed block. It sends the routing flow table according to the calculation results and configures the sending data window of the Helper node and the programmable switch aggregator resources.

[0040] Furthermore, the repair control module includes:

[0041] The modeling unit is used to model the problem of accelerating the repair of failed blocks. During the modeling process, only the data transmission between the upper network and each rack is abstracted as the uplink and downlink bandwidth from the upper network to the ToR switch. During the data aggregation process, the programmable switch processes the data with the maximum packet aggregation throughput PAT, and the rest of the data is directly forwarded to the destination repair node. , the same rack uniformly uses intra-network aggregation or server aggregation mode;

[0042] The solving unit is used to solve the problem of accelerating the repair of failed blocks.

[0043] An electronic device comprises a processor and a memory, wherein the memory is used to store a computer program, and the processor is used to execute the computer program to perform the above method.

[0044] A computer-readable storage medium storing a computer program, wherein the computer program implements the above method when executed by a processor.

[0045] Compared with the prior art, the advantages of the present invention are:

[0046] 1. The present invention aims to maximize the repair throughput of failed blocks, calculates the optimal repair strategy for failed blocks according to the currently available network bandwidth resources and the number of programmable switch aggregators, and adopts the repair mode of in-network repair and server combination. It can select the repair mode and optimize the configuration of routing paths according to the heterogeneous network bandwidth resources and the number of switch aggregators, efficiently integrate the two modes of in-network repair and server repair, integrate the advantages of in-network aggregation and server aggregation, and effectively improve the repair throughput of failed blocks.

[0047] 2. The present invention implements encoding and decoding operations of data blocks based on programmable switches through an in-network repair mode, which can significantly reduce data block transmission traffic and reduce repair delay. When the number of aggregators of the programmable switch is insufficient, the server repair mode can be used to avoid the repair being downgraded to the most original incast transmission mode, thereby maximizing the repair throughput of failed data blocks when the number of switch aggregators and link bandwidth resources are limited. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 It is a schematic diagram of the principles of traditional restoration, PPR, RP, and Paint restoration modes.

[0049] Figure 2 is an instructive example schematic diagram of the HFRA of the present invention.

[0050] Figure 3 Schematic diagram of the workflow of the HFRA method of the present invention.

[0051] Figure 4 It is a schematic diagram of the implementation process of the hybrid acceleration method for repairing failed blocks of erasure codes in this embodiment.

[0052] Figure 5 It is a schematic diagram of the network topology structure used in a specific application embodiment.

[0053] Figure 6 It is a schematic diagram of test results of the effect of the number of switch aggregators on the repair throughput in a specific application embodiment.

[0054] Figure 7 It is a schematic diagram of test results of the effect of injected traffic on repair throughput in a specific application embodiment.

[0055] Figure 8 It is a schematic diagram of the test results of the repair throughput of different encoding methods in a specific application embodiment.

[0056] Fig. 9 It is a schematic diagram of test results of the impact of different encoding methods on bandwidth overhead in a specific application embodiment.

[0057] Fig.10 It is a schematic diagram of test results of the effect of limiting the number of racks on the repair throughput in a specific application embodiment. DETAILED DESCRIPTION

[0058] The present invention is further described below in conjunction with the accompanying drawings and specific preferred embodiments, but the protection scope of the present invention is not limited thereby.

[0059] For ease of understanding, the relevant technical background of the present invention is first introduced by way of example.

[0060] 1. Erasure Code Storage System

[0061] Erasure codes (such as RS codes) have been widely used in major storage systems (such as HDFS and Ceph). Erasure codes are a linear coding theory, so the repair operation can be performed through Galois field addition. Specifically, The code will Data block encoding generation Redundancy check blocks: ( ),in It is a data block Corresponding to the check block The coding coefficient of Blocks are distributed and stored in Any failed block can be obtained by obtaining any other The blocks of storage nodes (Helper) are repaired and can tolerate up to The failure of a block (storage node). Figure 1 For example, Figure 1 (a) corresponds to the traditional repair mode, (b) corresponds to the PPR (PostPackage Repair) repair mode, (c) corresponds to the RP (Recovery Mode) repair mode, and (d) corresponds to the Paint repair mode. Encoding, when the data block needs to be read When a node fails, By reading the data block , Sum check block , Data Block This process is called a degraded read and introduces additional latency compared to normal data reads.

[0062] The encoding and decoding operations of erasure codes have linear properties: Each data block / check block in the code can be The Galois Field addition of the coded data blocks is performed by bitwise XOR operation. The linear nature of the encoding and decoding operation includes the following two characteristics:

[0063] Feature 1: The encoding and decoding operation does not change the data size. The bitwise XOR data block encoding and decoding operation can ensure that the encoding and decoding result has the same size as the original block. In other words, the original block , check block and other partial summations, such as The sizes are the same.

[0064] Property 2: Encoding and decoding operations are associative. The order of Galois Field addition does not change the result, for example, and Can be coded . Similarly, the decoding process also conforms to the associative law.

[0065] From the above feature 1, it can be seen that in the in-network repair mode, the original multiple data streams are bitwise XORed on the programmable switch, and only the aggregated result (the same size as a single data stream) is transmitted to the next hop, which can effectively reduce network traffic. Based on feature 2, the in-network repair mode can flexibly perform encoding and decoding of data blocks in any order.

[0066] 2. Programmable switches and in-network aggregation

[0067] Programmable switches have the ability to parse and operate on data packet fields. They abstract functions into a match-action pipeline through the Protocol Independent Switch Architecture (PISA). Users can complete the calculation process by programming the processing logic of data packets using some high-level languages. For example, Barefoot Tofino is a representative product of programmable switches. Programmable switches can process millions of packets per second at line speed while supporting sub-microsecond packet processing delays.

[0068] 3. Definition of on-network computing: On-network computing refers to network elements (such as programmable switches and smart network cards) performing programmed processing of data packets on the basis of forwarding traffic to implement programs that usually run on terminal hosts. Taking data-parallel model distributed training as an example, the gradient data aggregation operation of the working node is offloaded from the parameter server to the programmable switch for execution. When the data packet arrives at the programmable switch, the switch points it to the corresponding aggregator based on the sequence number of the data packet, which merges gradient data packets from different working nodes. When the aggregator completes the gradient aggregation operation, it forwards the aggregation results to the downstream device. In the above way, the gradient data streams from multiple working nodes can be aggregated into one data stream, which effectively reduces network traffic and data transmission delay, and accelerates the transmission of model gradients and the iteration of distributed training.

[0069] 4. The aggregation throughput is linearly related to the number of programmable switch aggregators: The packet aggregation throughput (PAT) of the programmable switch is defined to describe the maximum capacity of its aggregation operation. The aggregation function is an XOR operation on the encoded data. In a typical transmission protocol, the sender sends a window of data packets to the network in one RTT. When a programmable switch processes traffic, an aggregator needs to be assigned to each data packet in the flow. In one RTT, if the number of switch aggregators 𝑀 (in units of one data packet) is greater than the window size, the programmable switch can process all data packets in the window; otherwise, the switch can only process at most 𝑀 data packets in the window. Therefore, the aggregation throughput is linearly related to the number of programmable switch aggregators. The switch can aggregate traffic at a maximum rate of 𝑀 / 𝑇, and the traffic that cannot be aggregated will be directly forwarded by the switch to the destination node.

[0070] 5. In distributed jobs, in order to maintain the synchronization of operations, all nodes are usually required to maintain a consistent communication pace to avoid computing waits and communication delays caused by different sending rates. That is, the data sending rate of each distributed node is the same to avoid mutual waiting and aggregation delays.

[0071] When it comes to erasure code storage for fault-tolerant large model states, one solution is to enhance the model parameter protection method by designing a two-layer structure, which includes snapshot management process protection against software failures and erasure code storage protection against node failures. Compared with the traditional checkpoint method, this type of two-layer protection method can improve the survival rate of parameters and enhance the reliability of the system. At the same time, the hybrid method of erasure code and replica technology can also effectively update the model's redundant parameters and provide efficient data access. However, this type of two-layer solution does not take into account the time overhead of repairing failed blocks, which will cause the failure repair process to be affected by the uplink and downlink bandwidth of the communication link, thereby affecting the repair throughput.

[0072] In the prior art, the repair method for failed blocks in erasure code storage systems usually improves the repair throughput by improving the encoding method or optimizing the repair process. The improvement of the encoding method is mainly through the development of new encoding technology or hardware improvement to improve the repair performance. For example, one type is Locally Repairable Code (LRCs), which allows the use of fewer blocks to repair failed blocks, but requires additional storage space to save local parity blocks; the other type is regeneration code, which explores functional repair to reduce storage or network overhead. However, due to its highly restricted theoretical structure, it is difficult to apply to actual distributed storage systems.

[0073] Optimizing the repair process is another direction to improve repair efficiency, which mainly focuses on reducing the amount of data transmitted or improving transmission parallelism. Figure 1 As shown in (a), the traditional failed block repair method will cause congestion at the ingress port of the repair node. Aggregation is a failure repair technology used to reduce the amount of data transmission. For example, multiple data blocks are aggregated before the failed block decoding calculation to reduce the bandwidth resources occupied by failed data repair. The balanced distribution of stored data can also effectively improve the repair throughput, such as balancing the data layout within the node and rack through two orthogonal arrays. In addition, by integrating the aggregation method, the amount of data transmission across racks can be further reduced and the data repair efficiency can be improved.

[0074] In terms of improving data parallelism, one solution is to use a partial parallel repair mechanism to avoid congested transmission. Figure 1 As shown in (b), the blocks are multiplied by the corresponding decoding coefficients and then transmitted to other storage nodes at different timestamps. Figure 1As shown in (c), the failure repair is transformed into a one-by-one aggregation process. The storage node aggregates the data after receiving it from the upstream node, and then sends the aggregation result to the next node. In addition, each involved block is divided into many small-sized units and repaired in a pipelined parallel manner to further improve the degraded read performance. However, the pipeline method still requires considerable bandwidth overhead to support the required repair speed. In-network aggregation can avoid congestion of the ingress port of the repair node, and build multiple parallel tree pipelines to increase the repair speed. However, the above schemes do not take into account the resource limitations of the programmable switch aggregator. When the switch resources are insufficient, part of the traffic may degenerate to the traditional incast transmission mode, causing concurrent traffic to congest the network link.

[0075] Checkpoint technology can periodically back up the training state of large models from GPU memory to disk to avoid the loss of training state, but it cannot achieve model state recovery in the event of node failure. Erasure coding technology can achieve efficient fault tolerance for the failure of a few nodes at a low storage cost. However, since the repair of failed blocks of erasure coding requires obtaining corresponding storage blocks from multiple nodes, the repair process will bring high repair costs and time overhead, which in turn affects the state recovery process of large models.

[0076] Programmable switches (such as Tofino switches) provide powerful computing and caching capabilities for traditional storage and forwarding networks. In order to accelerate the distributed training of large models, one solution is to use in-network aggregation (INA). By completing the aggregation operation of model gradients on the programmable switch, the network traffic of the bottleneck link can be effectively reduced. Inspired by this, the present invention is based on the in-network repair mode. By offloading the repair operation of the failed data blocks of the erasure code storage system to the programmable switch for execution, the data blocks can be partially encoded in the network, and only the encoded data blocks are transmitted to the downstream link, thereby avoiding network congestion at the inbound port of the repair node and reducing the degraded read delay. In addition, the repair throughput of the failed block is only limited by the uplink bandwidth of each related node, thereby effectively improving the theoretical upper limit of the repair throughput.

[0077] However, in the actual in-network repair process, the efficiency of repairing failed data blocks will be affected by link bandwidth resources and the number of programmable switch aggregators. For example, when the number of switch aggregators is insufficient, the in-network repair mode will degenerate to the most primitive incast transmission. The parallel transmission of multiple data streams will occupy scarce bandwidth resources and cause the inbound ports of the repair nodes to be blocked, significantly reducing the repair throughput. In addition, due to the heterogeneity of network resources and the number of aggregators, the construction of routing paths for in-network aggregation is challenging and complex. Even the path planning algorithm that only considers bandwidth resource constraints is an NP-hard problem.

[0078] The present invention proposes a hybrid acceleration method HFRA (Hybrid Failure Repair Acceleration) for erasure code failure block repair for large model state fault tolerance, which efficiently integrates the two modes of in-network repair and server repair to improve the repair throughput of failed blocks. The in-network repair mode has the theoretically optimal bandwidth overhead and repair speed, but when the number of aggregators is insufficient, the aggregation mode will degenerate into traditional data; the server aggregation mode can avoid a large amount of concurrent incast transmission traffic, but its theoretical upper limit will be additionally affected by the downlink bandwidth of the relevant link. The HFRA of the present invention aims to maximize the repair throughput of failed blocks, calculates the optimal repair strategy for failed blocks according to the currently available network bandwidth resources and the number of programmable switch aggregators, adopts the repair mode of in-network repair and server combination, selects the repair mode and optimizes the configuration of routing paths according to the heterogeneous network bandwidth resources and the number of switch aggregators, fully integrates the two modes of in-network repair and server repair, and uses the in-network repair mode to implement the encoding and decoding operation of data blocks based on the programmable switch, which can significantly reduce the data block transmission traffic and reduce the repair delay. At the same time, when the number of aggregators of the programmable switch is insufficient, the server repair mode can be used to avoid the repair being downgraded to the most original incast transmission mode, thereby maximizing the repair throughput of failed data blocks under limited resources.

[0079] The optimization goal of the present invention is to determine the k storage nodes (Helpers) participating in the repair from the k+m-1 storage nodes when a failed block repair task occurs in a storage cluster equipped with a programmable switch, and calculate the optimal failed block repair mode, routing path, number of occupied switch aggregators, and data transmission rate of all Helper nodes based on the currently available network bandwidth resources and the number of programmable switch aggregators, so as to maximize the repair throughput of the failed block.

[0080] like Figure 2 As shown, the failed data block D 1 Need to be on the node The number of switch aggregators is indicated by black numbers, and the uplink and downlink bandwidths of each link are indicated by red and green numbers respectively. Encoding, when the node In case of failure, Select 5 Helper nodes to invalidate the data blocks Repairing the node Repair is performed at the location. Parameters such as the repair mode, routing path, number of occupied switch aggregators, and data transmission rate of all Helper nodes are determined by the upstream and downstream bandwidth resources of the relevant links and the number of aggregators of the programmable switch. The optimization goal is to maximize the repair throughput of the failed block.

[0081] However, insufficient number of aggregators will cause the in-network repair mode to degenerate into the most primitive incast transmission, causing a large amount of traffic to congest the network transmission link. Although in-network repair can significantly reduce the traffic flowing through the entire network, the number of switch aggregators will affect the aggregation throughput of data blocks for XOR operations. When the number of aggregators is insufficient, the in-network repair mode degenerates into the most primitive incast transmission. The concurrent transmission of multiple data streams will occupy a large amount of bandwidth resources and cause the inbound port of the repair node to be blocked.

[0082] like Figure 2 As shown, assuming , , , is selected as a Helper node and sends data at a rate of 6. If both racks use the intra-network aggregation mode, due to S 2 The number of aggregators can only support 3 units of data aggregation, and the node The downlink data transmission capacity is 3+3 (6-3)=12 (including 3 units of aggregated traffic, and S 1 , P 1 , P 2 The non-aggregated traffic exceeds its bandwidth limit (10). To address the above problems, the present invention adopts a mode combining server aggregation (such as RP) with intra-network aggregation to reduce the limitation of uplink and downlink bandwidth on aggregate throughput and avoid the large amount of data incast transmission that may be caused by the intra-network repair mode. Figure 2 For example, if the rack where switch S1 is located implements the intra-network aggregation mode and the rack where switch S2 is located uses the server aggregation mode, when all Helper nodes send data at a rate of 6, the nodes The downlink data transmission capacity is only 3+2 (6-3)=9 (including 3 units of aggregated traffic, and S 1 , S 2 non-aggregated traffic).

[0083] The present invention will be further described below in conjunction with specific embodiments.

[0084] like Figure 3As shown, it is applicable to repairing failed blocks in a storage cluster equipped with a programmable switch. The detailed steps of the hybrid acceleration method for repairing failed blocks with erasure codes in this embodiment include:

[0085] Step S01: Monitor the data damage status in the controlled network, and when receiving the failed repair node When a failure repair request is sent, the process goes to step S02 to start the repair control;

[0086] Step S02. Query the metadata information of the stored data and, based on the current network bandwidth and the number of available aggregators of the programmable switch, The storage nodes that are determined to participate in the repair Helper node information, k and m represent the number of original data blocks and the number of redundant data blocks respectively (the original data can be restored by any k of the k+m data blocks), and maximize the repair throughput of failed blocks. According to the currently available network bandwidth resources and the number of programmable switch aggregators, the optimal repair strategy for failed blocks is calculated. The repair strategy includes the repair mode of failed blocks, routing paths, the number of occupied switch aggregators, and the transmission rate of corresponding storage nodes. The repair mode is a combination of some Helper nodes using the in-network aggregation mode and some Helper nodes using the server aggregation mode. The routing flow table is issued according to the determined repair strategy.

[0087] Step S03. Repair strategy execution: Control each Helper node to send data according to the determined repair strategy, and perform aggregation operations at the programmable switch. Repair the invalid data blocks according to the received data packets.

[0088] like Figure 4 As shown in the figure, by using a controller as the core component of HFRA, its main function is to provide flexible network traffic scheduling and control capabilities. Specifically, the controller grasps the configuration information such as the network topology structure and centrally manages the currently available network bandwidth resources and the number of switch aggregators. In addition, the controller also saves the metadata information of the stored data. When the data is damaged, the failed repair node The controller queries the metadata information of the stored data and selects the aggregator from the current network bandwidth and the number of available aggregators of the programmable switch. Determine the storage nodes involved in the repair The controller obtains information about each Helper node② and calculates the repair mode, routing path, number of occupied switch aggregators, and transmission rate of all related storage nodes of the failed block. Based on this, the controller sends the routing flow table and configures the sending data window of the Helper node and the programmable switch aggregator resources③.

[0089] After the repair strategy is determined in step S02, the repair is performed according to the determined repair strategy in step S03. In this embodiment, step S03 specifically includes:

[0090] Step S301. When the Helper node participating in the repair receives the repair instruction, the Helper node sends data according to the specified sending rule;

[0091] Step S302. Each Helper node is based on its local block and its coding coefficient Directly calculate the encoding block , and assign them to corresponding sending queues, each sending queue sends data on the corresponding routing path with a determined throughput and identifies the sent data according to the sending order;

[0092] Step S303. Perform aggregation operation on the data packet payloads with the same identifier at the programmable switch;

[0093] Step S304: Repair the failed node Repair the invalid data blocks according to the received data packets.

[0094] by Figure 4 For example, after receiving the repair instruction from the controller, each Helper node sends data according to the sending rule④. In order to avoid the high packet processing delay caused by the codec performing complex multiplication operations on the programmable switch, the Helper sends data according to its local block and its coding coefficient Directly calculate the encoding block , and assigned to the corresponding sending queue. Each sending queue sends data on the corresponding routing path with a determined throughput, and each data packet is identified by PSN (Packet Sequence Number) in the order of sending. The data packet loads from different Helpers identified as the same PSN are aggregated at the programmable switch⑤. The aggregation throughput is linearly related to the number of aggregators assigned to the failed block repair task. Data packets that cannot be aggregated are forwarded directly on the switch. Finally, the failure repair node The invalid data block⑥ can be repaired according to the received data packet.

[0095] Considering that the search space of the failed block repair task will grow exponentially with the expansion of the cluster scale, it is relatively complex and difficult to determine the optimal failed data repair mode, routing path, resource occupancy, and data transmission rate due to the heterogeneity of network resources and the number of aggregators. Even the path planning algorithm that only considers bandwidth resource constraints is a type of NP-hard problem. In addition, as the cluster scale expands, the search space of the optimal repair strategy will also grow exponentially. In response to the above problems, the HFRA of the present invention first models the failed block repair acceleration problem, and then uses an efficient heuristic algorithm to solve it, so that a better failed block repair strategy can be quickly determined with lower time complexity.

[0096] The steps of modeling the failed block repair acceleration problem in this embodiment specifically include:

[0097] Assume that only the in-network aggregation protocol is deployed on the ToR (Top of Rack) switch. This assumption meets the demand for computing offload nearby. In addition, higher-level network topologies usually forward data based on the ECMP (Equal-Cost Multi-Path Routing) routing strategy to achieve multi-path load balancing. Therefore, this embodiment does not plan and design the upper-layer network transmission, but only abstracts the data transmission between the upper-layer network and each rack into the uplink and downlink bandwidth from the upper-layer network to the ToR switch. Figure 2 As shown, the network above the ToR switch is summarized as network, and the upstream and downstream bandwidth resources of the programmable switch S1 and the upper network are abstracted as 10 and 9 respectively.

[0098] When the intra-network aggregation is limited to the ToR switch, the data of any Helper node can be aggregated on at most two programmable switches due to the current intra-network aggregation protocol. This embodiment uses a best-effort approach to aggregate data, that is, the programmable switch processes the data with its maximum packet aggregation throughput PAT, and the rest of the data will be directly forwarded to the destination repair node .

[0099] When considering the data aggregation mode, this embodiment assumes that the same rack uniformly uses the intra-network aggregation or server aggregation mode. Using two modes in the same rack will generate three types of mixed traffic, including intra-network aggregated traffic, intra-network unaggregated traffic, and server aggregated traffic, which will cause great difficulty in protocol deployment and resource coordination overhead. In addition, due to the linear nature of the erasure code encoding and decoding, the aggregated traffic is the same size as the single traffic. Therefore, mixed traffic will cause additional data transmission overhead and affect the failed block repair throughput.

[0100] Based on the above assumptions, this embodiment models the problem of accelerating failed block repair. Table 1 describes the symbols involved and their meanings.

[0101] Table 1 Symbols and meanings

[0102]

[0103] In this embodiment, the failed block is placed at the node The problem modeling for repair is as follows:

[0104] , st,

[0105] (1)

[0106] (2)

[0107] (3)

[0108] (4)

[0109] (5)

[0110] (6)

[0111] (7)

[0112] (8)

[0113] (9)

[0114] in, They represent the aggregated traffic within the network, the non-aggregated traffic within the network, and the aggregated traffic of the server. Indicates ToR switches for each rack Uplink bandwidth, represents the aggregate throughput of all racks, Indicates a failed repair node The downlink bandwidth of the link. is the number of Helper nodes, are decision variables respectively. When the aggregation mode in rack r is intra-network aggregation 1, when the aggregation mode is server aggregation 0, rack Nodes within When selected as a Helper node 1. When not selected as a Helper node 0. In the superscript and subscript of the above parameters, i represents the serial number of the storage node, r represents the serial number of the server rack, and up and down represent the uplink and downlink respectively.

[0115] The above formula (1) indicates that the server aggregation mode is subject to the upstream bandwidth of the corresponding link, and the intra-network aggregation mode is subject to the upstream and downstream bandwidth of the corresponding link at the same time; formula (2) indicates that in the intra-network aggregation mode, for the racks that can perform intra-network aggregation ( >1), the aggregation throughput INAr depends on the minimum value of the data transmission rate and the aggregation capability of the programmable switch (maximum data packet aggregation throughput); Formula (3) indicates that the throughput UINAr of the unaggregated traffic is the part where the data transmission rate is greater than the aggregation capability of the programmable switch. The larger the data transmission rate and the more Helper nodes participate in the aggregation, the higher the throughput of the unaggregated traffic; Formula (4) indicates that for the server aggregation mode, the data throughput is the data transmission rate x.

[0116] Specifically, the goal of the failed block repair task is to maximize the repair throughput, that is, the rate at which all Helper nodes send data satisfies constraint (1) (Formula (1)), which limits the failed block repair throughput to not exceed the maximum link bandwidth allowed in each repair mode. That is, affected by the data transmission path, the maximum intra-rack network repair throughput is the minimum uplink bandwidth of all related nodes, and the server aggregation maximum throughput is the minimum value of the uplink and downlink bandwidths of all nodes. The actual aggregation throughput through the programmable switch is subject to the data transmission rate and aggregation capacity to satisfy constraint (2) (Formula (2)), while the throughput of unaggregated data is also related to the number of Helper nodes participating in the aggregation (corresponding to Formula (3)). The throughput in the server aggregation mode is the actual data transmission rate (corresponding to Formula (4)).

[0117] As shown in formula (5), for any rack Regardless of the selected aggregation mode, the amount of data after aggregation cannot exceed the corresponding ToR switch. Uplink bandwidth Then, as shown in formula (6), all relevant data will flow to the failure repair node Corresponding ToR switch , the total traffic does not exceed the upper network topology to Downlink bandwidth The above traffic includes three categories: intra-network aggregate traffic , Unaggregated traffic in the network and server aggregate traffic Traffic from all racks in intra-network aggregation mode ( and ) need to be summed, and for server aggregation mode, data traffic reaches the switch The aggregation has been completed one by one before, so the data transmission throughput is only As shown in formula (7), after the maximum polymerization capacity is After the ToR switches, the aggregate throughput from all racks No higher than the failure repair node Downlink bandwidth of the link Finally, as shown in formula (8), the number of Helper nodes Satisfy the task of repairing the failed blocks of the erasure code; as shown in formula (9), the decision variable It must be a 0-1 integer variable.

[0118] Based on the above optimization objectives and constraint conditions (1) to (9), even if the limitation on the number of switch aggregators is not considered and only the link bandwidth resources are used as constraints, the failed block repair problem is still relaxed into a mixed integer nonlinear programming (MINP) problem. The MINP problem is a type of NP-hard problem that cannot be solved in polynomial time. When the cluster size becomes larger, the solution time will become increasingly unacceptable. Therefore, the present invention solves the above failed block repair acceleration problem by adopting an efficient heuristic method to determine the optimal repair strategy, so that a better failure repair method can be quickly obtained under a reasonable time complexity, thereby improving the failure repair throughput.

[0119] In this embodiment, the specific steps of using the failed block repair acceleration heuristic algorithm to solve the failed block repair acceleration problem include:

[0120] Step S201. Obtain the storage nodes related to the failed block , Failure repair node , enter each storage node Uplink and downlink bandwidth and And ToR switches Aggregation capacity (maximum packet aggregate throughput) And each ToR switch Uplink and downlink bandwidth and ;

[0121] Step S202. Select a target Helper node from all relevant storage nodes of the failed block according to the uplink bandwidth between the node and the ToR switch;

[0122] Step S203. The servers in the same rack use the network aggregation or server aggregation mode uniformly, wherein the rack in the network aggregation mode The maximum aggregate throughput limit is defined as confine_INA_r=min( ), rack in server aggregation mode The maximum aggregate throughput limit is defined as confine_RP_r=min( , ), if the maximum throughput of the intra-network repair is the minimum value of the uplink bandwidth between the Helper node and the ToR switch, the number of switch aggregators, and the uplink bandwidth between the switch and the upper network topology, the corresponding Helper node selects the intra-network aggregation mode; if the maximum throughput of the server aggregation is the minimum value of the uplink and downlink bandwidths between the Helper node and the ToR switch, the corresponding Helper node selects the server aggregation mode;

[0123] Step S204. Each Helper node sends data at the minimum throughput to coordinate the maximum aggregate throughput of each rack, that is, each Helper node coordinates the maximum aggregate throughput of each rack, and uses the minimum value of the maximum aggregate throughput in each rack as the data transmission rate to test the data transmission rate. Select the ToR switch with the most aggregators at present Re-aggregation is performed, that is, a programmable switch with stronger aggregation capability is re-selected to utilize its remaining aggregation resources for another aggregation;

[0124] Step S205. Traffic after two rounds of aggregation from step S201 to step S204 The traffic after two rounds of aggregation is minimized. The size of the reduced aggregate flow is sufficient to repair the node The bandwidth of the two downlinks in the rack and The constraint directly determines the maximum sending throughput of the Helper node and the aggregation mode mode_r in each rack; otherwise, when any link bandwidth does not meet the constraint, the aggregation mode of all racks is directly set to server aggregation, and the data transmission rate is set to The minimum value of Racks in server aggregation mode The maximum aggregate throughput of

[0125] Aggregate the data between each rack one by one and save it on the switch The failed block is repaired, and the minimum throughput of the aggregated data is sent to the repair node. .

[0126] In a specific application embodiment, the hybrid acceleration algorithm for repairing failed blocks of erasure codes of Algorithm 1 below may be used:

[0127] Algorithm 1. Hybrid acceleration algorithm for repairing invalid blocks of erasure codes.

[0128] The input to the algorithm includes the encoding method used by the erasure code storage system , related storage nodes , failed repair node , any node Uplink and downlink bandwidth and ,switch Aggregation Capacity Its uplink and downlink bandwidth and The output of the algorithm includes the selection of Helper nodes, the data aggregation mode within each rack, and the maximum throughput of failed block repair. .

[0129] ① select k helpers from with the largest , construct and ;

[0130] ② for each rack containing helpers

[0131] ③confine_INA_r=min( );

[0132] ④confine_RP_r=min( , );

[0133] ⑤if confine_INA_r confine_RP_r

[0134] ⑥mode_r is INA in rack ;

[0135] ⑦else

[0136] ⑧mode_r is RP in rack ;

[0137] ⑨end if

[0138] ⑩ end for

[0139] ;

[0140] while )

[0141] change mode_r from INA to RP with the minimum (confine_INA_r-confine_RP_r);

[0142] end while

[0143] choose the second-layer aggregation switch with the maximum and sufficient ;

[0144] compute ;

[0145] if ||

[0146] change mode_r to RP in all racks

[0147] x min( );

[0148] end if

[0149] First, the combination of Helper node selection for all related storage nodes of the failed block will greatly increase the complexity of the algorithm. Encoding will appear combinations. Traversing each combination and designing the corresponding aggregation mode, resource allocation, and path planning will result in a large time overhead. Therefore, in order to select a suitable Helper node while reducing the time complexity, the above Algorithm 1 directly uses the common constraint of the intra-network aggregation or server aggregation mode: the uplink bandwidth of the node and the ToR switch as the selection indicator, as shown in step ① of Algorithm 1. The time complexity of the above Helper selection method is O((k+m-1)log(k+m-1)), and it avoids the impact of lower uplink bandwidth on the overall failure repair rate. and Respectively represent racks A collection of uplink and downlink bandwidth resources selected as the Helper node.

[0150] Secondly, consider using the intra-network aggregation or server aggregation mode for servers in the same rack to avoid mixed traffic and reduce the difficulty of protocol deployment and resource coordination overhead. When selecting the aggregation mode in each rack, the intra-network aggregation mode is limited by the uplink bandwidth of each Helper node and ToR switch, the number of switch aggregators, and the uplink bandwidth of the switch and the upper network topology, as shown in constraints (1) (Equation (1)) and (5) (Equation (5)); while the server aggregation mode is limited by the uplink and downlink bandwidth of the Helper node and the ToR switch, as shown in constraint (4) (Equation (4)). Therefore, Algorithm 1 tightens the constraints and sets the racks in the intra-network aggregation mode. The maximum aggregate throughput limit is defined as confine_INA_r=min( ) (Step ③); put the rack in server aggregation mode The maximum aggregate throughput limit is defined as confine_RP_r=min( , ) (step ④). The time complexity of selecting the aggregation mode is O(|r|). Even when In the case of The throughput of sending data is more than The part will be transformed into the traditional data forwarding mode. Using this forwarding mode will bring a large concurrent flow, which is subject to many restrictions on resources such as the aggregation capacity and link bandwidth of downstream switches and repair nodes. These restrictions determine the theoretical upper limit of the execution of the aggregation mode of each rack to a certain extent. Therefore, in order to avoid multiple rounds of iterations of the algorithm, the above algorithm 1 directly determines the aggregation mode through the maximum aggregation throughput in each rack (steps ⑤-⑨).

[0151] Then, each Helper node coordinates the maximum aggregate throughput of each rack and tests the data sending rate with its minimum value (step ). Since the failed repair node The aggregation resources of the ToR switch at the location may be limited, and the aggregation protocol of the existing technology only supports two-layer aggregation. The above algorithm 1 selects the ToR switch with the most sufficient number of current aggregators. Perform the second round of aggregation (rack Considering the bottleneck of the upper network and the downlink link of the switch, some racks in the in-network aggregation mode can be converted to server aggregation mode (step - ), the time complexity of the algorithm is O(|r| 2 ). Traffic after two rounds of aggregation The aggregated traffic after reduction needs to satisfy the repair node The bandwidth of the two downlinks in the rack are and When the bandwidth constraint is met, the maximum sending throughput of the Helper node and the aggregation mode mode_r in each rack can be directly determined. Otherwise, when the bandwidth of any link does not meet the constraint, the aggregation mode of all racks will be directly set to server aggregation, and the data sending rate will be set to The minimum value (step - ). At this point, the data between each rack will be aggregated one by one and stored on the switch. The failed blocks are repaired, and the repair throughput meets the requirements of all bandwidth resources, and the algorithm ends.

[0152] by Figure 2 When selecting the Helper node, it is considered that the aggregation mode of each rack is limited by the uplink bandwidth, such as the node The uplink bandwidth is only 1, which means that the maximum repair throughput of the rack will not be higher than 1 regardless of whether it uses the intra-network aggregation or server aggregation mode. , , , , Selected as a Helper node. For switches For rack 1, the maximum throughput of its in-network repair is , , Uplink bandwidth, switch The number of aggregators The minimum uplink bandwidth is (min(7,9,8,7,10)=7), and the maximum server aggregate throughput is , , The link bandwidth Therefore, rack 1 selects the intra-network aggregation mode, and rack 2 selects the server aggregation mode. Subsequently, each Helper node sends data at a throughput of min(6,7)=6, and the switch The aggregated data is sent to the repair node with a throughput of 6. . Because the switch and nodes The downlink bandwidth resources are all greater than 6, so the final failed block repair throughput is 6.

[0153] The present invention innovatively applies in-network computing to an erasure code storage system. Based on the dependence of the in-network repair throughput on programmable switch resources, the present invention integrates the in-network repair and server repair modes to form a hybrid acceleration method for repairing failed blocks of erasure codes suitable for large model state fault tolerance. The repair throughput of failed blocks can be improved, and because only the linear correlation of the erasure code is used in the repair process, it can also be flexibly applied to other erasure code encoding methods (such as LRC, regeneration code, etc.), and can be easily expanded to a parallel repair mode to further improve the repair throughput.

[0154] In order to verify the effectiveness of the present invention, the performance of the HFRA of the present invention was tested in a prototype system and large-scale simulation under a real environment and compared with other traditional repair methods. The test platform of the prototype system consists of 3 FPGA devices and 8 servers. Each FPGA device is equipped with an Intel Arria 10 FPGA chip and is equipped with 4 10GbE network interfaces. This embodiment implements the in-network aggregation logic of ATP on these FPGA devices and uses them as programmable switches. The above 3 FPGA devices are installed on a workstation equipped with 2 Intel Xeon Platinum 8124M processors, 128GB memory and a 500GB solid-state drive. The remaining 7 workstations correspond to the 7 hosts of H1-H7, each equipped with 2 Intel Xeon Platinum 8124M processors, 128GB memory, 500GB solid-state drive and an Intel 8259910GbE network card. The host and FPGA device are connected by a 10Gbps physical link, and its topology is as follows Figure 5 shown.

[0155] A GPT2-ML large model is used to train the model state after 220,000 steps on 15GB of cleaned text, and the Jerasure library is used to generate RS (5,3) coding blocks, each block size is set to 256MB. The experiment randomly specifies a failed block, and then randomly selects 5 coding blocks from the remaining blocks and places them on nodes H1-H5. Node H6 is set as a repair node, which needs to obtain these 5 related coding blocks from nodes H1-H5 to rebuild the failed block. In addition, the iperf tool is used to make node H7 send unidirectional traffic to node H1 at a specified rate to simulate different downstream bandwidths on the storage node. In addition, in order to make the server aggregation mode compatible with the aggregation logic in the network, this embodiment also organizes multiple aggregators on the server memory. Since the server's memory resources are sufficient, the total size of these end-side aggregators can accommodate the complete coding block data. Therefore, the message arriving at the server can be directly hashed to the corresponding aggregator and perform aggregation operations, and then the aggregation results will be immediately passed to the next node (switch or server). All aggregation operations are performed based on XOR operations, and the payload size of the message is set to 1024 bytes.

[0156] Finally, a Fat-tree network is built to conduct a large-scale simulation experiment, which includes 32 racks and a total of 128 servers. The bandwidth of each link is set to 100 Gbps, and the aggregation capacity of the ToR switch on each rack is randomly generated in the range of [0,100] Gbps. In order to simulate different link residual bandwidths, this embodiment generates network background traffic between servers based on the gravity model, and based on this, the residual bandwidth of the corresponding link is counted, and five erasure code encoding methods RS (3, 2), RS (5, 3), RS (6, 3), RS (9, 3) and RS (12, 4) are selected for the experiment. The storage nodes and repair nodes of all coding blocks are randomly generated within the cluster.

[0157] This embodiment compares the HFRA of the present invention with three traditional repair methods. The first is the traditional repair method Conv. When a block fails, Conv randomly selects a node from the available storage nodes. k Helper nodes are connected and the data blocks stored in them are transmitted to the requesting node for decoding and reconstruction. The second comparison method is RP, which connects the involved Helper nodes according to the routing path with the highest throughput and performs step-by-step aggregation of the linear pipeline. The third comparison method is INA, which uses the intra-network aggregation mode to repair the failed blocks and supports two rounds of data aggregation on the programmable switch according to the latest aggregation logic. The test results of the prototype system are the average of 10 runs, and the large-scale simulation results are the average of 100 runs.

[0158] This embodiment first tests the impact of the number of switch aggregators on the repair throughput. Figure 6 As shown in (a), when the number of aggregators of the programmable switch S3 is fixed at 15, as the aggregation capacity of S2 increases, the repair throughput of the HFRA method of the present invention is always higher than that of other traditional repair methods, and when the S2 aggregator data is 70, it achieves twice the repair throughput of the INA method. This is because the HFRA of the present invention can select the aggregation mode more flexibly. The increase in the S2 aggregation capacity enables the left rack to use intra-network aggregation and the right rack to use server aggregation, avoiding the limitation of the limited number of aggregators of S3 on the repair throughput. The throughput of the traditional RP and Conv methods remains unchanged because they do not use programmable switches to process data packets. After the number of aggregators is greater than 30, the throughput of the INA method no longer changes. This is because the limited number of aggregators on the right rack limits the maximum data transmission rate.

[0159] like Figure 6 (b) shows the failure repair throughput when the number of S2 aggregators is fixed to 45. The HFRA method of the present invention always adopts the method of intra-network aggregation in the left rack and server aggregation in the right rack, so that its throughput is maintained at 8.96 Gbps. For the INA method, the left and right racks always choose the intra-network aggregation mode. The initial throughput of the INA method is affected by the S3 aggregator, and stabilizes at 8.24 Gbps after the number of aggregators is greater than 45. This is because the second-layer aggregation is limited by the first-layer aggregation capacity due to the limitations of the current aggregation protocol. The data that is not aggregated in the first layer will be forwarded directly to the destination node. Therefore, the downlink transmission from S3 to H6 is congested, making the throughput of the INA method lower than that of the HFRA method of the present invention. The throughput of the Conv method is always maintained at 2 Gbps. This is because the traffic of the five Helper nodes divides the 10 Gbps link bandwidth resources of S3-H6 equally.

[0160] Subsequently, this embodiment tests the effect of injected traffic on the repair throughput. Node H7 sends unidirectional traffic to node H1 at a specified rate. The increase in injected traffic will occupy the downlink bandwidth from S2 to H1, thereby simulating the failed block repair performance under different link congestion conditions. Figure 7As shown in (a), when the number of aggregators of S2 and S3 is fixed at 45 and 15, the throughput of INA and Conv methods remains unchanged as the injected traffic increases. This is because these two methods only use the uplink bandwidth of each link. Downlink congestion has the most serious impact on the RP method, and its repair throughput drops from 9.78Gbps to only 1.53Gbps. When no traffic is injected, the HFRA method of the present invention selects the server aggregation mode for both the left and right racks, and its aggregation throughput is consistent with the RP method. When the S2-H1 downlink link is congested, the HFRA method of the present invention combines the intra-network aggregation and server aggregation modes, so that its repair throughput is finally stabilized at 8.75Gbps.

[0161] like Figure 6 (b) The number of aggregators in S2 and S3 is adjusted to 30 and 30, and the trend is similar to Figure 7 (a) is roughly the same, but the advantage of the HFRA method of the present invention over the INA method becomes smaller. This is because the HFRA method of the present invention selects the intra-network aggregation and server aggregation modes in the left and right racks respectively, causing the aggregation capability of the programmable switch S2 on the left to become a performance bottleneck.

[0162] This embodiment further uses a large-scale simulation experiment to evaluate the performance of the HFRA of the present invention and other comparative methods. First, Figure 8 The repair throughput of different erasure code encoding methods is compared. k and m As the value of gets larger, the repair throughput of each method gradually decreases. Although larger erasure code parameters can guarantee the fault tolerance performance of multiple node failures to a certain extent, the number of data blocks required to repair a data block is greater, and the racks involved are more dispersed. In addition, the INA and Conv methods are greatly affected by the erasure code parameters, because both of them will generate some cross-rack forwarding traffic, which is easily affected by blocked links and limits the data sending rate. For the MFRA and RP methods, since they use the server aggregation mode to aggregate multiple cross-rack traffic into a single traffic, they are less affected by the erasure code parameters.

[0163] Then, Fig. 9 The bandwidth resource overhead within and between racks under different encoding methods is shown. As the erasure code parameters increase, the bandwidth resource overhead within and between racks gradually increases. This is because the data block size used in this embodiment is fixed at 256MB, and more Helper node data will lead to larger data traffic. Fig. 9As shown in (a), for the INA and Conv methods, data transmission within the rack only requires the uplink bandwidth from the node to the switch, so the bandwidth overhead is the same and is only about 50% of the MFRA and RP methods (which use the uplink and downlink of the nodes within the rack). In addition, the intra-rack transmission strategy of the INA, Conv, and RP methods is determined under a given number of Helper nodes, so their inter-rack bandwidth overhead remains unchanged.

[0164] For data transmission between racks, such as Fig. 9 As shown in (b), the INA method occupies less bandwidth resources than Conv due to its advantage of intra-network aggregation. However, since the RP method requires repeated server aggregation between racks, its bandwidth overhead is about twice that of the INA method. The RP inter-rack bandwidth overhead under RS(12,4) encoding even reaches 5.93GB. The inter-rack bandwidth overhead of the HFRA method of the present invention is slightly lower than that of the RP method. This is because some racks adopt the intra-network aggregation mode, and the aggregated traffic only needs to be transmitted directly to the destination node.

[0165] at last, Fig.10 The effect of limiting the number of racks involved in the repair of failed blocks on the repair throughput under RS(9,3) coding is shown. All traffic of the Conv method is not aggregated, which will cause the inbound port blocking problem of the failed repair node. Under RS(9,3) coding, 9 data streams are congested in the downlink of the failed repair node, resulting in a repair throughput much lower than that of other comparison methods, only 0.88Gbps. In addition, the repair throughput of the Conv method is only related to the coding method and has nothing to do with the number of racks involved. When the number of racks where the Helper nodes are located is limited from 3 to 9, the locations of the Helper nodes are more dispersed, resulting in an increase in cross-rack traffic, which gradually reduces the repair throughput of the INA method. For the MFRA method and the RP method, since they adopt the server aggregation mode, a large amount of cross-rack traffic is converted into a single traffic on the downlink of the repair node. Therefore, the throughput of the two methods is less affected by the number of racks.

[0166] In summary, the HFRA method of the present invention can integrate the advantages of in-network aggregation and server aggregation, and obtain a higher repair throughput than other current methods under the conditions of limited number of switch aggregators and link bandwidth resources. Specifically, under RS ​​(9, 3) coding, the repair throughput of the HFRS method of the present invention is 2.5 times that of the INA method and 6.5 times that of the Conv method, which plays a vital role in the rapid recovery of the large model training state under node failure, and can ensure the fault tolerance performance of the system and ensure the smooth progress of large model training.

[0167] The hybrid acceleration method HFRA for repairing failed blocks of erasure codes for state fault tolerance of large models in the present invention selects the in-network repair mode and the server repair mode and optimizes the configuration of the routing path according to the heterogeneous network bandwidth resources and the number of switch aggregators. It can maximize the repair throughput of failed data blocks under limited network resources, and effectively ensure the state fault tolerance and rapid repair of large model training in the event of node failure.

[0168] This embodiment further provides a computer device, including a processor and a memory, wherein the memory is used to store a computer program, and the processor is used to execute the computer program to perform the above method.

[0169] It is understandable that the above method of this embodiment can be executed by a single device, such as a computer or server, etc., and can also be applied to a distributed scenario and completed by multiple devices in cooperation with each other. In the case of a distributed scenario, one of the multiple devices can only execute one or more steps in the above method of this embodiment, and multiple devices interact to complete the above method. The processor can be implemented in the form of a general-purpose CPU, a microprocessor, an application-specific integrated circuit, or one or more integrated circuits, etc., for executing related programs to implement the above method of this embodiment. The memory can be implemented in the form of a read-only memory ROM, a random access memory RAM, a static storage device, and a dynamic storage device. The memory can store an operating system and other applications. When the above method of this embodiment is implemented by software or firmware, the relevant program code is stored in the memory and called and executed by the processor.

[0170] This embodiment further provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the above method is implemented.

[0171] The above is only a preferred embodiment of the present invention, and does not limit the present invention in any form. Although the present invention has been disclosed as a preferred embodiment, it is not intended to limit the present invention. Therefore, any simple modification, equivalent change and modification made to the above embodiment according to the technical essence of the present invention without departing from the content of the technical solution of the present invention shall fall within the scope of protection of the technical solution of the present invention.

Claims

1. A hybrid acceleration method for repairing failed blocks in erasure codes, suitable for repairing failed blocks in a storage cluster equipped with a programmable switch, characterized in that: The steps of the method include: Step S01: Monitor the data damage status in the controlled network, and when receiving the failed repair node When a failure repair request is sent, the process goes to step S02 to start the repair control; Step S02. Repair strategy determination: query the metadata information of the stored data and determine the number of available aggregators from the current network bandwidth and the programmable switch. The storage nodes that are determined to participate in the repair Helper node information, k , m Respectively represent the number of original data segmentation blocks and the number of redundant data blocks, and take maximizing the repair throughput of the failed block as the goal, calculate the optimal repair strategy for the failed block according to the currently available network bandwidth resources and the number of programmable switch aggregators, the calculation of the optimal repair strategy for the failed block according to the currently available network bandwidth resources and the number of programmable switch aggregators includes modeling the failed block repair acceleration problem and solving the failed block repair acceleration problem to determine the optimal repair strategy for the failed block, the determination of the optimal repair strategy for the failed block by solving the failed block repair acceleration problem includes: servers in the same rack uniformly use the in-network aggregation or server aggregation mode, if the maximum throughput of the in-network repair is If the uplink bandwidth between the Helper node and the ToR switch, the number of switch aggregators, and the minimum value of the uplink bandwidth between the switch and the upper network topology, the corresponding Helper node selects the intra-network aggregation mode; if the maximum throughput of the server aggregation is the minimum value of the uplink and downlink bandwidths of the Helper node and the ToR switch, the corresponding Helper node selects the server aggregation mode; the repair strategy includes the repair mode, the routing path, the number of occupied switch aggregators, and the transmission rate of the corresponding storage node. The repair mode is a combination mode in which some Helper nodes use the intra-network aggregation mode and some Helper nodes use the server aggregation mode, and the routing flow table is issued according to the determined repair strategy; Step S03. Repair strategy execution: Control each Helper node to send data according to the determined repair strategy, and perform aggregation operations at the programmable switch. Repair the invalid data blocks according to the received data packets.

2. The hybrid acceleration method for repairing erasure code failed blocks according to claim 1, characterized in that: Step S03 includes: When each Helper node involved in the repair receives the repair instruction, each Helper node sends data according to the specified sending rules; Each Helper node calculates the coding block according to the local block and coding coefficient, and assigns it to the corresponding sending queue. Each sending queue sends data on the corresponding routing path with a determined throughput and marks the sent data according to the sending order. performing an aggregation operation on the data packet payloads having the same identifier at the programmable switch; Repairing a failed node Repair the invalid data blocks according to the received data packets.

3. The hybrid acceleration method for repairing erasure code failed blocks according to claim 1 or 2, characterized in that: In step S02, in the process of modeling the problem of accelerating the repair of failed blocks, the data transmission between the upper network and each rack is abstracted as the uplink and downlink bandwidth from the upper network to the ToR switch. In the process of data aggregation, the programmable switch aggregates the throughput of the maximum data packet. PAT Processing data, exceeding maximum packet aggregate throughput PAT The rest of the data is forwarded directly to the destination repair node , the same rack uses the intra-network aggregation or server aggregation mode.

4. The hybrid acceleration method for repairing erasure code failed blocks according to claim 3, characterized in that: The expression for modeling the problem of accelerating the repair of failed blocks is: , s.t., in, Indicates the data sending rate of each Helper node. Indicates that the slave storage node Uplink bandwidth to the ToR switch, Indicates the distance from the ToR switch to the storage node The downlink bandwidth is Indicates the rack The aggregation mode within, when it is the first value, it is the network aggregation, when it is the second value, it is the server aggregation; Indicates the rack Nodes within Whether it is selected as a Helper node, where the value is the first value if it is selected as a Helper node, and the value is the second value if it is not selected as a Helper node; They represent the aggregated traffic within the network, the unaggregated traffic within the network, and the aggregated traffic of the server. Indicates the rack The aggregate server throughput of Indicates ToR switches for each rack Uplink bandwidth, represents the aggregate throughput of all racks, Indicates a failed repair node The downlink bandwidth of the link. Repairing failed nodes The uplink bandwidth of the link. is the number of Helper nodes, and x represents the data sending rate of each Helper node.

5. The hybrid acceleration method for repairing erasure code failed blocks according to claim 4, characterized in that: In step S02, the solving of the failed block repair acceleration problem to determine the optimal repair strategy for the failed block also includes: Enter each storage node Uplink and downlink bandwidth and And ToR switches Maximum aggregate packet throughput And each ToR switch Uplink and downlink bandwidth and ; Select the target Helper node from all related storage nodes of the failed block based on the uplink bandwidth between the node and the ToR switch; Each Helper node sends data at the minimum throughput to coordinate the maximum aggregate throughput of each rack and select the ToR switch with the most aggregators. Repolymerization is performed; Determine the size of the aggregated traffic after two rounds of aggregation. If the size of the aggregated traffic meets the requirements, repair the node. The bandwidth of the two downlinks in the rack and The maximum sending throughput of the Helper node and the aggregation mode mode_r in each rack are directly determined based on the constraints; otherwise, if any link bandwidth does not meet the constraints, the aggregation mode of all racks is directly set to server aggregation, and the data sending rate is set to , The minimum value in Racks in server aggregation mode The maximum aggregate throughput of Aggregate the data between each rack one by one and save it on the switch The failed block is repaired, and the minimum throughput of the aggregated data is sent to the repair node. .

6. The hybrid acceleration method for repairing erasure code failed blocks according to claim 5, characterized in that: In-network aggregation mode rack The maximum aggregate throughput is defined as confine_INA_r=min( ), rack in server aggregation mode The maximum aggregate throughput is defined as confine_RP_r=min( , ).

7. A hybrid acceleration device for repairing failed blocks of erasure codes, characterized in that: include: The repair control startup module is used to monitor the data damage status in the controlled network and when it receives the failed repair node When a failure repair request is sent, the repair strategy determination module is transferred to start the repair control; The repair strategy determination module is used to query the metadata information of the stored data and select the repair strategy from the current network bandwidth and the number of available aggregators of the programmable switch. The storage nodes that are determined to participate in the repair Helper node information, k , m Respectively represent the number of original data segmentation blocks and the number of redundant data blocks, and take maximizing the repair throughput of the failed block as the goal, calculate the optimal repair strategy for the failed block according to the currently available network bandwidth resources and the number of programmable switch aggregators, the calculation of the optimal repair strategy for the failed block according to the currently available network bandwidth resources and the number of programmable switch aggregators includes modeling the failed block repair acceleration problem and solving the failed block repair acceleration problem to determine the optimal repair strategy for the failed block, the determination of the optimal repair strategy for the failed block by solving the failed block repair acceleration problem includes: servers in the same rack uniformly use the in-network aggregation or server aggregation mode, if the maximum throughput of the in-network repair is If the uplink bandwidth between the Helper node and the ToR switch, the number of switch aggregators, and the minimum value of the uplink bandwidth between the switch and the upper network topology, the corresponding Helper node selects the intra-network aggregation mode; if the maximum throughput of the server aggregation is the minimum value of the uplink and downlink bandwidths of the Helper node and the ToR switch, the corresponding Helper node selects the server aggregation mode; the repair strategy includes the repair mode, the routing path, the number of occupied switch aggregators, and the transmission rate of the corresponding storage node. The repair mode is a combination mode in which some Helper nodes use the intra-network aggregation mode and some Helper nodes use the server aggregation mode, and the routing flow table is issued according to the determined repair strategy; The repair strategy execution module is used to control each Helper node to send data according to the determined repair strategy, and perform aggregation operations at the programmable switch. Repair the invalid data blocks according to the received data packets.

8. The hybrid acceleration device for repairing failed blocks of erasure codes according to claim 7, characterized in that: The repair strategy determination module includes: The modeling unit is used to model the problem of accelerating the repair of failed blocks. During the modeling process, only the data transmission between the upper network and each rack is abstracted as the uplink and downlink bandwidth from the upper network to the ToR switch. During the data aggregation process, the programmable switch processes the data with the maximum packet aggregation throughput PAT, and the rest of the data is directly forwarded to the destination repair node. , the same rack uniformly uses intra-network aggregation or server aggregation mode; The solving unit is used to solve the problem of accelerating the repair of failed blocks and determine the optimal repair strategy.

9. An electronic device comprising a processor and a memory, wherein the memory is used to store a computer program, wherein: The processor is configured to execute the computer program to perform the method according to any one of claims 1 to 6.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Control method and device for data block repair

    CN111385200A

  • Multi-node scheduling repair method and system based on erasure codes

    CN113721848A