In-network Aggregation Erasure Code Recovery System and Method Based on UDP Protocol
Through the programmable switching system based on the UDP protocol, data block aggregation at the network level, the problems of network transmission bottlenecks and CPU load in the existing technology are solved, efficient data recovery and throughput improvement are achieved, and various data distribution models are suitable for big data centers.
Patent Information
- Application Number
- CN202310210597.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-07
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2043-03-07
AI Technical Summary
The existing erasure code recovery scheme is affected by network transmission bottlenecks and computing load on storage nodes with limited CPU performance. The TCP protocol-based solution increases packet processing delay in a stable network, making it impossible to recover data efficiently.
Data block aggregation is performed using a programmable switching system based on the UDP protocol, and N network streams are aggregated into one stream through a programmable switch, reducing network transmission bottlenecks and reducing the CPU load of the target node, and increasing throughput using a custom protocol.
Effectively reduce network transmission bottlenecks, reduce storage node computing overhead, significantly improve the throughput and delay of the recovery process, and adapt to various data distribution models in big data centers.
Smart Images

Figure CN116192988B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a technology in the field of distributed file processing, and specifically to an in-network aggregation erasure code recovery system and method based on the UDP protocol. Background Art
[0002] Modern data centers use erasure codes to encode the stored data to ensure data reliability in a way with lower storage overhead. The most mainstream erasure code used at present stage is the Reed-Solomon code (RS code). Most erasure codes will split the data into independent data blocks. Its advantage is that in the case of data loss, the aggregated data can be obtained by transmitting multiple remaining data blocks to a single target node. However, the data recovery speed is largely limited by the network bandwidth of the receiver, and during the recovery process, the receiving node needs to perform a large amount of calculations and requires memory buffering to store the intermediate data of the calculations. Especially for some storage nodes with limited CPU performance and the number of cores, the above disadvantages are particularly obvious. Therefore, we can move the erasure code calculation to be executed in the switch, which can be not affected by the network transmission bottleneck and can reduce the load on the CPU of the storage node. In recent years, in-network aggregation data recovery schemes have been proposed, but the existing schemes have deficiencies in terms of performance and network adaptability. The main reason is that the existing schemes are all based on the TCP protocol, and the TCP flow control and acknowledgment functions are useless in a stable network environment, but instead increase the processing delay of data packets. Summary of the Invention
[0003] In view of the above deficiencies of the prior art, the present invention proposes an in-network aggregation erasure code recovery system and method based on the UDP protocol. Compared with the existing RS erasure code, transmitting N data blocks to a single target node will cause network congestion. The present invention uses a programmable switching system to replace the data aggregation process performed by the storage node to connect all storage nodes, so that all block data can be aggregated during the transmission process. Through the programmable switch, N network flows are effectively aggregated into one network flow to implement the erasure code recovery process, reduce the influence of the network transmission bottleneck and reduce the load on the CPU of the target node. In addition, the present invention uses a custom protocol based on UDP, which increases the effective throughput compared with the TCP-based solution. And it can adapt to various data distribution models of the current big data center.
[0004] The present invention is realized through the following technical solutions:
[0005] The present invention relates to an in-network aggregation erasure code recovery system based on the UDP protocol, including: a programmable switching system (P4EC), an auxiliary process module provided at a target node, and a sender process module provided at a sender node, wherein: the auxiliary process module forwards the recovery request of the local distributed file system in the form of a recovery request to a plurality of sending nodes and waits for the aggregated data from the programmable switching system; the sender process module reads data blocks from the local operating system according to the recovery request of the auxiliary process and sends them to the programmable switching system; after aggregating the data bytes in the data blocks fetched from the sender nodes, the programmable switching system sends them to the target node in the form of data packets according to a custom protocol based on UDP.
[0006] Preferably, when the auxiliary process module detects packet loss through the listening mode, it will send a data packet retransmission request to all sender process modules.
[0007] The programmable switching system includes: a control module, a parser, a routing table module, a checksum module, a metadata management module, and a byte processor, wherein: the control module performs aggregation processing through the byte processor, and directly forwards the aggregated data to the target node when the number of data blocks participating in the calculation meets the requirement, otherwise stores the obtained temporary result in a register first; the parser parses each field in the protocol header according to the network protocol information to detect the data packets of the custom protocol; the routing table module forwards the data packets to the correct destination according to the mapping between the network address and the switch port; the checksum module updates the UDP checksum field according to the data packets; the metadata management module records the count of the data blocks that have been received for each index through a bitmap according to the data packet index and the block number field, and determines whether the current data packet is the last data packet for a given index; the byte processor performs aggregation processing on individual data bytes according to the index and block number information, and stores the temporary result in the corresponding register.
[0008] The temporary result refers to the result obtained by performing coefficient multiplication and summation on data blocks with a quantity less than N, where N is the RS parameter. When the data output of less than N blocks is sent to the programmable switching system, the calculation result is temporarily stored on the switch until the data of other blocks is received.
[0009] The aggregated data refers to the result obtained by performing coefficient multiplication and summation on N data blocks with a quantity meeting the RS parameter.
[0010] The custom protocol includes: an additional 3-byte protocol header set based on the UDP protocol header and 64 bytes of data. During the transmission process, the data blocks are gradually sent by the sender process using the custom protocol. For each subsequent data packet, its index is incremented by 1, and the block number remains unchanged.
[0011] The protocol header includes: a 2-byte data index field and a 1-byte ID field.
[0012] The result of the aggregation processing for each received data byte is stored in the internal register of the switch. The calculation specifically includes: coefficient multiplication calculation and summation calculation, wherein: coefficient multiplication calculation multiplies each received byte with the coefficient related to the sender, and the result is obtained by querying the 26×26 byte multiplication table already in the switch memory. Therefore, this operation only requires one memory search; summation calculation adds the result of multiplication to the value stored in the switch register, and when the summation result is a temporary result, it is stored back in the register, otherwise it is forwarded to the target node.
[0013] The monitoring mode means that after the recovery process is started, the auxiliary process module sends a recovery request to all sender nodes. After initiating the recovery request, the auxiliary process module will enter the monitoring mode and wait to receive the recovered data. Even if the data packets are out of order, the original data blocks can be re-aggregated from the received data by referencing the index field of the custom protocol. The distributed file system is notified after the data blocks are fully re-aggregated.
[0014] Preferably, the auxiliary process module sets a countdown for each of the five data packets with the lowest index value expected to be received. When no data packet is received within the specified time, the auxiliary process module sends a data packet retransmission request to all sending nodes.
[0015] The target node and the sender node are both provided with the distributed file system for collaboratively controlling the encoding of data and the distribution of data blocks between nodes. When data loss is detected, the distributed file system will transmit a recovery request to the auxiliary process to start the recovery process.
[0016] The recovery request includes: the IP addresses of all sending nodes participating in the recovery and the names of the data blocks to be sent. The auxiliary process module sends the recovery request to all sending nodes according to the recovery request. The sending process module responds to the recovery request by reading the data blocks to be sent and sending them to the programmable switching system.
[0017] The data packet retransmission request includes: the name of the data block to be sent and the UDP packet of the index of the data packet lost.
[0018] Technical Effects
[0019] The present invention utilizes the UDP protocol to perform data block recovery within a programmable switch, and all data transmissions during the recovery process are based on the UDP protocol. Compared with the prior art, the present invention can be integrated with existing distributed file systems and adapt to various data distribution models in data centers. The switch does not require data block distribution information for packet aggregation. Compared with the method based on the TCP protocol, the present invention can significantly increase throughput, reduce latency, and reduce the computational overhead of the programmable switching system. Description of the Drawings
[0020] Figure 1 It is a schematic diagram of the application scenario of the embodiment;
[0021] Figure 2 It is a flowchart of the recovery process involved in the present invention;
[0022] Figure 3 It is a schematic diagram of the topological structure of the P4EC program module involved in the present invention;
[0023] Figure 4 It is a schematic diagram of the custom recovery protocol structure involved in the present invention;
[0024] Figure 5 It is a schematic diagram of the in-network aggregation mechanism of the present invention. Detailed Embodiment
[0025] As Figure 1 shown, it is an in-network aggregation erasure code recovery system based on the UDP protocol according to this embodiment, including: a programmable switching system disposed in a programmable switch, N sender nodes respectively connected thereto through a network, and a target node. An auxiliary process module is provided in the target node, a sender process module is provided in each sender node, and a distributed file system is provided in both the sender nodes and the target node, where: N is an RS code parameter.
[0026] As Figure 3As shown in the figure, the programmable switching system includes: a control module, a parser, a routing table module, a checksum module, a metadata management module, and a byte processor, where: The control module aggregates the data included in the data packet through the byte processor, and directly forwards the aggregated data to the target node when the number of data blocks participating in the calculation is met; otherwise, the obtained temporary result is first stored in the register. The parser parses each field on the protocol header according to the network protocol information to detect data packets of the custom protocol. The routing table module forwards the data packet to the correct destination according to the mapping between the network address and the switch port. The checksum module updates the UDP checksum field according to the data packet. The metadata management module records the count of the data blocks that have been received for each index through a bitmap according to the data packet index and the block number field, and determines whether the current data packet is the last data packet for the given index. The byte processor performs aggregation processing on a single data byte according to the index and block number information, and stores the result in the corresponding register.
[0027] As Figure 2 shown in the figure, the in-network aggregation erasure code recovery based on the UDP protocol in this embodiment based on the above system includes:
[0028] 1) When the target node detects data loss, the distributed file system of the node will pass the recovery request to its auxiliary process module.
[0029] 2) The auxiliary process module sends a recovery request to all sender nodes according to the information included in the recovery request.
[0030] 3) All sender process modules that receive the recovery request read the data blocks from the local operating system according to the recovery request, and send the data blocks to the programmable switching system according to the custom protocol.
[0031] 4) After the programmable switching system aggregates the data bytes in the data packet to be processed from the sender node, it forwards the aggregated data to the target node.
[0032] 5) The auxiliary process module of the target node re-aggregates the lost data blocks according to the received aggregated data, and notifies the distributed file system that the data blocks have been recovered.
[0033] As Figure 4 shown in the figure, the custom protocol refers to: adding an additional 3-byte protocol header to the UDP protocol header, which includes a data index field (2 bytes) and an ID field (1 byte); during the transmission process, the data blocks are gradually sent by the sender process using the custom protocol. For each subsequent data packet, the index is incremented by 1, and the block number is fixed. When the index exceeds 2 16 , it will be reset to 0.
[0034] AsFigure 5 As shown, the aggregation process refers to: data packets of different data blocks are aggregated according to their indexes, and each index is associated with a register and a bitmap on the switch; when the programmable switching system receives data of all blocks with a given index, the result is sent to the target node, otherwise, the calculated result is stored at the position identified by the index field of its register, and the UDP checksum of the data packet is updated through the checksum module before sending, and then the data packet is forwarded to the correct network address through the routing table.
[0035] In this embodiment, the programmable switching system is implemented using the P4_16 framework. By compiling the P4_16 program into a Match-Action Unit (MAU) pipeline, it is applied to all received data packets. The MAU specifies checking specific fields from each data packet to match the corresponding operations to be performed subsequently. To ensure the performance of the program, all calculations performed on the data packets are determined at compile time.
[0036] The coefficient multiplication in the aggregation process described above is the multiplication on the Galois field related to the RS code; preferably, all 256×256 multiplication results are pre-computed and written into the memory when the programmable switching system starts up.
[0037] The summation operation in the aggregation process described above is the binary exclusive OR operation between two values, where: the first value is the result of the coefficient multiplication, and the second value is read from the switch register at the given index.
[0038] When the distributed file system uses RS(6, 3) (N = 6) coding and contains 10 nodes. Each node should run the distributed file system process, as well as an auxiliary process and a sender process. Although it is not required for the nodes to be directly connected to the programmable switching system, all node communications should be carried out through the programmable switching system.
[0039] When a user stores a 6MB file in the file system. The file is split into 9 1MB blocks, located at nodes 0 - 8. When the data block of node 0 is lost. Node 1 can be selected as the target node, and nodes 2 - 7 can be selected as the senders. The auxiliary process of node 1 will send a recovery request to nodes 2 - 7. Nodes 2 - 7 will read the specified data blocks and send them using a custom protocol. Each node will send 15625 packets, with the index range from 0 to 15624.
[0040] Data transmitted from nodes 2 - 7 will all be aggregated through the programmable switching system. For each index, when the data from 6 blocks is aggregated, the programmable switching system will forward the aggregated data to the target node.
[0041] After virtual environment experiments, a distributed file system using the above method was deployed on a test server with 20 CPU cores and 125 GB of memory. With an RS(6, 3) encoding configuration and a 1 Mbps link bandwidth, the recovery process of 1 MB data blocks was initiated. The experimental data obtained was as follows: the amount of data sent through the target node link was 1703455 B, and the recovery process time was 13.6 seconds.
[0042] As shown in Table 1, it is a comparison of the amounts of data transmitted through the target node link among the method of the present invention, the method based on the TCP protocol, and the original RS code recovery method.
[0043] Recovery method Data volume (lower) Data volume (upper) Total data volume Transmission overhead Original RS code recovery method 7432999 106861 7539860 7.5× Method based on TCP protocol 1968750 1031250 3000000 3.0× The present invention 1703125 330 1703455 1.7×
[0044] In summary, compared with the prior art, by connecting all storage nodes to a programmable switching system for packet aggregation, the present invention greatly reduces the impact of network bottlenecks in the RS code recovery process and reduces the computational overhead of storage nodes. The recovery time can theoretically be reduced by up to N times, where N is the number of senders.
[0045] Those skilled in the art can make partial adjustments to the above specific implementation in different ways without departing from the principles and purposes of the present invention. The protection scope of the present invention is subject to the claims and is not limited by the above specific implementation. All implementation solutions within its scope are subject to the present invention.
Claims
1. An in-network aggregation erasure code recovery system based on the UDP protocol, characterized in that, Including: A programmable switching system, an auxiliary process module disposed at a target node, and a sender process module disposed at a sender node, where: The auxiliary process module forwards a recovery request of a local distributed file system in the form of a recovery request to a plurality of sending nodes and waits for the aggregated data from the programmable switching system; The sender process module reads data blocks from the local operating system according to the recovery request of the auxiliary process and sends them to the programmable switching system; The programmable switching system aggregates the data bytes in the read data blocks from the sender nodes and then sends them to the target node in the form of data packets according to a custom protocol based on UDP; The programmable switching system includes: a control module, a parser, a routing table module, a checksum module, a metadata management module, and a byte processor, where: The control module performs aggregation processing through the byte processor, and directly forwards the aggregated data to the target node when the number of data blocks participating in the calculation is satisfied, otherwise stores the obtained temporary result in a register first; The parser parses each field on the protocol header according to the network protocol information to detect data packets of the custom protocol; The routing table module forwards the data packet to the correct destination according to the mapping between the network address and the switch port; The checksum module updates the UDP checksum field according to the data packet; The metadata management module records the count of the data blocks that have been received for each index through a bitmap according to the data packet index and the block number field, and determines whether the current data packet is the last data packet for a given index; The byte processor performs aggregation processing on individual data bytes according to the index and block number information, and stores the temporary result in the corresponding register; The custom protocol includes: an additional 3-byte protocol header set based on the UDP protocol header and 64 bytes of data.
2. The in-network aggregation erasure code recovery system based on the UDP protocol according to claim 1, characterized in that When the auxiliary process module detects packet loss through the listening mode, it will send a data packet retransmission request to all sender process modules; The data packet retransmission request includes: a UDP packet of the name of the data block to be sent and the index of the lost data packet.
3. The in-network aggregation erasure code recovery system based on the UDP protocol according to claim 1, wherein The temporary result refers to: the result obtained by performing coefficient multiplication and summation on data blocks with a quantity less than N, where N is an RS parameter. When the data output of less than N blocks is sent to the programmable switching system, the calculation result is temporarily stored on the switch until the data of other blocks is received; The result of aggregating and processing each received data byte is stored in the internal register of the switch. The calculation specifically includes: coefficient multiplication calculation and summation calculation, where: the coefficient multiplication calculation multiplies each received byte by the coefficient related to the sender, and the result is obtained by querying the byte multiplication table already existing in the switch memory. Therefore, this operation only requires one memory lookup; the summation calculation adds the result of the multiplication to the value stored in the switch register. When the summation result is a temporary result, it is stored back in the register, otherwise it is forwarded to the target node.
4. The in-network aggregation erasure code recovery system based on the UDP protocol according to claim 2, wherein The listening mode refers to: After the recovery process is started, the auxiliary process module sends a recovery request to all sender nodes. After sending the recovery request, the auxiliary process module will enter the listening mode and wait to receive the recovered data. Even if the data packets are disordered, the original data blocks can be re-aggregated from the received data by referring to the index field of the custom protocol, and the distributed file system is notified after the data blocks are completely re-aggregated.
5. The in-network aggregation erasure code recovery system based on the UDP protocol according to claim 1, characterized in that, The distributed file system is provided in both the target node and the sender node, and is used to cooperate in controlling the encoding of data and the distribution of data blocks between nodes. When data loss is detected, the distributed file system will pass the recovery request to the auxiliary process to start the recovery process.
6. The in-network aggregation erasure code recovery system based on the UDP protocol according to claim 4, wherein The recovery request includes: the IP addresses of all sender nodes participating in the recovery and the names of the data blocks to be sent. The auxiliary process module sends a recovery request to all sender nodes according to the recovery request, and the sender process module responds to the recovery request by reading the data blocks to be sent and sending them to the programmable switching system.
7. A UDP protocol-based in-network aggregation erasure code recovery method for the system according to any one of claims 1-6, characterized in that, including: 1) When the target node detects data loss, the distributed file system of the node will pass the recovery request to its auxiliary process module; 2) The auxiliary process module sends a recovery request to all sender nodes according to the information included in the recovery request; 3) All sender process modules that receive the recovery request read the data blocks from the local operating system according to the recovery request and send the data blocks to the programmable switching system according to the custom protocol; 4) After aggregating the data bytes in the packets to be processed from the sender nodes, the programmable switching system forwards the aggregated data to the target node; 5) The auxiliary process module of the target node re-aggregates the lost data blocks according to the received aggregated data and notifies the distributed file system that the data blocks have been recovered.
8. The method for in-network aggregation erasure code recovery based on the UDP protocol according to claim 7, characterized in that, The aggregation process mentioned above means that the data packets of different data blocks are aggregated according to their indexes, and each index is associated with a register and a bitmap on the switch; when the programmable switching system receives all the data of the blocks with a given index, the result is sent to the target node, otherwise, the calculated result is stored at the position identified by the index field of its register, and the UDP checksum of the data packet is updated by the checksum module before sending, and then the data packet is forwarded to the correct network address through the routing table.
9. The method for in-network aggregation erasure code recovery based on the UDP protocol according to claim 8, wherein The coefficient multiplication in the aggregation process refers to: multiplication on the Galois field related to the RS code, where all multiplication results are pre-computed and written to the memory when the programmable switching system starts up.
10. The in-network aggregation erasure code recovery method based on the UDP protocol according to claim 8, wherein The summation operation in the aggregation process mentioned above is the binary exclusive OR operation between two values, where: the first value is the result of coefficient multiplication, and the second value is read from the switch register at the given index.
Citation Information
Patent Citations
Erasure code repairing method based on network computing, erasure code updating method based on network computing and system
CN110190926A