A distributed storage system replica replication and multicasting method

Through the multicast RDMA technology on the Internet, programmable switches and multicast daemon work together, the problem of data redundancy and excessive CPU overhead in multi-replica replication in distributed storage systems is solved, and more efficient data replication and lower performance overhead are achieved.

CN115834263BActive Publication Date: 2025-05-06NANJING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211457213.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-21
Publication Date
2025-05-06
Estimated Expiration
2042-11-21

AI Technical Summary

Technical Problem

The existing distributed storage system based on RDMA has problems with data redundancy and separate replication status during the multi-replica replication process, resulting in wasted network bandwidth and excessive CPU overhead at the sending end, which has become a performance bottleneck.

Method used

The multicast RDMA is designed on the Internet through the multicast daemon working in concert with the programmable switch, and the matching action table is configured to realize the replication of data packets and the aggregation of confirmation packets, and the multicast replication position algorithm is optimized to reduce bandwidth waste.

Benefits of technology

It effectively reduces network bandwidth waste and sending CPU overhead during multi-replica replication, improves the throughput of multi-replica replication and the overall performance of distributed storage systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115834263B_ABST
    Figure CN115834263B_ABST
Patent Text Reader

Abstract

The present invention provides a distributed storage system replica replication on-line multicast method, including a server as a data sending end and a receiving end and a handshake of a multicast daemon running on a certain server; the multicast daemon configures a matching action table to a programmable switch through a control plane interface, and the control plane sends the multicast replication and aggregation table items to the programmable switch Tofino chip; the sending end server sends a data message to the receiving end server, and the programmable switch configured with a multicast group copies the message to multiple outbound ports, and modifies the corresponding fields of the message in the outbound ports; the receiving end server receives the data and sends a confirmation message, and the programmable switch performs an aggregation operation on the confirmation message from the receiving end server after receiving it, and sends it back to the sending end server after aggregation. The present invention effectively saves bandwidth in the server and the network, reduces server CPU overhead, and improves the throughput of distributed storage replica replication.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of distributed computing systems, and in particular relates to a distributed storage system replica replication and network multicasting method. Background Art

[0002] Distributed storage system, as a new storage form, is widely used in private and public data centers. In distributed storage system, multi-copy backup is the most basic backup strategy, which is to store the same data on multiple different servers to achieve data security, reliability and high availability.

[0003] Remote Direct Memory Access (RDMA), as a high-performance network technology, is widely used in distributed storage systems. RDMA provides a method to directly access remote non-volatile memory (NVM) without going through the target host's operating system to copy and modify data, achieving low latency and high throughput.

[0004] However, remote procedure calls only support unicast operations, and RDMA-based data replication has two problems: data redundancy and separate replication status. On the one hand, the redundancy of multiple replicated data causes a huge waste of network bandwidth, and on the other hand, the separate replication status forces the sender's CPU to handle multiple replica receiver connections at the same time, including sending requests, polling completion events, and keeping track of the sending status of each connection, resulting in several times of CPU overhead and becoming a performance bottleneck. Summary of the invention

[0005] The present invention aims at the deficiencies in the prior art and provides a distributed storage system replica replication and multicasting method.

[0006] The present invention provides a distributed storage system replica replication and network multicasting method, comprising:

[0007] Any server in the multicast group is used as a sender, and the remaining servers are used as receivers, so as to obtain a sender server and multiple receiver servers; the sender server and multiple receiver servers are connected to the server running the multicast daemon process in the multicast group through RDMA.

[0008] The multicast daemon sends a matching action table for the multicast sending phase and a matching action table for the confirmation message aggregation phase to the programmable switch; the control plane of the programmable switch sends the matching action table for the multicast sending phase and the confirmation message aggregation phase to the Tofino chip on the programmable switch; the matching action table includes the IP address of the sending end server, the virtual destination queue pair number of the sending end server, the destination IP address of each receiving end server, and the queue pair number of each receiving end server;

[0009] In the multicast transmission phase, the transmitting server sends a data message to the receiving server; the programmable switch configured with the matching action table copies the data message to multiple egress ports, and modifies the destination IP address and queue number field of the data message in the multiple egress ports;

[0010] In the multicast confirmation message aggregation stage, the receiving server receives the data message from the sending server and sends a confirmation message to the sending server; the confirmation message arrives at the programmable switch directly connected to the sending server; the programmable switch directly connected to the sending server configures the matching action table required for the confirmation message aggregation stage, performs an aggregation operation on the confirmation message after receiving the confirmation message from the receiving server, and sends the aggregated confirmation message back to the sending server; wherein, the aggregation operation is to use the matching action table to search the multicast group ID and the offset number of the receiving server in the multicast group according to the source IP address and the destination queue pair address; and read and modify the confirmation sequence number of the confirmation message according to the multicast group ID and the offset number of the receiving server in the multicast group.

[0011] Furthermore, in the multicast sending stage, the sending end server sends a data message to the receiving end server; the programmable switch configured with the matching action table copies the data message to multiple egress ports, and modifies the destination IP address and queue number field of the data message in the multiple egress ports, including:

[0012] After receiving the message at the inbound port, the programmable switch matches the action table in the multicast sending phase, matches a multicast group ID according to the IP address of the sending server and the virtual destination queue number of the sending server, and copies the message to the port where the destination IP of all receiving servers is located;

[0013] The programmable switch searches the multicast group ID and the copied offset number in the egress port according to the matching action table to find the corresponding destination IP address and queue pair number, and modifies the destination IP address and queue pair number fields of the message to the found destination IP address and queue pair number.

[0014] Furthermore, each programmable switch checks the receiving server and routing table in the corresponding multicast group in turn; if the outbound ports of two receiving servers in the multicast group are inconsistent in the routing table, a copy is performed on the corresponding programmable switch to obtain all the locations for copying a multicast group.

[0015] Furthermore, the control plane of the programmable switch specifies the offset number of the copy in the multicast group when configuring the multicast group; after the transmission manager of the Tofino chip of the programmable switch copies the message, it checks the offset number of the message at the output port of the programmable switch, modifies the message field according to the multicast group ID and the offset number, and recalculates the checksum.

[0016] Furthermore, the programmable switch matches the multicast group ID and the offset number within the multicast group according to the source IP and destination queue numbers of the confirmation message; the register of the programmable switch is used to save the currently confirmed sequence number of each receiving server of each multicast group; for each confirmation message, the programmable switch updates the confirmed sequence number of the current receiving server register and reads the confirmed sequence numbers of the remaining receiving server registers; the minimum value of the confirmed sequence number is the sequence number of the confirmation message currently sent to the sending server, indicating that all receiving servers have received the completed sequence number.

[0017] Furthermore, the SALU calculation unit of the programmable switch is used to simultaneously read and write a register storing a confirmed sequence number of a multicast group, write a new value, and read out the old value; if the read old value is equal to the current new value, the confirmation message corresponding to the confirmation sequence number is discarded.

[0018] The present invention provides a distributed storage system replica replication in-network multicast method, comprising: a server as a data sending end and a receiving end shakes hands with a multicast daemon process running on a server to exchange queue pair information; the multicast daemon process configures a matching action table to a programmable switch through a control plane interface, and the control plane sends the multicast replication and aggregation table items to a Tofino chip of the programmable switch; the sending end server sends a data message to the receiving end server, and the programmable switch configured with a multicast group copies the message to multiple output ports, and modifies the corresponding fields of the message in the output ports; the receiving end server receives the data and sends a confirmation message, and the programmable switch performs an aggregation operation on the confirmation message after receiving the confirmation message from the receiving end server, and sends the aggregation back to the sending end server.

[0019] The present invention introduces a programmable switch to design a reliable on-net multicast RDMA, which realizes the transparency of the peer network card and compatibility with the existing RDMA reliability and congestion control mechanism. On-net multicast RDMA improves the bandwidth waste in the network and the CPU overhead of the server on the end during multi-copy replication, and improves the throughput of multi-copy replication and the performance of multi-copy backup of distributed storage systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solution of the present invention, the drawings required for use in the embodiments are briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0021] Figure 1 A flowchart of a distributed storage system replica multicast method provided by an embodiment of the present invention;

[0022] Figure 2 MC-RDMA protocol architecture diagram provided for an embodiment of the present invention;

[0023] Figure 3 A schematic diagram of bandwidth waste in a distributed storage system provided by an embodiment of the present invention;

[0024] Figure 4 A processing flow structure diagram of a programmable switch provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0025] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0026] like Figure 1 As shown, an embodiment of the present invention provides a distributed storage system replica replication in-network multicast method, comprising:

[0027] Step 101, use any server in the multicast group as a sender, and the remaining servers as receivers, to obtain a sender server and multiple receiver servers; establish RDMA connections between the sender server and the multiple receiver servers and the server running the multicast daemon process in the multicast group.

[0028] Before sending data, multicast group members need to establish a connection, that is, exchange information through handshake. In a multicast group, there is a server as the data sender and the remaining servers as the data receivers. The servers as data senders and receivers establish an RDMA connection with the multicast daemon running on one (it can be any one) server. The connection needs to exchange queue pair (QP) information through handshake, including IP address, QP number of each server and initial sequence number. Optionally, the process is that the sending server and the receiving server send their own IP address and QP number to the multicast daemon. The multicast daemon sends the virtual QP number of the multicast group and the IP address of the first receiving server received to the sending server, and sends the received sending server address and sending server QP number to the receiving server. The multicast daemon sends this information to each server to establish an RDMA connection. At the same time, this information is used to sort out the information required for the match-action table of the programmable switch multicast sending stage: the sending server IP address, the sending server virtual destination QP number, and the destination IP address of each receiving server, and the QP number of each receiving server. In this way, the programmable switch can match the multicast group according to the message sent by the sending server, and then perform the operation of copying the message. In addition, it is necessary to generate the information required for the match-action table of the multicast confirmation message aggregation stage: the key used for matching: the destination IP address of each receiving server with the Nth receiving server IP address and the sending server QP number as parameters, and automatically generate the sequence number (1, 2, 3, ..., m) and the offset number N in increasing order.

[0029] Step 102, the multicast daemon sends a matching action table for the multicast sending phase and a matching action table for the confirmation message aggregation phase to the programmable switch; the control plane of the programmable switch sends the matching action table for the multicast sending phase and the confirmation message aggregation phase to the Tofino chip on the programmable switch; the matching action table includes the sending server IP address, the sending server virtual destination queue pair number, the destination IP address of each receiving server, and the queue pair number of each receiving server.

[0030] After the RDMA connection is established, the multicast daemon needs to configure the multicast group information to the programmable switch. This is completed by the multicast group management module of the multicast daemon. The multicast group management module must first calculate the fork location. The latest fork location is the best, that is, the "lazy fork strategy". The later the message is copied, the more effectively the bandwidth can be saved. The multicast replication location algorithm is an algorithm based on a greedy strategy. Its content is: each programmable switch checks the receiving end in its own multicast group and its own routing table in turn. If the outgoing ports of the two receiving end servers of the multicast group are inconsistent in the routing table, a replication is required on this programmable switch, otherwise it is handed over to the next programmable switch to continue the same operation. Finally, all the locations that need to be replicated for a multicast group are obtained. After obtaining the location that needs to be replicated, the multicast group management module calls the control plane interface located on the programmable switch, such as Figure 4 As shown in FIG. 1 , configuration information is sent to the programmable switch, including multicast group matching in the sending phase and the confirmation phase, multicast group IP address and QP number, etc. For RDMA unilateral operation, additional memory offset information needs to be sent.

[0031] For the multicast sending stage, the programmable switch ingress port needs to be configured with a matching action table for searching the multicast group ID. The key is the sending server IP address and the virtual destination QP number of the sending server, and the parameter is the multicast group ID. The packet replication engine (PRE) of the programmable switch traffic management (TM) component needs to be configured to replicate to multiple ports according to the multicast group ID. The port configured here is the port where all the receiving servers of this multicast group are located. Each port has a replication offset number, which represents the number of messages that are replicated in this multicast. The programmable switch egress port needs to be configured with a matching action table for searching the IP address and QP number of the receiving server according to the multicast group ID and the replication offset number. Here, the destination IP address and QP number of the replicated message are modified to the IP address and QP number of the matching receiving server. For the multicast confirmation message aggregation stage, the programmable switch egress port needs to be configured with a matching action table for searching the multicast group ID and the offset number of the receiving server in the multicast group according to the source IP address and the destination QP address. This matching action table is used to perform aggregation operations in the confirmation message aggregation stage. Among them, the programmable switch that needs to be configured with a matching action table in the multicast sending phase is calculated by the multicast replication location algorithm. The programmable switch that needs to be configured with a matching action table in the multicast confirmation message aggregation phase is the switch directly connected to the sending server, which is usually called a top of rack (ToR) switch in the data center.

[0032] Step 103, in the multicast sending phase, the sending server sends a data message to the receiving server; the programmable switch configured with the matching action table copies the data message to multiple egress ports, and modifies the destination IP address and queue number field of the data message in the multiple egress ports.

[0033] After receiving the message at the ingress port of the programmable switch, the action table is matched in the multicast sending stage. According to the IP address of the sending server and the virtual destination QP number of the sending server, a multicast group ID is matched to copy the message to the port where the destination IP of all receiving servers is configured. The programmable switch searches the multicast group ID and the copied offset number in the egress port according to the matching action table to find the corresponding destination IP address and QP number, and modifies the destination IP address and QP number fields of the message to the found destination IP address and QP number.

[0034] like Figure 4 As shown, in the multicast sending phase, the sending server sends a sending message to the RDMA network card, and the network card sends data to the corresponding multicast group QP. After the programmable switch receives the Infiniband message, if the correct multicast group is sent in step 102, it will match the corresponding multicast group, and then the Packets Replication Engine (PRE) in the Traffic Manager (TM) module of the programmable switch will copy the message, and the number of copies is obtained by the multicast replication position algorithm in step 102. The copied message directly arrives at the egress port of the programmable switch, and the message fields, including the QP number and the destination IP address, will be modified accordingly in the egress port to make them consistent with the corresponding receiving end QP information and meet the specifications of the RDMA network card, so that the receiving end server network card can receive it normally. For RDMA unilateral operation, it is also necessary to modify the memory address offset.

[0035] When configuring a multicast group, the control plane of the programmable switch specifies the offset number of the replica in the multicast group; after the transmission manager of the Tofino chip of the programmable switch copies the message, it checks the offset number of the message at the egress port of the programmable switch, modifies the message field according to the multicast group ID and the offset number, and recalculates the checksum field.

[0036] Use the DCQCN algorithm for congestion control. If congestion occurs, the ECN running on the programmable switch will mark the message. After the receiving end receives the message with the ECN mark, it will send a congestion notification packet (CNP) to the sending end to notify the sending end to reduce the speed. The programmable switch also needs to aggregate the CNP message, that is, change the source IP of the message to the multicast group IP to which the sending end establishes a connection.

[0037] Step 104, in the multicast confirmation message aggregation stage, the receiving server receives the data message from the sending server and sends a confirmation message to the sending server; the confirmation message arrives at the programmable switch directly connected to the sending server; the programmable switch directly connected to the sending server configures the matching action table required for the confirmation message aggregation stage, and performs an aggregation operation on the confirmation message after receiving the confirmation message from the receiving server, and the aggregated confirmation message is sent back to the sending server; wherein, the aggregation operation is to use the matching action table to find the multicast group ID and the offset number of the receiving server in the multicast group according to the source IP address and the destination queue pair address, and read and modify the confirmation sequence number of the confirmation message according to the multicast group ID and the offset number of the receiving server in the multicast group.

[0038] The programmable switch matches the multicast group ID and the offset number within the multicast group according to the source IP and destination queue numbers of the confirmation message; the programmable switch register is used to save the currently confirmed sequence number of each receiving server of each multicast group; for each confirmation message, the programmable switch updates the confirmed sequence number of the current receiving server register and reads the confirmed sequence numbers of the remaining receiving server registers; the minimum value of the confirmed sequence number is the sequence number of the confirmation message currently sent to the sending server, indicating that all receiving servers have received the completed sequence number.

[0039] There is a memory resource in the programmable switch to store the currently confirmed sequence number of each receiving end server of the multicast group. Each storage unit is called a register. The register is addressed by the multicast group ID and the offset number of the receiving end server in the multicast group. The received confirmation message writes to its own register, modifies the content of the sequence number field of this message, reads the registers of other receiving ends of this multicast group ID, and compares them with the size of its own confirmation sequence number. The smallest sequence number is retained for each comparison, and finally the sequence number field of the current confirmation message is modified to the smallest sequence number, indicating the sequence number that all current receiving ends have confirmed. For example, a multicast group with an ID of 101 has three receiving ends, with offsets of 0, 1, and 2 respectively. The current confirmation message belongs to the second one, that is, the offset number is 1. The operation is to write to the register with an offset of 1 for the multicast group with an ID of 101, modify the content therein, read the registers with offsets of 0 and 2 for the multicast group with an ID of 101, and modify the sequence number field of the confirmation message to the smallest value among them.

[0040] Then, to prevent the sending server from receiving duplicate confirmation messages, the programmable switch performs a deduplication operation. The deduplication operation also requires a similar register, which indicates the confirmation message sequence number of the multicast group ID that has been sent to the sending server. The register is only addressed by the multicast group ID. After completing the operation of modifying the message sequence number in the previous step, the SALU calculation unit of the programmable switch is used to read and write the register that stores the sequence number that has been confirmed by a multicast group at one time, write a new value, and read the old value, so as to break the limitation that a resource of the programmable switch can only be accessed once. If the new value is equal to the old value, it means that the confirmation message has been sent to the sending server and is a duplicate. The current message is discarded and the sending server will not receive a duplicate confirmation message; if the new value is not equal to the old value, it means that the confirmation message has not been sent to the sending server and is new. The current message is retained, and the aggregation operation is completed. The aggregated confirmation message is sent back to the sending server.

[0041] For NAK messages, since NAK messages indicate that all messages before the sequence number carried in the NAK message have been confirmed, the receiving server expects to receive data messages with this sequence number. Therefore, during the aggregation phase of the programmable switch, the NAK message sequence number needs to be minus one to update the register, and the NAK sequence number is updated to the confirmed minimum value plus one. In this way, for the sending server, the confirmed message will not be NAKed, ensuring correctness. After receiving the confirmation message, the sending server will continue to send data, and in this way, the application data can be sent in a cycle, thereby completing multi-copy backup.

[0042] like Figure 2 As shown in the figure, there are some components on the host and programmable switch respectively. Among them, the main functions of the host include: establishing RDMA multicast connections, sending data, reliability guarantee and congestion control; the main functions of the programmable switch include copying data messages in the sending phase and aggregating confirmation messages in the confirmation phase. In order to realize these functions, the host and programmable switch each have separate modules to be responsible for the corresponding functions.

[0043] like Figure 3 As shown, Figure 3 The figure shows the data flow pattern of two traditional replication methods, direct replication and chain replication, using three replicas as an example and adopting the Y-type write strategy. The Y-type write means that the client first writes the data to the master replica, and the master replica then sends the same data to other replicas to achieve backup. Direct replication means that the server where the master replica is located directly establishes RDMA connections with the two slave replicas. This method brings two problems. The first is bandwidth waste. Figure 3It can be seen that one of the traffic sent from the master replica to the slave replica is duplicated. If the topology is larger and the number of hops is greater, there will be more duplicate traffic data; when congestion occurs, the waste of bandwidth will greatly affect the throughput. In addition, this also limits the export bandwidth of the sending server. For example, if the hardware export of the sending network card has 100Gbps, when sending to two replicas, the throughput will be limited to 50Gbps, and more replicas will be limited to a smaller bandwidth. The second is the CPU overhead of the sending server. The master replica establishes connections with the two slave replicas respectively, which will cause the sending server CPU to send sending events and polling completion events, and the overhead of processing data increases exponentially.

[0044] Chain replication is a process where a master copy establishes a connection with a slave copy, and then the slave copies establish connections in sequence. The master copy sends data to the first slave copy, and the first slave copy immediately sends the same data to the next slave copy after receiving the data. The last slave copy sends a confirmation message to the sender. This method can ensure that the throughput of the sender is not limited by the number of copies, but Figure 3 As can be seen from the figure, there is still extra bandwidth waste in the network. In addition, chain replication also brings extra latency because the replication process is serial. The more replicas there are, the greater the extra latency. In-network RDMA multicast implemented with programmable switches effectively solves this problem. Since the replication point is on the switch in the network, bandwidth is saved and there is no extra latency like chain replication. In addition, since the CPU does not need to maintain multiple connections, the CPU overhead can also be effectively saved.

[0045] The present invention has been described in detail above in conjunction with specific implementations and exemplary examples, but these descriptions cannot be understood as limiting the present invention. Those skilled in the art understand that, without departing from the spirit and scope of the present invention, a variety of equivalent substitutions, modifications or improvements may be made to the technical solution of the present invention and its implementation methods, all of which fall within the scope of the present invention. The scope of protection of the present invention shall be subject to the attached claims.

Claims

1. A distributed storage system replica replication in-network multicast method, characterized in that: include: Any server in the multicast group is used as a sender, and the remaining servers are used as receivers, so that a sender server and multiple receiver servers are obtained; Establish RDMA connections between the sending server and multiple receiving servers and the server running the multicast daemon in the multicast group; The multicast daemon sends a matching action table for the multicast sending phase and a matching action table for the confirmation message aggregation phase to the programmable switch; the control plane of the programmable switch sends the matching action table for the multicast sending phase and the confirmation message aggregation phase to the Tofino chip on the programmable switch; the matching action table includes the IP address of the sending end server, the virtual destination queue pair number of the sending end server, the destination IP address of each receiving end server, and the queue pair number of each receiving end server; In the multicast transmission phase, the transmitting server sends a data message to the receiving server; the programmable switch configured with the matching action table copies the data message to multiple egress ports, and modifies the destination IP address and queue number field of the data message in the multiple egress ports; In the multicast confirmation message aggregation stage, the receiving server receives the data message from the sending server and sends a confirmation message to the sending server; the confirmation message arrives at the programmable switch directly connected to the sending server; the programmable switch directly connected to the sending server configures the matching action table required for the confirmation message aggregation stage, performs an aggregation operation on the confirmation message after receiving the confirmation message from the receiving server, and sends the aggregated confirmation message back to the sending server; wherein, the aggregation operation is to use the matching action table to search the multicast group ID and the offset number of the receiving server in the multicast group according to the source IP address and the destination queue pair address; and read and modify the confirmation sequence number of the confirmation message according to the multicast group ID and the offset number of the receiving server in the multicast group.

2. The distributed storage system replica replication in-network multicast method according to claim 1, characterized in that: In the multicast sending stage, the sending end server sends a data message to the receiving end server; the programmable switch configured with the matching action table copies the data message to multiple outbound ports, and modifies the destination IP address and queue number field of the data message in the multiple outbound ports, including: After receiving the message at the inbound port, the programmable switch matches the action table in the multicast sending phase, matches a multicast group ID according to the IP address of the sending server and the virtual destination queue number of the sending server, and copies the message to the port where the destination IP of all receiving servers is located; The programmable switch searches the multicast group ID and the copied offset number in the egress port according to the matching action table to find the corresponding destination IP address and queue pair number, and modifies the destination IP address and queue pair number fields of the message to the found destination IP address and queue pair number.

3. The distributed storage system replica replication in-network multicast method according to claim 1, characterized in that: Each programmable switch checks the receiving server and routing table in the corresponding multicast group in turn; if the outgoing ports of two receiving servers in the multicast group are inconsistent in the routing table, a copy is made on the corresponding programmable switch to obtain all the locations for the copy of a multicast group.

4. The distributed storage system replica replication in-network multicast method according to claim 1, characterized in that: When configuring a multicast group, the control plane of the programmable switch specifies the offset number of the replica in the multicast group; after the transmission manager of the Tofino chip of the programmable switch copies the message, it checks the offset number of the message at the egress port of the programmable switch, modifies the message field according to the multicast group ID and the offset number, and recalculates the checksum.

5. The distributed storage system replica replication in-network multicast method according to claim 1, characterized in that: The programmable switch matches the multicast group ID and the offset number within the multicast group according to the source IP and destination queue number of the confirmation message; the register of the programmable switch is used to save the sequence number currently confirmed by each receiving end server of each multicast group; For each confirmation message, the programmable switch updates the confirmed sequence number of the current receiving end server register and reads the confirmed sequence numbers of the remaining receiving end server registers; The minimum value of the confirmed sequence number is the sequence number of the confirmation message currently sent to the sending server, indicating the sequence number that all receiving servers have received.

6. The distributed storage system replica replication in-network multicast method according to claim 1, characterized in that: The SALU computing unit of the programmable switch is used to simultaneously read and write the register that stores the confirmed sequence number of a multicast group, write a new value, and read out the old value; if the old value read out is equal to the current new value, the confirmation message corresponding to the confirmation sequence number is discarded.

Citation Information

Patent Citations

  • Distributed data storage

    CN102301367A

  • Dynamic multistage flow control method based on programmable switching chip

    CN112637090A