Data center cross-domain link disaster recovery backup method based on DPU
By adopting a DPU-based dynamic IP aggregation-hierarchical encryption sharding-intelligent flow reorganization collaborative architecture in the data center network, the problems of sharding and encryption separation, flow state explosion and handover response hysteresis in cross-domain link disaster recovery backup are solved, and high-reliability and low-latency transmission is achieved, and through intelligent handover strategies, core services are migrated to the optimal path without interruption.
Patent Information
- Application Number
- CN202510424199.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-06-20
AI Technical Summary
In the cross-domain link disaster recovery backup, existing data center networks have bottlenecks such as sharding and encryption separation, flow state explosion and handover response hysteresis, which are difficult to meet the transmission needs of high reliability and low latency.
Adopting a dynamic IP aggregation-hierarchical encryption sharding-intelligent stream reorganization collaborative architecture based on DPU, the GSR of the gateway service router is built through the DPU's programmable data plane and hardware acceleration engine, and the gateway service router GSR is realized to achieve end-to-end encrypted transmission, and through bidirectional network state perception and dynamic health assessment, efficient switching of the main and backup links is achieved.
It effectively solves the problems of high delay in shard restructuring, limited encryption throughput and explosive streaming state, and realizes high-reliability and low-latency transmission of cross-domain links, and ensures that core services are migrated to the optimal path without interruption through intelligent handover strategies.
Smart Images

Figure CN120186071A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data center networks, and in particular to a cross-domain link disaster tolerance and backup method for data centers based on DPU. Background Art
[0002] Modern data centers have evolved from simple data storage facilities to the core computing power hubs that support cloud computing and artificial intelligence. With the rapid increase in the scale of cross-domain long-haul links, traditional disaster tolerance and backup solutions face two major contradictions: the soaring cost of dedicated line bandwidth and the insufficient reliability of public network links. The current mainstream solutions still rely on software sharding or synchronous mirroring technologies, and there are bottlenecks such as high shard recombination latency and limited encryption throughput.
[0003] As the third major computing power carrier in the data center, the data processing unit DPU has achieved breakthrough performance with a shard recombination latency <50 μs and an encryption throughput of 200 Gbps by virtue of its programmable data plane and heterogeneous acceleration engine, greatly improving the efficiency compared with traditional CPU solutions. Its core capabilities include: releasing CPU / GPU computing power by hardware-accelerating network virtualization, storage acceleration, and security encryption tasks; dynamically triggering route switching based on link health, and supporting seamless adaptation of protocols such as EBGP; hardware-level multi-tenant isolation and zero-trust architecture to block the lateral spread of attacks. Although DPU technology has significantly improved shard recombination and encryption efficiency, there are still three major unsolved problems in existing solutions: 1) Flow state explosion: The mapping of millions of IP flows leads to an exponential increase in the storage overhead of the TCAM flow table, restricting large-scale deployment. 2) Cross-layer protocol conflict: The encryption header of the encryption protocol IPSec and the bit width of the shard identification field preempt, resulting in failed checksums for some sharded packets. 3) Insufficient multi-scenario collaboration ability: The current technology focuses on disaster tolerance within the data center and lacks the deep collaboration ability with cross-domain long-haul links and edge computing. Summary of the Invention
[0004] The purpose of the present invention is to overcome the disadvantages and deficiencies of the prior art, and propose a cross-domain link disaster tolerance and backup method for data centers based on DPU, which breaks through the core bottlenecks such as the separation of sharding and encryption, flow state explosion, and slow switching response in traditional solutions, deeply integrates the programmable data plane and hardware acceleration engine of DPU, and constructs a "dynamic IP aggregation - hierarchical encryption sharding - intelligent flow recombination" collaborative architecture to meet the high-reliability and low-latency transmission requirements in cross-domain connection disaster tolerance scenarios.
[0005] To achieve the above purpose, the technical solution provided by the present invention is: A cross-domain link disaster tolerance and backup method for data centers based on DPU, comprising the following steps:
[0006] 1) Build a gateway service router GSR using a switch and a DPU, deploy it under the border router CBR, and use it as the core processing node of the backup link. Set the backup link route and use the autonomous system path prepending (AS-PathPrepending) technology to adjust the backup link priority.
[0007] 2) Fragment the data packets of the traffic passing through the DPU based on the ARM core of the DPU to meet the maximum transmission unit (MTU) limit in the public network transmission. At the same time, to meet the memory limit of the DPU, add controllable IP pairs to reduce the number of flows during the fragmentation and reassembly process.
[0008] 3) Use the GSR built in step 1) to implement end-to-end encrypted transmission of the backup link traffic through an encryption engine.
[0009] 4) Dynamically sense network anomalies by real-time monitoring the public network link quality and service transmission status. Based on the preset hierarchical disaster tolerance policy, prioritize ensuring that the core service traffic quickly switches to a low-latency and highly reliable dedicated line link, while non-critical traffic is automatically degraded to the backup path, achieving millisecond-level seamless switching.
[0010] Furthermore, in step 1), build a gateway service router GSR based on a switch and a DPU, deploy it under the public network border router CBR of the data center, and use it as the core processing node of the cross-domain backup link. GSR is composed of a multi-port high-speed switch and a DPU. The DPU adopts a heterogeneous computing architecture, integrating an ARM multi-core control plane and a P4 programmable data plane, and realizes high-speed interconnection between the switch and the DPU through a PCIe Gen5 interface.
[0011] For one area in two data centers, there is a core service router CSR in the area. CSR is connected to the servers in the data center and externally connected to a border router BSR. BSR is connected to BSRs in other areas through a cross-domain dedicated line. CSR is also connected to the public network border router CBR, and CBR is connected to the public network ISP to access the public network.
[0012] Between the two areas, GSR publishes its public IP address to the public network ISP through CBR and establishes an EBGP neighbor relationship with the connected area. The GSR in this area learns the internal routes of this area from CBR and transmits them to the GSRs in other areas through EBGP. After receiving the transmitted routes, the GSRs in other areas need to transmit them to CBR and perform AS-path override to ensure that the public network transmission under normal circumstances directly goes through CBR instead of choosing the GSR backup path, thus forming a routing channel for the cross-region backup link.
[0013] For the routing policy of the DPU, the internal routing daemon of the DPU is responsible for writing the learned routing entries of other regions into an independent routing table in the Linux kernel. This routing table is designed to store only the encrypted tunnel routes required for cross-DPU communication. Through the kernel namespace isolation mechanism, conflicts with the regular routing table are avoided, ensuring the accurate classification and directional forwarding of service traffic. The business process pds_dp_app listens for changes in the routing table in real time by subscribing to the kernel network link socket netlink socket. When a routing entry is added or deleted, the routing target network, next-hop address, and outgoing interface information carried in the routing message are parsed, and the routing rules are sent to the hardware forwarding plane by calling the dedicated hardware programming interface of the DPU. During this process, each route in the control plane is dynamically bound with an IPSec encryption policy, and the open-source routing management component fpmsyncd is customized. The modified fpmsyncd process only filters the routing update events of the underlay network and synchronizes such routes to other DPU nodes in the cluster through the Redis database message queue function.
[0014] Further, in step 2), large packets existing in the data center traffic are fragmented and reorganized for transmission in the public network. To suppress the explosion of the number of fragmented flows, a pair of controllable IP addresses is dynamically added to the outer layer of the original packet to avoid the problem of flow state explosion, forming an extended data unit: [Controllable IP header] + [Original IP header] + [Original data payload]. The controllable IP pair is generated through a preset rule or hash algorithm: controllable IP pair = H(source IP ⊕ destination IP) mod N, where H is a hash function and N is the scale of the controllable IP address pool. Through this mapping, millions of real IP flows are compressed into a limited controllable IP space. After adding the controllable IP pair, fragmentation and encryption operations are performed. The extended data unit [Controllable IP header] + [Original IP header] + [Original data payload] is fragmented according to the MTU limit, and the fragmentation size meets the MTU. Each fragment is independently encrypted in IPSec tunnel mode, and the encryption range includes the controllable IP header, the original IP header, and the data fragment payload, generating an ESP encapsulation structure: [IPSec ESP header] + [Encrypted controllable IP header] + [Encrypted original IP header fragment] + [Encrypted data fragment payload]. Then, a UDP header and a public network IP header are added to form the final transmission unit: [Public network IP header] + [UDP header] + [IPSec ESP header] + [Encrypted controllable IP header] + [Encrypted original IP header fragment] + [Encrypted data fragment payload];
[0015] For the receiving end, extract IPSec ESP data according to the public network IP header and UDP header, decrypt it to obtain the controllable IP header pair, use this as the flow identifier instead of the original IP header fragment, maintain the reassembly context in the DPU memory with the controllable IP header pair as the key. After all the shards arrive, splice the controllable IP header pair, the original IP header and the data payload according to the offset. After verifying the integrity, strip the controllable IP header pair and restore the original data packet.
[0016] Further, in step 3), based on the backup link constructed in step 1), it is necessary to encrypt and transmit the traffic on the backup link. The specific steps are as follows:
[0017] Packet reception and parsing: The IP packet enters the DPU via CBR. The DPU extracts the Ethernet frame header, IP header and transport layer information through the hardware parser; the matching processing unit MPU matches the source IP, destination IP, protocol, source port, destination port five-tuple according to the pre-programmed P4 rules. If the encryption policy is hit, look up the corresponding security policy SA entry through the content addressable memory CAM to obtain the key material, security parameter index SPI and anti-replay window parameters;
[0018] Payload segmentation and padding: Split the IP packet according to the selected encryption algorithm AES-GCM-256 block size. Generate a 12-byte random initialization vector IV by the random number generator, and combine it with the 32-bit sequence number to construct the complete counter initial value J0:
[0019] J0 = IV||0 31 1
[0020] AES-GCM encryption: The encryption engine executes the following atomic operations through the full pipeline architecture: for each 16-byte plaintext block P i Generate the ciphertext block C i :
[0021]
[0022] where CTR i = J0 + i mod 2 128 , i is the block number, K is the key of AES-256; then calculate the integrity check value of the payload in parallel through the GMAC authentication algorithm;
[0023] ESP encapsulation and reassembly: The encrypted data is reassembled into an ESP packet: add an ESP header, including SPI + sequence number, and an ESP tail, including padding length + next header type: where the SPI size is 4 bytes, which is used to identify the security association; the sequence number size is 4 bytes, which is used for anti-replay attack; at the same time, append the 16-byte ICV to the end of the packet; finally, modify the original IP protocol field to 50 to identify ESP, and recalculate the IP header checksum;
[0024] Decryption: For the received packet, first perform the fragmentation assembly operation, then search for the SA based on the destination IP address, ESP protocol, and SPI. Secondly, check the sequence number for replay attacks. Finally, perform the ICV checksum verification. Ultimately, according to the algorithm and key parameters specified in the SA, decrypt the encrypted part of the data, remove the padding part, and reconstruct the original IP packet.
[0025] Furthermore, in step 4), based on two-way network status awareness and dynamic health assessment, achieve efficient switching of the primary and backup links:
[0026] Two-way network status awareness synchronously collects key metrics of the local and peer cross-domain routers every 2 seconds to ensure the real-time and two-way nature of status awareness. Active probing of the link is achieved through two-way ICMP probing: The local router sends a 1500-byte probe packet to the peer, and the peer router synchronously sends a probe packet of the same specification in the reverse direction to form a two-way detection channel. 20 probe packets are sent per cycle, and the two-way packet loss rate and round-trip delay RTT are statistically calculated. The packet loss rate calculation adopts the worst-case value principle, that is, take the larger value of the forward and reverse packet loss rates to ensure sensitivity to link degradation. At the same time, obtain the real-time traffic count of the local and peer router interfaces through the SNMPv3 protocol, and calculate the interface bandwidth utilization rate:
[0027]
[0028] In the formula, ifHCOutOctets is the cumulative number of bytes sent by the interface output by the SNMP counter, ifHCOutOctets last is the latest number of bytes sent, and the sampling interval is fixed at 2 seconds. At the same time, obtain in-depth metrics through sFlow sampling, including TCP retransmission rate and out-of-order packet rate, to assist in evaluating the link quality;
[0029] The health score S comprehensively reflects the link quality, and its calculation integrates four core metrics:
[0030]
[0031] In the formula, Local_Load and Remote_Load are the interface utilization rates of the local and peer routers respectively, with weights of 20% each, reflecting the link congestion degree; Loss_Rate is the maximum value of the two-way packet loss rate, with a weight of 30%, directly reflecting the transmission reliability; Measured_Delay is the actual value of the detected RTT, Delay_Threshold is the upper limit of the delay tolerated by the service, and the ratio reflects the delay health, with a weight of 30%;
[0032] When S < 0.7 lasts for 3 consecutive cycles, it is determined that the primary link enters the sub-healthy state, triggering the handover evaluation process. To suppress instantaneous jitter, the exponential weighted moving average EWMA is introduced in the score calculation for smoothing:
[0033] S smoothed = 0.7×S currecnt + 0.3×S previous
[0034] In the formula, S currecnt is the current link health score, and S previous is the average value of the previous link health scores. Only when the smoothed link health score S smoothed continues to be lower than the threshold, subsequent actions are triggered;
[0035] The handover strategy is divided into two modes: full-scale handover and partial traffic handover, which are dynamically selected based on the score and link status:
[0036] Full-scale handover: When S smoothed < 0.5 and the primary link packet loss rate ≥ 10% or the delay ≥ 2 times the threshold, it is determined that the primary link fails; rapid handover is achieved through BGP routing policy updates: remove the AS-Path Prepending modification of the backup link routing entry, increase its local priority parameter Local Preference to 250 (the default for the primary link is 200), and at the same time append 3 times of AS-Path Prepending to the primary link to reduce the priority. This process relies on the BGP route reflector, and the convergence time can be compressed within 5 seconds; if an SDN controller is deployed, traffic can be further forced to switch to the backup link interface through batch redirection of OpenFlow flow tables to ensure the priority recovery of critical services;
[0037] Partial traffic handover: When 0.5 ≤ S smoothed < 0.7 and the remaining bandwidth of the backup link ≥ 20%, the load balancing mode is started; the equal-cost multi-path method ECMP is used to split traffic according to service priorities: high-priority traffic such as database synchronization and transaction requests is retained on the primary link, and low-priority traffic such as log backup and software download is proportionally allocated to the backup link; weighted random early detection WRED is implemented on the backup link to randomly discard non-critical traffic based on the queue length and service category to avoid congestion deterioration.
[0038] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0039] 1. Through the DPU heterogeneous computing architecture and programmable plane, the hardware offloading of fragmentation recombination and encryption tasks is realized, reducing the dependence on network device resources.
[0040] 2. Compress the flow state scale using the controllable IP mapping technology to avoid the memory explosion problem during fragmentation and reassembly.
[0041] 3. Combine two-way network state awareness to achieve dynamic evaluation of link health. Through the progressive drainage mechanism, ensure the smooth and reliable switching process between the primary and backup links. And based on the hierarchical switching strategy of business priorities, prioritize ensuring that core services are migrated to the optimal path without interruption. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is the network topology diagram between data centers of the present invention.
[0043] Figure 2 It is the schematic diagram of GSR construction of the present invention.
[0044] Figure 3 It is the schematic diagram of the encryption and fragmentation / reassembly strategy of the present invention.
[0045] Figure 4 It is the schematic diagram of the traffic sending path in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] The present invention will be further described in detail below in conjunction with the embodiments and the drawings, but the embodiments of the present invention are not limited thereto.
[0047] As Figures 1 to 4 shown, this embodiment discloses a DPU-based cross-domain link disaster recovery and backup method for data centers, including the following steps:
[0048] 1) Use switches and DPU to construct a gateway service router GSR, and deploy it under the border router CBR as the core processing node of the backup link. Set the backup link routing, and use the autonomous system path prepending (AS-PathPrepending) technology to adjust the backup link priority, specifically as follows:
[0049] Construct a gateway service router GSR based on switches and DPU, and deploy it under the public network border router CBR of the data center as the core processing node of the cross-domain backup link. GSR is composed of a multi-port high-speed switch and DPU. Among them, DPU adopts a heterogeneous computing architecture, integrating an ARM multi-core control plane and a P4 programmable data plane, and realizes high-speed interconnection between the switch and DPU through a PCIe Gen5 interface;
[0050] For one area in two data centers, there is a core service router CSR in the area. CSR is connected to the servers in the data center and externally connected to a border router BSR. BSR is connected to the BSRs in other areas through cross-domain dedicated lines. CSR is also connected to the public network border router CBR, and CBR is connected to the public network ISP to access the public network;
[0051] Between two regions, the GSR publishes its public IP address to the public network ISP through CBR and establishes an EBGP neighbor relationship with the connected regions. The GSR within this region learns the internal routes of this region from CBR and forwards them to the GSRs in other regions through EBGP. After receiving the transmitted routes, the GSRs in other regions need to forward them to CBR and perform AS-path override to ensure that in normal cases, the public network transmission directly goes through CBR instead of choosing the GSR backup path, thus forming a routing channel for the cross-region backup link.
[0052] For the routing policy of the DPU, the internal routing daemon of the DPU is responsible for writing the learned route entries of other regions into an independent routing table in the Linux kernel. This routing table is designed to only store the encrypted tunnel routes required for cross-DPU communication. Through the kernel namespace isolation mechanism, conflicts with the regular routing table are avoided, thereby ensuring the accurate classification and directional forwarding of service traffic. The service process pds_dp_app listens for route table change events in real time by subscribing to the kernel network link socket netlink socket. When a route entry is added or deleted, the route target network, next-hop address, and outgoing interface information carried in the route message will be parsed, and the routing rules will be sent to the hardware forwarding plane by calling the dedicated hardware programming interface of the DPU. During this process, the control plane dynamically binds an IPSec encryption policy to each route, and the control plane has customized the open-source routing management component fpmsyncd. The modified fpmsyncd process only filters the route update events of the underlay network and synchronizes such routes to other DPU nodes in the cluster through the Redis database message queue function.
[0053] 2) Based on the ARM core of the DPU, packet fragmentation is performed on the traffic passing through the DPU to meet the maximum transmission unit MTU limit in public network transmission. At the same time, in order to meet the memory limit of the DPU, controllable IP pairs are added to reduce the number of flows during the fragmentation and reassembly process, as follows:
[0054] Fragment and reassemble large packets in the data center traffic for transmission over the public network. To suppress the explosion of the number of fragmented flows, a pair of controllable IP addresses is dynamically added to the outer layer of the original data packet to avoid the problem of flow state explosion, forming an extended data unit: [Controllable IP header] + [Original IP header] + [Original data payload]. The controllable IP pair is generated through a preset rule or a hash algorithm: controllable IP pair = H(source IP ⊕ destination IP) mod N, where H is a hash function and N is the scale of the controllable IP address pool. Through this mapping, millions of real IP flows are compressed into a limited controllable IP space. After adding the controllable IP pair, perform fragmentation and encryption operations. Fragment the extended data unit [Controllable IP header] + [Original IP header] + [Original data payload] according to the MTU limit, and the fragmentation size meets the MTU. Independently implement IPSec tunnel mode encryption for each fragment, and the encryption scope includes the controllable IP header, the original IP header, and the data fragment payload, generating an ESP encapsulation structure: [IPSec ESP header] + [Encrypted controllable IP header] + [Encrypted original IP header fragment] + [Encrypted data fragment payload]. Then add a UDP header and a public network IP header to form the final transmission unit: [Public network IP header] + [UDP header] + [IPSec ESP header] + [Encrypted controllable IP header] + [Encrypted original IP header fragment] + [Encrypted data fragment payload];
[0055] For the receiving end, extract the IPSec ESP data according to the public network IP header and the UDP header, and obtain the controllable IP header after decryption. Use this as the flow identifier instead of the original IP header fragment. Use the controllable IP pair as the key to maintain the reassembly context in the DPU memory. After all fragments arrive, splice the controllable IP header, the original IP header, and the data payload according to the offset. After verifying the integrity, strip the controllable IP header to restore the original data packet.
[0056] 3) On the basis of the backup link constructed in step 1), it is necessary to encrypt and transmit the traffic on the backup link. The specific steps are as follows:
[0057] Packet reception and parsing: The IP packet enters the DPU via CBR. The DPU extracts the Ethernet frame header, IP header, and transport layer information through the hardware parser; the matching processing unit MPU matches the five-tuple of source IP, destination IP, protocol, source port, and destination port according to the pre-programmed P4 rules. If the encryption policy is hit, look up the corresponding security policy SA entry through the content addressable memory CAM to obtain the key material, security parameter index SPI, and anti-replay window parameter;
[0058] Payload segmentation and padding: Segment the IP packet according to the selected encryption algorithm AES-GCM-256 block size. Generate a 12-byte random initialization vector IV by the random number generator, and combine it with a 32-bit sequence number to construct the complete counter initial value J0:
[0059] J0 = IV||0 31 1
[0060] AES - GCM Encryption: The encryption engine performs the following atomic operations through a fully pipelined architecture: for each 16 - byte plaintext block P i Generate ciphertext block C i :
[0061]
[0062] where CTR i = J0 + i mod 2 128 , i is the block number, K is the key of AES - 256; then parallelly calculate the integrity check value of the payload through the GMAC authentication algorithm;
[0063] ESP Encapsulation and Reassembly: The encrypted data is reassembled into an ESP packet: Add an ESP header, including SPI + sequence number, and an ESP trailer, including padding length + next - header type: where the SPI size is 4 bytes, which is used to identify the security association; the sequence number size is 4 bytes, which is used for anti - replay attack; at the same time, append a 16 - byte ICV to the end of the packet; finally, modify the original IP protocol field to 50 to identify ESP and recalculate the IP header checksum;
[0064] Decryption: For the received packet, first perform the fragmentation assembly operation, then find the SA according to the destination IP address, ESP protocol, and SPI, secondly check the sequence number for anti - replay attack, and finally verify the ICV check value; finally, decrypt the encrypted part of the data according to the algorithm and key parameters specified in the SA, and remove the padding part to reconstruct the original IP packet.
[0065] 4) By real - time monitoring the public network link quality and service transmission status, dynamically perceive network anomalies, and based on a preset hierarchical disaster - tolerance strategy, preferentially ensure that core service traffic quickly switches to a low - latency and highly reliable dedicated line, and non - critical traffic is automatically degraded to the backup path to achieve millisecond - level seamless switching; among them, based on two - way network state perception and dynamic health assessment, achieve efficient switching between the primary and backup links:
[0066] Bidirectional network status awareness takes 2 seconds as a cycle to synchronously collect key metrics of the local and peer cross-domain routers, ensuring the real-time and bidirectional nature of status awareness. Active probing of the link is achieved through bidirectional ICMP probing: the local router sends a 1500-byte probe packet to the peer, and the peer router synchronously sends a probe packet of the same specification in the reverse direction, forming a bidirectional detection channel. 20 probe packets are sent per cycle, and the bidirectional packet loss rate and round-trip delay RTT are statistically calculated. The packet loss rate is calculated using the worst-case value principle, that is, taking the larger value of the forward and reverse packet loss rates to ensure sensitivity to link degradation. At the same time, the real-time traffic count of the local and peer router interfaces is obtained through the SNMPv3 protocol, and the interface bandwidth utilization rate is calculated:
[0067]
[0068] where ifHCOutOctets is the cumulative number of bytes sent by the interface output by the SNMP counter, and ifHCOutOctets last is the latest number of bytes sent, and the sampling interval is fixed at 2 seconds; at the same time, deep metrics are obtained through sFlow sampling, including TCP retransmission rate and out-of-order packet rate, which are used to assist in evaluating the link quality;
[0069] The health score S comprehensively reflects the link quality, and its calculation integrates four core metrics:
[0070]
[0071] where Local_Load and Remote_Load are the interface utilization rates of the local and peer routers respectively, with weights of 20% each, reflecting the degree of link congestion; Loss_Rate is the maximum value of the bidirectional packet loss rate, with a weight of 30%, directly reflecting the transmission reliability; Measured_Delay is the actual value of the detected RTT, and Delay_Threshold is the upper limit of the delay tolerated by the service. The ratio reflects the delay health, with a weight of 30%;
[0072] When S < 0.7 for 3 consecutive cycles, it is determined that the primary link enters the sub-healthy state, and the handover evaluation process is triggered. To suppress instantaneous jitter, the exponential weighted moving average EWMA is introduced for smoothing in the score calculation:
[0073] S smoothed = 0.7×S currecnt + 0.3×S previous
[0074] where S currecnt is the current link health score, and S previous is the average value of the previous link health scores. Only when the smoothed link health score S smoothedWhen the value remains below the threshold, subsequent actions are triggered;
[0075] The switching strategy is divided into two modes: full switching and partial traffic switching, which are dynamically selected based on the score and link status:
[0076] Full switch: When S smoothed <0.5 and the packet loss rate of the main link is ≥10% or the delay is ≥2 times the threshold, the main link is determined to be failed; fast switching is achieved through BGP routing policy update: remove the AS-Path Prepending modification of the backup link routing entry, increase its local priority parameter Local Preference to 250, the default main link is 200, and add 3 AS-Path Prepending to the main link to reduce the priority. With the help of BGP routing reflector, the convergence time can be compressed to within 5 seconds; if the SDN controller is deployed, the OpenFlow flow table batch redirection can be further used to force the traffic to switch to the backup link interface to ensure priority recovery of key services;
[0077] Partial flow switching: when 0.5≤S smoothed <0.7 and the remaining bandwidth of the backup link is ≥20%, the load balancing mode is started; the equal-cost multi-path method ECMP is used to divert traffic according to business priority: high-priority traffic such as database synchronization and transaction requests is retained on the primary link, and low-priority traffic such as log backup and software download is proportionally allocated to the backup link; weighted random early detection WRED is implemented on the backup link, and non-critical traffic is randomly discarded based on queue length and business category to avoid worsening congestion.
[0078] The following is a practical example of the DPU-based data center cross-domain link disaster recovery backup method. The most advanced AMD DPU is used to build a disaster recovery backup solution for network links between data centers, which includes the following steps:
[0079] 1) This implementation adopts a hardware architecture based on the AMD Elba series DSC-200DPU to build a gateway service router GSR. GSR consists of a multi-port 100Gbps high-speed switch tama and a DSC-200DPU interconnected through a PCIe Gen5 x16 interface, where the DPU is equipped with a 16-core ARM Cortex-A78 control plane and a P4 programmable data plane. The DPU has a built-in hardware encryption engine that supports IPSec ESP / AES-GCM-256 algorithms, and integrates an sFlow sampling module and a BGP routing processing accelerator. GSR is deployed at the lower layer of the public network border router CBR in the data center. CBR is connected to the public network ISP via optical fiber, and each GSR node is equipped with dual redundant power supplies and hot-swappable fan modules.
[0080] 2) Construction of cross-domain backup links:
[0081] Establish a backup link between two regions, Region A and Region B: The core service routers (CSRs) in each region are connected to the border routers (BSRs) through 25 Gbps links, and the BSRs are interconnected through 10 Gbps cross-domain dedicated lines. The GSR devices are connected to the lower layer of the CBR in bypass mode, establish EBGP neighbor relationships with the GSRs in the peer region through the BGP protocol, and use the AS-Path Prepending technology to modify the backup path routing entries, appending the local AS number three times to reduce the priority, ensuring that normal traffic prefers to be transmitted through the direct connection path of the CBR for public network transmission.
[0082] Inside the DSC-200 DPU, the zebra routing daemon is responsible for writing the learned routing entries of other regions into the independent routing table 100 of the Linux kernel. The modified fpmsyncd process synchronizes routing updates to other DPU nodes in the cluster through the Redis PUB / SUB channel. When routing is distributed, the dpu_route_programming() interface of the DPU SDK is called to map the next-hop address and the outgoing interface to the TCAM entries of the P4 pipeline.
[0083] 3) When the original IP flow enters the DPU, the data plane performs the following operations: Extract the five-tuple to calculate the hash value and map it to the controllable IP address pool; Insert an extended IP header and perform fragmentation on the data unit according to the MTU; Call the hardware encryption engine to encrypt with the AES-GCM-256 algorithm to generate an ESP encapsulation structure; Add a public network IP header, where the source IP is the public network IP of the GSR, the destination IP is the public network IP of the peer GSR and the UDP header, and set the IP protocol field to 50.
[0084] After the data packet arrives at the peer DPU, it goes through the following hardware pipeline stages: Maintain the reassembly context in the on-chip HBM memory of the DPU with the controllable IP pair as the key; When all the fragments arrive, verify the ICV value and perform the memcpy operation to splice the original data packet. Transfer the reassembled data packet to the kernel protocol stack through the DMA engine; After restoring it to the original IP packet, transfer it to the CBR for forwarding inside the data center.
[0085] 4) The CSR will perform health detection, send 20 ICMP probe packets of 1500 bytes to the peer every 2 seconds, calculate the bidirectional RTT and packet loss rate. The RTT measurement uses hardware timestamps. One sample is extracted from every 512 data packets to analyze the TCP retransmission rate and out-of-order degree, and finally calculate the health:
[0086]
[0087] Perform traffic switching according to the health.
[0088] As described above, it is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claimed rights.
Claims
1. A data center cross-domain link disaster recovery backup method based on DPU, characterized in that: The following steps are involved: 1) Use switches and DPUs to build a gateway service router GSR, and deploy it under the border router CBR as the core processing node of the backup link. Set the backup link route and use the autonomous system path prepending AS-Path Prepending technology to adjust the backup link priority. 2) The DPU-based ARM core fragments the traffic passing through the DPU to meet the maximum transmission unit (MTU) limit in public network transmission. At the same time, in order to meet the memory limit of the DPU, a controllable IP pair is added to reduce the number of flows in the fragment reassembly process; 3) Using the GSR constructed in step 1), the backup link traffic is transmitted end-to-end encrypted through the encryption engine; 4) Through real-time monitoring of the quality of public network links and business transmission status, network anomalies can be dynamically perceived. Based on the preset hierarchical disaster recovery strategy, the core business traffic is prioritized to be quickly switched to low-latency, highly reliable dedicated links, and non-critical traffic is automatically downgraded to the backup path, achieving millisecond-level seamless switching.
2. According to claim 1, a DPU-based data center cross-domain link disaster recovery backup method is characterized by: In step 1), a gateway service router GSR is built based on the switch and DPU, and deployed at the lower layer of the data center public network border router CBR as the core processing node of the cross-domain backup link. The GSR is composed of a multi-port high-speed switch and a DPU. The DPU adopts a heterogeneous computing architecture, integrates an ARM multi-core control plane and a P4 programmable data plane, and realizes high-speed interconnection between the switch and the DPU through a PCIe Gen5 interface. For one area in the two data centers, there is a core service router CSR in the area. The CSR is connected to the servers in the data center and is externally connected to a border router BSR. The BSR is connected to the BSRs in other areas through cross-domain dedicated lines. The CSR is also connected to the public network border router CBR. The CBR is connected to the public network ISP and accesses the public network. Between two areas, GSR publishes its public IP address to the public ISP through CBR and establishes EBGP neighbor relationship with the connected area. GSR in the area learns internal routes from CBR and transmits them to GSR in other areas through EBGP. After receiving the transmitted routes, GSR in other areas needs to transmit them to CBR and perform AS-path override to ensure that public network transmission in normal situations directly goes through CBR instead of GSR backup path, thus forming a routing channel for cross-area backup link; For the routing strategy of DPU, the internal routing daemon of DPU is responsible for writing the learned routing entries of other areas into the independent routing table of Linux kernel. The routing table is designed to store only the encrypted tunnel routes required for cross-DPU communication, and avoid conflicts with the conventional routing table through the kernel namespace isolation mechanism, so as to ensure the accurate classification and directional forwarding of business traffic. The business process pds_dp_app monitors the change events of the routing table in real time by subscribing to the kernel network link socket netlink socket. When adding or deleting routing entries, the routing target network, next hop address and outbound interface information carried by the routing message will be parsed, and the DPU-specific hardware programming interface will be called to send the routing rules to the hardware forwarding plane. In this process, each route of the control plane is dynamically bound to the IPSec encryption policy, and the control plane has customized the open source routing management component fpmsyncd. The modified fpmsyncd process only filters the routing update events of the underlay network, and synchronizes such routes to other DPU nodes in the cluster through the Redis database message queue function.
3. According to claim 2, a DPU-based data center cross-domain link disaster recovery backup method is characterized in that: In step 2), the large packets in the data center traffic are fragmented and reassembled for transmission in the public network. In order to suppress the expansion of the number of fragmented flows, a pair of controllable IP addresses are dynamically added to the outer layer of the original data packet to avoid the problem of flow state explosion, forming an extended data unit: [controllable IP pair header] + [original IP header] + [original data payload]. The controllable IP pair is generated by a preset rule or a hash algorithm: controllable IP pair = H (source IP ⊕ destination IP) mod N, where H is a hash function and N is the size of the controllable IP address pool. Through this mapping, the real IP flow of millions of levels is compressed into a limited controllable IP space. After adding the controllable IP pair, fragmentation and encryption operations are performed. The extended data unit [controllable IP pair header] + [original IP header] + [original data payload] is fragmented according to the MTU limit. The fragment size meets the MTU. IPSec tunnel mode encryption is independently implemented for each fragment. The encryption range includes the controllable IP pair header, the original IP header and the data fragment payload, and the ESP encapsulation structure is generated: [IPSec ESP header] + [encrypted controllable IP header] + [encrypted original IP header fragment] + [encrypted data fragment payload], and then add the UDP header and public network IP header to form the final transmission unit: [public network IP header] + [UDP header] + [IPSec ESP header] + [encrypted controllable IP header] + [encrypted original IP header fragment] + [encrypted data fragment payload]; For the receiving end, IPSec ESP data is extracted based on the public network IP header and UDP header, and the controllable IP pair header is obtained after decryption. This is used as the flow identifier instead of the original IP header fragment. The controllable IP pair is used as the key to maintain the reassembly context in the DPU memory. After all fragments arrive, the controllable IP pair header, original IP header and data payload are spliced according to the offset. After verifying the integrity, the controllable IP pair header is stripped off to restore the original data packet.
4. According to claim 3, a DPU-based data center cross-domain link disaster recovery backup method is characterized in that: In step 3), based on the backup link constructed in step 1), the traffic on the backup link needs to be encrypted for transmission. The specific steps are as follows: Message reception and parsing: IP messages enter the DPU via CBR, and the DPU extracts the Ethernet frame header, IP header, and transport layer information through the hardware parser; the matching processing unit MPU matches the five-tuple of source IP, destination IP, protocol, source port, and destination port according to the pre-programmed P4 rules, hits the encryption policy, and searches for the corresponding security policy SA entry through the content addressable memory CAM to obtain the key material, security parameter index SPI, and anti-replay window parameters; Payload segmentation and padding: The IP packet is segmented according to the selected encryption algorithm AES-GCM-256 block size, and a 12-byte random initialization vector IV is generated by a random number generator, which is combined with a 32-byte sequence number to construct the complete counter initial value J0: J0=IV||0 31 1 AES-GCM encryption: The encryption engine performs the following atomic operations through a fully pipelined architecture: i Generate ciphertext block C i : Where, CTR i =J0+i mod 2 128 , i is the block number, K is the AES-256 key; then the integrity check value of the payload is calculated in parallel through the GMAC authentication algorithm; ESP encapsulation and reassembly: The encrypted data is reassembled into an ESP message: an ESP header is added, including SPI+sequence number and ESP tail, including padding length+next header type: the SPI size is 4 bytes, which identifies the security association; the sequence number size is 4 bytes, which is used to resist replay attacks; at the same time, a 16-byte ICV is appended to the end of the message; finally, the original IP protocol field is modified to 50 to identify ESP, and the IP header checksum is recalculated; Decryption: For the received packet, first perform fragment assembly operation, then find the root destination IP address, ESP protocol and SPI to find SA, then check the sequence number for revisit attacks, and finally verify the ICV check value; finally, according to the algorithm and key parameters specified in the SA, decrypt the encrypted part of the data, remove the padding part, and reconstruct it back to the original IP packet.
5. According to claim 4, a DPU-based data center cross-domain link disaster recovery backup method is characterized in that: In step 4), efficient switching of the primary and backup links is achieved based on bidirectional network status perception and dynamic health evaluation: Bidirectional network status perception uses a 2-second cycle to synchronously collect key indicators of local and peer cross-domain routers to ensure real-time and bidirectional status perception. Active detection of links is achieved through bidirectional ICMP detection: the local router sends a 1500-byte detection packet to the peer router, and the peer router synchronously sends a detection packet of the same specification in the reverse direction to form a bidirectional detection channel. 20 detection packets are sent per cycle to count the bidirectional packet loss rate and round-trip delay RTT. The packet loss rate is calculated using the worst value principle, that is, the larger value of the forward and reverse packet loss rates is taken to ensure sensitivity to link degradation. At the same time, the real-time traffic counts of the local and peer router interfaces are obtained through the SNMPv3 protocol to calculate the interface bandwidth utilization: Where, ifHCOutOctets is the cumulative number of bytes sent by the interface output by the SNMP counter, ifHCOutOctets last It is the latest number of bytes sent, and the sampling interval is fixed at 2 seconds. At the same time, sFlow sampling is used to obtain deep indicators, including TCP retransmission rate and out-of-order packet rate, to assist in evaluating link quality. The health score S comprehensively reflects the link quality. Its calculation integrates four core indicators: In the formula, Local_Load and Remote_Load are the interface utilization rates of the local and remote routers respectively, with a weight of 20% each, reflecting the link congestion level; Loss_Rate is the maximum value of the two-way packet loss rate, with a weight of 30%, directly reflecting the transmission reliability; Measured_Delay is the actual value of the detected RTT; Delay_Threshold is the upper limit of the delay tolerated by the service, and the ratio reflects the delay health, with a weight of 30%; When S<0.7 lasts for 3 consecutive cycles, the main link is judged to be in a sub-healthy state, triggering the switching evaluation process. To suppress instantaneous jitter, the score calculation introduces the exponentially weighted moving average EWMA for smoothing: S smoothed =0.7×S currecnt +0.3×S previous In the formula, S currecnt is the current link health score, S previous is the average of the previous link health scores. Only when the smoothed link health score S smoothed When the value is continuously below the threshold, subsequent actions are triggered; The switching strategy is divided into two modes: full switching and partial traffic switching, which are dynamically selected based on the score and link status: Full switch: When S smoothed <0.5 and the packet loss rate of the main link is ≥10% or the delay is ≥2 times the threshold, the main link is determined to be failed; fast switching is achieved through BGP routing policy update: remove the AS-Path Prepending modification of the backup link routing entry, increase its local priority parameter Local Preference to 250, the default main link is 200, and add 3 AS-Path Prepending to the main link to reduce the priority. With the help of BGP routing reflector, the convergence time can be compressed to within 5 seconds; if the SDN controller is deployed, the OpenFlow flow table batch redirection can be further used to force the traffic to switch to the backup link interface to ensure priority recovery of key services; Partial flow switching: when 0.5≤S smoothed When the traffic rate is less than 0.7 and the remaining bandwidth of the backup link is ≥ 20%, the load balancing mode is started; the ECMP method is used to split traffic according to business priority: high-priority traffic such as database synchronization and transaction requests is retained on the primary link, and low-priority traffic such as log backup and software download is allocated to the backup link in proportion; Weighted Random Early Detection (WRED) is implemented on the backup link to randomly discard non-critical traffic based on queue length and service category to avoid worsening congestion.
Citation Information
Cited By
Cross-domain bidirectional data transmission method and device, and storage medium
CN120614217A
Data backup and recovery optimization method and system of distributed control system
CN121635182A
Distributed edge computer room computing power integration method and system based on DPU
CN121957830A