Lossless switching system, method and equipment and storage medium

By actively discovering link resources, real-time fault detection, redundancy and packet loss retransmission, and cross-core processing, delay and data loss problems in link handover are solved, efficient and lossless link handover is achieved, and the reliability and real-time of the network system are improved.

CN120263624APending Publication Date: 2025-07-04BEIJING COMPUTER NETWORK TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510406895.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing link switching technology has problems such as detection delay, packet loss, inflexible switching methods, and inability to meet high real-time requirements, which affects the reliability and performance of the network.

Method used

The resource management and discovery module is adopted to actively discover link resources, and the real-time fault detection and switching module quickly detects and switches to the backup link. The zero packet loss module ensures data integrity through redundancy and packet loss retransmission technologies. The link state is the same asynchronous processing module supports flexible switching. The low-latency processing module reduces delay through a distributed processing architecture across cores and protocols.

Benefits of technology

It realizes fast and lossless link switching, ensures data integrity and low latency, improves the reliability and real-timeness of the network system, and is suitable for high-real-time application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120263624A_ABST
    Figure CN120263624A_ABST
Patent Text Reader

Abstract

The invention discloses a lossless switching system, which comprises a resource management and discovery module configured to actively discover or configure available link resources, form a resource pool, actively identify the state of the link resources and select an optimal main link according to the link quality; the real-time fault detection and switching module is used for selecting a proper detection packet type according to the flow load magnitude when a network has a fault based on an efficient link fault discovery protocol; the zero data packet loss module is used for ensuring data integrity through redundancy and lost packet retransmission technologies in the link switching process; the link state synchronous and asynchronous processing module is used for processing synchronous or asynchronous operation of link switching; the low-delay processing module enables the delay of single-point data processing not to exceed 0.5 millisecond through a cross-kernel and protocol distributed processing architecture; according to the system, the reliability, the data integrity and the real-time performance of a network system are effectively improved, and the problems of faults, time delay and packet loss possibly occurring in traditional network switching are avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a lossless switching system, method, device and storage medium. Background Art

[0002] With the continuous expansion of the network scale, especially the popularization of emerging technologies such as cloud computing, 5G communication, and the Internet of Things, the requirements for network reliability and real-time performance are becoming increasingly strict. In modern communication systems, link switching technology has become an important means to ensure network stability and efficiency. However, existing link switching technologies have several technical defects, which affect the performance and reliability of the network, specifically manifested in the following aspects:

[0003] In traditional link switching schemes, the detection of link failures usually relies on periodic heartbeat detection or probes based on specific protocols. These methods have detection delays and cannot initiate switching immediately after a failure occurs. Especially when a failure occurs in the network, there may be a long recovery time during the switching process, resulting in service interruption or affecting the user experience. Traditional methods cannot dynamically select the appropriate type of fault detection packet according to the real-time traffic load, which further exacerbates the response time of network fault switching.

[0004] During the link switching process, packet loss is a common problem, especially in the absence of redundancy mechanisms and packet loss retransmission technologies. Due to the lack of an effective packet loss retransmission strategy in traditional switching mechanisms, data loss during link switching directly affects the quality of network services, resulting in incomplete data or transmission interruption, and cannot meet the application requirements of high reliability and high real-time performance.

[0005] In a multi-link environment, link switching operations usually need to select synchronous or asynchronous switching methods according to different traffic types (such as northbound traffic and southbound traffic). Traditional link switching technologies have limited support in this regard, and different switching methods have a greater impact on system performance. Synchronous switching may cause higher latency and affect network response time; while asynchronous switching may lead to inconsistent network states, thus affecting the transmission integrity of data. Existing technologies cannot flexibly and efficiently perform adaptive switching according to traffic types, resulting in instability and performance bottlenecks during the switching process.

[0006] Applications with high real-time requirements (such as video conferencing, online games, financial transactions, etc.) have extremely strict requirements for the latency of link switching. Traditional link switching technologies rely on a single-point data processing architecture, and the processing process may cause significant latency, especially when the network load is high or the link quality is unstable. The single-point processing method is also difficult to meet the low-latency requirements in a distributed architecture. Therefore, existing technologies cannot provide timely service recovery in some critical application scenarios.

[0007] In a multi-link environment, how to efficiently manage and schedule link resources to ensure dynamic adjustment of link quality is the key to improving network stability and performance. However, existing link resource management systems usually can only select resources based on static configurations and are unable to detect changes in link resources and quality fluctuations in a timely manner. The lack of a flexible link quality monitoring and intelligent resource scheduling mechanism makes it difficult for the system to switch to the optimal link in a timely manner when link load is uneven or link quality deteriorates, resulting in a decline in network performance. Summary of the Invention

[0008] The purpose of the present invention is to provide a lossless switching system, method, device, and storage medium, which effectively improve the reliability, data integrity, and real-time performance of the network system, and avoid problems such as failures, delays, and packet losses that may occur in traditional network switching.

[0009] The technical solution adopted by the present invention to solve its technical problems is as follows:

[0010] A lossless switching system includes:

[0011] A resource management and discovery module configured to actively discover or configure available link resources to form a resource pool, actively identify the status of link resources, and select the best primary link according to link quality;

[0012] A real-time fault detection and switching module, based on an efficient link fault discovery protocol, selects an appropriate probe packet type according to the traffic load level when a network fault occurs for real-time link fault detection, and immediately initiates link switching after detecting a fault;

[0013] A zero-packet-loss module for ensuring data integrity through redundancy and packet loss retransmission technologies during link switching;

[0014] A link state synchronous and asynchronous processing module for processing synchronous or asynchronous operations of link switching and providing fast link detection capabilities to support synchronous and asynchronous switching of north-south traffic;

[0015] A low-latency processing module, through a cross-kernel and protocol distributed processing architecture, enables the latency of single-point data processing not to exceed 0.5 milliseconds.

[0016] Another technical problem to be solved by the present invention is to provide a method for implementing lossless switching, including the following steps:

[0017] Actively discover and configure available link resources to form a resource pool, and select the best primary link according to link quality;

[0018] Use an efficient link fault discovery protocol for real-time fault detection, and quickly switch to a backup link when a fault occurs;

[0019] Achieve zero packet loss during link switching through redundant links and packet loss retransmission;

[0020] Provide the ability to switch link status synchronously and asynchronously to avoid the spread of network failures caused by single-point failures;

[0021] Ensure that the latency of network operations is less than or equal to 0.5 milliseconds through cross-kernel and protocol distributed processing.

[0022] Preferably, the method for actively discovering and configuring link resources and selecting the best primary link is as follows:

[0023] Identify and configure available link resources in the network through an active discovery mechanism, and select the optimal link based on link quality evaluation. The link quality score Qi is determined by the following parameters:

[0024]

[0025] Select the best link as the primary link:

[0026]

[0027] where w1, w2, w3, w4 are weight coefficients, Bandwidth i is the bandwidth, Latency i is the latency, Loss Rate i is the packet loss rate, and Load i is the load.

[0028] Preferably, the method for using an efficient link failure discovery protocol for real-time failure detection and quickly switching to a backup link when a failure occurs is as follows:

[0029] Set the health Hi of each link and calculate its real-time health through dynamic parameters such as latency, packet loss rate, and bandwidth:

[0030]

[0031] where α, β, γ are weighting coefficients that adjust the evaluation criteria for health. When the link health is lower than the threshold, the link is considered to have failed, i.e.:

[0032] H i <Threshold

[0033] At this time, trigger link switching and select the backup link Backup Link:

[0034]

[0035] Preferably, the method for achieving zero packet loss during link switching through redundant links and packet loss retransmission is as follows:

[0036] When a link failure is detected, cache the currently transmitted data packet and transmit it after switching to the backup link. Let the delay of link switching be Δt and the caching time be t cache , and the amount of cached data packets is N cache , and use the following method to manage packet loss and retransmission:

[0037] Cache Time=Δt and N cache =Throughput·Δt

[0038] Use the TCP protocol for packet loss retransmission so that the data packets lost during the handover are retransmitted.

[0039] Preferably, the method for providing synchronous and asynchronous handover capabilities of the link state is as follows:

[0040] Synchronous handover: All switches or routers switch to the backup link at the same time;

[0041] Asynchronous handover: Make the data flow consistent during the link handover process through the multi-path protocol and routing calculation;

[0042] During synchronous handover, the device waits for all routers / switches to confirm the handover. During asynchronous handover, each device independently judges the link handover.

[0043] Preferably, the method for ensuring that the delay of network operations is less than or equal to 0.5 milliseconds through cross-kernel processing is as follows:

[0044] Suppose there are N computing cores, the processing time of each core is Tcore, and the tasks are evenly distributed, then the total delay T total is expressed as:

[0045]

[0046] Among them:

[0047] T task is the processing time of a single task, N is the number of cores, and T communication is the communication delay between cores.

[0048] Preferably, the method for ensuring that the delay of network operations is less than or equal to 0.5 milliseconds through protocol distributed processing is as follows:

[0049] By directly accessing the remote memory without going through the operating system kernel;

[0050] Use RDMA, and the network delay T network is expressed as:

[0051] T network =TNIC +T transport +T protocal overhead

[0052] Among them, T NIC is the delay of the network interface card, and T transport is the delay of the transmission protocol, and T protocoloverhead is the overhead of the protocol stack.

[0053] Another technical problem to be solved by the invention is to provide an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the lossless switching method described in any one of the above is implemented.

[0054] Another technical problem to be solved by the invention is to provide a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the lossless switching method described in any one of the above is implemented.

[0055] The beneficial effects of the present invention are as follows:

[0056] Through the real-time fault detection and switching module, based on an efficient link fault discovery protocol and traffic load level, the detection packet type is intelligently selected, enabling rapid detection of network faults and prompt switching to the standby link. The timeliness of fault detection and the efficiency of switching effectively solve the service interruption problem during link faults; through the zero-packet-loss module, redundant and packet-loss retransmission technologies are used to ensure that data is not lost during link switching. This technology avoids the common data loss problem during link switching by ensuring the complete transmission of data during the switching process.

[0057] The link state synchronous and asynchronous processing module in the solution can flexibly handle synchronous and asynchronous switching, providing fast link detection capabilities. For north-south and south-north traffic, it can support different switching methods for different traffic, ensuring the flexibility and efficiency of the switching operation, avoiding inconsistencies and performance bottlenecks; through the low-latency processing module, adopting a cross-kernel and protocol distributed processing architecture, the data processing latency does not exceed 0.5 milliseconds. This not only ensures the efficiency of link switching but also ensures that the latency of single-point processing remains at a very low level during distributed processing, meeting the requirements of high-real-time scenarios; through the resource management and discovery module, available link resources are actively discovered and configured to form a resource pool, and the best primary link is selected according to the link quality. This module monitors the link quality and automatically adjusts the selection and allocation of links to ensure timely switching when the link quality deteriorates, avoiding performance bottlenecks. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 is a flowchart of a lossless switching method of the present invention. Specific implementation method

[0060] The principles and features of the present invention will be described below. The examples given are only for explaining the present invention and are not intended to limit the scope of the present invention. The present invention will be described more specifically by way of example in the following paragraphs. The advantages and features of the present invention will be clearer according to the following description and claims.

[0061] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. The terms used in the description of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items. Embodiment

[0062] Refer to Figure 1 As shown, a non-destructive switching system includes:

[0063] A resource management and discovery module, configured to actively discover or configure available link resources, form a resource pool, actively identify the status of link resources, and select the best primary link according to link quality;

[0064] A real-time fault detection and switching module, based on an efficient link fault discovery protocol, selects an appropriate probe packet type according to the traffic load magnitude when a network fault occurs, performs real-time link fault detection, and immediately initiates link switching after detecting a fault;

[0065] A zero data packet loss module, used to ensure data integrity through redundancy and packet loss retransmission technologies during link switching;

[0066] A link state synchronous and asynchronous processing module, used to process synchronous or asynchronous operations of link switching, and provide fast link detection capabilities, supporting synchronous and asynchronous switching of north-south traffic;

[0067] A low-latency processing module, through a cross-kernel and protocol distributed processing architecture, makes the latency of single-point data processing not exceed 0.5 milliseconds.

[0068] Through redundancy and packet loss retransmission technologies, this solution can ensure data integrity during link switching and ensure that no data is lost. This is crucial for applications with high real-time requirements, such as video conferencing, online transactions, etc.; the system can quickly switch to a backup link after a link fault occurs without affecting the complete transmission of data.

[0069] The resource management and discovery module forms a dynamic and flexible resource pool by actively discovering and configuring available link resources. It can effectively identify the status of link resources, automatically select the primary link with the best link quality, and ensure the stability and efficiency of the link. The system can dynamically adjust the primary link according to the real-time link quality to avoid affecting network performance due to link quality fluctuations.

[0070] The real-time fault detection and switching module, based on an efficient link fault discovery protocol, can detect faults immediately at the moment of fault occurrence and select an appropriate probe packet type according to the traffic load magnitude to ensure the accuracy and timeliness of detection. Fault detection and switching can be completed in an extremely short time, thus avoiding long-term service interruption or performance degradation.

[0071] The link status synchronous and asynchronous processing module provides flexible processing capabilities for synchronous and asynchronous switching, which is especially suitable for different types of traffic, such as different switching requirements for northbound traffic and southbound traffic. It supports automatically adjusting the switching mode according to the traffic type to ensure that the traffic always maintains the optimal path during the switching process. Flexible synchronous and asynchronous operations can be optimized in different scenarios to reduce the switching delay and avoid system instability.

[0072] The low-latency processing module, through a distributed architecture design across kernels and protocols, ensures that the latency of single-point data processing does not exceed 0.5 milliseconds. This low-latency requirement ensures that the system can still maintain high performance in scenarios facing high-concurrency traffic or high real-time requirements.

[0073] A method for realizing lossless switching includes the following steps:

[0074] Actively discover and configure available link resources, form a resource pool, and select the best primary link according to the link quality;

[0075] Use an efficient link fault discovery protocol for real-time fault detection and quickly switch to the backup link when a fault occurs;

[0076] Achieve zero packet loss during the link switching process through redundant links and packet retransmission;

[0077] Provide link status synchronous and asynchronous switching capabilities to avoid the spread of network faults caused by single-point failures;

[0078] Ensure that the latency of network operations is less than or equal to 0.5 milliseconds through cross-kernel and protocol distributed processing.

[0079] Redundant link and packet loss retransmission technologies can ensure that data is not lost during the process of link switching, and are especially suitable for high-demand applications such as real-time audio and video, online transactions, etc. This means that even if a link failure occurs, the system can seamlessly switch to ensure service continuity and data integrity; by actively discovering and configuring available link resources, the system can dynamically manage link resources, real-time evaluate the quality of links (such as latency, bandwidth, packet loss rate, etc.), and select the best primary link to improve network performance and reliability. The system can automatically adjust the primary link according to the change of link quality at any time to reduce the impact of link quality fluctuations on services.

[0080] Efficient link failure discovery protocols (such as BFD, LLDP, etc.) can quickly detect and switch to the backup link at the first time when a network failure occurs. This fast response mechanism can effectively avoid service interruption, especially when a network failure or link failure occurs, which can greatly improve the stability of the system; the provided synchronous and asynchronous switching capabilities of link status enable the system to flexibly adjust the switching method according to different service requirements. Synchronous switching can reduce traffic interruption, while asynchronous switching can minimize interference during the link switching process. This ability effectively avoids the propagation of network failures caused by single-point failures and improves the reliability of the network; through a cross-kernel and protocol distributed processing architecture, the data processing tasks are distributed to multiple nodes, and through an efficient protocol stack and hardware acceleration technology, it is ensured that the latency of network operations is less than or equal to 0.5 milliseconds. Such low latency guarantee is crucial for applications with extremely high requirements for timeliness (such as high-frequency trading, real-time video streaming, etc.).

[0081] Automatic discovery of link resources: By deploying link quality monitoring tools or protocols (such as ICMP, SNMP, etc.), actively discover available link resources and measure the status of links, including bandwidth, latency, packet loss rate, etc. Form a resource pool for dynamically selecting the best link; Link resource pool management: Introduce a network management system (NMS) to continuously monitor the link quality, establish a resource pool, and automatically select the best primary link according to the link quality. The SDN (Software Defined Network) architecture can be used to manage the selection and switching of links through a central control platform.

[0082] Efficient failure discovery protocol: Use protocols such as BFD to quickly detect link failures. BFD can discover link failures within milliseconds, reducing the time for failure detection and switching; Link switching strategy: When a link failure is detected, the system automatically selects a backup link according to the current load and service type and performs the link switching operation. It is necessary to ensure that the backup link has sufficient bandwidth and low latency to support the traffic load.

[0083] Redundant Link Design: By presetting redundant links or backup links, ensure quick switching when the primary link fails. The redundant links can be physical or virtual links, with the goal of ensuring that there is always an available path for traffic; Packet Loss Retransmission Mechanism: Adopt the retransmission mechanism of TCP or UDP protocols to compensate for lost data packets. During the link switching process, the system can detect data loss and request retransmission to ensure data integrity. For real-time traffic (such as video and voice), a UDP transmission method with redundancy (such as the RTP protocol) can be adopted to maintain smooth data transmission even in case of packet loss.

[0084] Synchronous Switching: For latency-sensitive traffic, adopt a synchronous switching strategy, i.e., during the link switching process, the transfer of traffic is immediate to avoid interruption as much as possible. Synchronous switching needs to ensure the rate and latency matching of the primary link and the backup link; Asynchronous Switching: For traffic that tolerates a certain amount of latency, adopt an asynchronous switching strategy. Asynchronous switching can reduce traffic interruption during the switching process but may slightly increase the switching time. A system supporting asynchronous switching can delay for a period of time after fault detection before switching to the backup link to reduce the impact on network performance.

[0085] Distributed Data Processing: Through a distributed architecture across kernels and protocols, avoid bottlenecks in the data processing process. For example, adopt technologies such as DPDK to accelerate the processing of data packets and reduce the latency of data packet forwarding and protocol stack processing; Hardware Acceleration: Utilize hardware acceleration technologies such as FPGA and ASIC to accelerate the processing of the network protocol stack, reduce the CPU load, and achieve extremely low latency in data forwarding speed; Kernel Optimization: Through kernel optimization of the network protocol stack (such as reducing system calls and improving interrupt response speed), further reduce the latency of network operations to ensure that the overall system latency is less than or equal to 0.5 milliseconds.

[0086] The method for actively discovering and configuring link resources and selecting the best primary link is as follows:

[0087] Identify and configure available link resources in the network through an active discovery mechanism, and select the optimal link based on link quality assessment. The link quality score Qi is determined by the following parameters:

[0088]

[0089] Select the best link as the primary link:

[0090]

[0091] where w1, w2, w3, w4 are weight coefficients, Bandwidth i is the bandwidth, Latency i is the latency, Loss Rate iis the packet loss rate, and Load i is the load.

[0092] Identify and configure the available link resources in the network through the active discovery mechanism to ensure that the system can update and manage the link resources in real time. This dynamic link selection method can evaluate the quality of each link according to the actual link performance conditions (bandwidth, delay, packet loss rate, load, etc.) to ensure that the best primary link is always selected; the link quality score (Qi) combines multiple key parameters (bandwidth, delay, packet loss rate, load, etc.), and the weight of each parameter (w1, w2, w3, w4) can be flexibly adjusted according to actual needs, enabling the system to preferentially select the link with the optimal performance and suitable for the current service load.

[0093] The system can dynamically adjust the weight coefficients (w1, w2, w3, w4) of each parameter according to different requirements of actual applications. For example, in a scenario with high bandwidth requirements, the weight of bandwidth can be increased to ensure that bandwidth becomes the main evaluation criterion; in a scenario sensitive to delay and packet loss, the weights of delay and packet loss rate can be increased to preferentially select a link with less delay and less packet loss. This flexibility can help network administrators optimize for different service requirements.

[0094] By dynamically evaluating the link quality and selecting the best link, the allocation of network resources can be optimized to the greatest extent. Load balancing and efficient utilization of resources avoid waste of link resources. For example, links with high load can be avoided, and links with low load can be enabled when needed, thus improving the overall utilization efficiency of network resources.

[0095] While actively discovering link resources, the system can monitor the status of the links in real time. Once a performance problem or fault occurs in a certain link, the system can quickly switch to another link with higher quality to ensure that the network does not experience serious faults. This proactive link quality assessment and selection strategy enhances the fault tolerance and robustness of the network.

[0096] For end users, by selecting the optimal link to provide services, the network quality can be significantly improved, reducing problems such as network latency, packet loss, and congestion, and optimizing the user experience, especially for applications with high requirements for network quality, such as video conferencing, online games, cloud computing, etc.; this solution is not only applicable to existing network architectures but also has strong scalability. When the network scale expands and the types of links increase, the system can continue to manage and evaluate the newly added links through the active discovery mechanism. The system will automatically select the best link according to the network conditions to ensure the stability and scalability of the system.

[0097] The method of using an efficient link fault discovery protocol for real-time fault detection and quickly switching to a backup link when a fault occurs is as follows:

[0098] Set the health Hi of each link, and calculate its real-time health through dynamic parameters such as delay, packet loss rate, and bandwidth:

[0099]

[0100] Among them, α, β, and γ are weighting coefficients that adjust the evaluation criteria of health. When the link health is lower than the threshold, it is considered that the link has failed, that is:

[0101] H i <Threshold

[0102] At this time, trigger link switching and select the backup link Backup Link:

[0103]

[0104] This solution calculates the link health (Hi) dynamically and monitors parameters such as delay, packet loss rate, and bandwidth of each link in real time. When the link health is lower than the preset threshold, the system can quickly identify the faulty link and trigger the link switching mechanism in a timely manner. This ensures that the network can recover as soon as possible when a link fails, reducing the impact of the failure on the business.

[0105] By adjusting the weighting coefficients (α, β, γ), the evaluation criteria of link health can be flexibly adjusted according to the needs of different applications. For example, in real-time video and voice communication, delay and packet loss rate may be more important than bandwidth, while in big data transmission, bandwidth may be the most critical factor. This can optimize the link evaluation according to specific needs to ensure the optimal link switching strategy.

[0106] By automatically switching to the backup link when a link fails, the system can greatly improve the reliability and fault tolerance of the network. Even if there is a link problem, the timely activation of the backup link can ensure that the network service is not interrupted, enhancing the stability of the network and its adaptability to faults, especially in high-demand business scenarios (such as financial transactions, video conferences, etc.).

[0107] The method to achieve zero packet loss during link switching through redundant links and packet retransmission is as follows:

[0108] When a link failure is detected, cache the currently sent data packets and transmit them after switching to the backup link. Let the delay of link switching be Δt and the cache time be t cache , and the amount of cached data packets is N cache , and use the following method to manage packet loss and retransmission:

[0109] Cache Time = Δt and N cache = Throughput·Δt

[0110] Use the TCP protocol for packet loss retransmission so that the data packets lost during the handover are retransmitted.

[0111] By caching the currently sent data packets and transmitting them after switching to the backup link, the packet loss phenomenon during the link switching process can be effectively avoided. Even if packet loss occurs on the backup link, through the TCP packet loss retransmission mechanism, it is ensured that the lost data packets can be retransmitted, ultimately achieving zero packet loss; when a link fails, the system switches through redundant links without interrupting data transmission. The caching mechanism and the smooth handover of the backup link ensure the continuity of data transmission and are transparent to the application layer users, avoiding service interruption or business loss.

[0112] Utilizing the reliability of redundant links and the TCP protocol, the system can maintain stable transmission even when a link fails. The retransmission mechanism of the TCP protocol can ensure that the data finally reaches the destination in case of packet loss on the backup link, thus significantly improving the reliability of the entire communication system; by caching data packets during the redundant link switch, the permanent loss of data packets is avoided. After the link is restored, the cached data packets will be transmitted immediately, thus completing the data transmission in a short time and reducing the delay of fault recovery.

[0113] When a link fails, the backup link can immediately take over data transmission instead of waiting for all the lost data packets to be re-requested. Redundant links can maximize the utilization of network bandwidth and maintain a high transmission efficiency during link failures; this solution can be applied to various network environments, such as wireless networks and satellite links where packet loss is likely to occur. Through the design of redundant links and the mechanism of the TCP protocol, it can flexibly cope with various fluctuations and failures in the network.

[0114] For real-time data transmission or critical services, packet loss will seriously affect the quality and performance. Through this solution, although packet loss may occur during link failures, the TCP retransmission mechanism ensures that the data will not be lost for a long time, guaranteeing the high availability and stability of the services; using the TCP protocol itself has high versatility, and most existing network devices and applications support the TCP protocol, so there is no need to make major changes to the existing system to achieve the zero-packet-loss solution. In addition, the configuration and management of redundant links are relatively simple and easy to deploy and maintain.

[0115] The method for providing synchronous and asynchronous handover capabilities of link status is as follows:

[0116] Synchronous handover: All switches or routers switch to the backup link at the same time;

[0117] Asynchronous handover: Make the data flow consistent during the link switching process through the multi-path protocol and routing calculation;

[0118] During synchronous handover, the device waits for all routers / switches to confirm the handover. During asynchronous handover, each device independently determines the link handover.

[0119] All devices switch to the standby link at the same moment, which can ensure the consistency of the entire network in case of link failure, avoiding data loss or inconsistency caused by some devices failing to switch in time. This synchronous handover guarantees the integrity and reliability of the network; each device independently determines the link handover, allowing different devices to complete the handover at different times. Even if some devices cannot switch immediately due to failures, other devices can still continue to work normally, enhancing the fault tolerance of the network.

[0120] Although synchronous handover requires waiting for coordination and confirmation among devices during handover, it can avoid data inconsistency or data loss caused by some devices not finishing the handover, ensuring a smooth network handover process; when each device independently determines the handover, it does not need to wait for confirmation from other devices, so the latency of link handover can be significantly reduced. Especially in a large-scale network, the handover process can be completed more quickly, reducing the time of network interruption.

[0121] For environments with extremely high requirements for consistency, synchronous handover can ensure consistency and coordination during the link handover process, and is suitable for application scenarios that are sensitive to handover latency and require consistency; through multi-path protocols and routing calculations, devices can independently determine link handover, which helps to flexibly perform path selection and optimization in complex network topologies, and can cope with various network topologies and dynamic changes, enhancing the scalability and flexibility of the network.

[0122] The method to ensure that the latency of network operations is less than or equal to 0.5 milliseconds through cross-kernel processing is as follows:

[0123] Suppose there are N computing cores, and the processing time of each core is Tcore. Assuming the tasks are evenly distributed, the total latency T total is expressed as:

[0124]

[0125] where:

[0126] T task is the processing time of a single task, N is the number of cores, and T communication is the communication latency between cores.

[0127] The method to ensure that the latency of network operations is less than or equal to 0.5 milliseconds through protocol distributed processing is as follows:

[0128] By directly accessing remote memory without going through the operating system kernel;

[0129] Using RDMA, the network latency T network is expressed as:

[0130] T network = T NIC + T transport + T protocol overhead

[0131] where T NIC is the latency of the network interface card, T transport is the latency of the transmission protocol, and T protocoloverhead is the overhead of the protocol stack.

[0132] By evenly distributing tasks across multiple computing cores, each core is responsible for independent processing tasks, thus reducing the burden on a single core and lowering the total latency. The total latency Ttotal is jointly determined by the task processing time, the communication latency between cores, and the allocation strategy. A reasonable core allocation can ensure latency minimization; The RDMA technology directly accesses remote memory without going through the operating system kernel, bypassing traditional kernel-level communication overheads (such as context switching, protocol stack processing, etc.), effectively reducing the network latency Tnetwork, and further reducing latency and increasing data transfer speed due to the absence of intermediate processing nodes.

[0133] By directly accessing remote memory without going through the operating system kernel, the time of kernel intervention is reduced, and the overhead of the traditional network protocol stack is avoided. This enables data to be transferred to the target device more quickly and directly, thus optimizing the latency performance of network operations and meeting the requirements for ultra-low latency applications (such as financial transactions, real-time data processing, etc.).

[0134] RDMA reduces the latency T NIC of the network interface card and the overhead T protocol overhead of the transmission protocol through an efficient network interface card (NIC) and an optimized transmission protocol. In addition, using RDMA can reduce the resources required for transmission protocol processing, lowering the overall network latency, which is particularly effective for high-frequency and large-scale data communications; RDMA optimizes the data transmission path by reducing the overhead T transport of the transport layer protocol, making network communication more efficient and ensuring a network latency below 0.5 milliseconds.

[0135] Since this solution relies on distributed processing (across multiple computing cores), it can flexibly adapt to different scales of computing requirements. As the number of computing cores in the system increases, the system can better handle high loads, and the latency can still be maintained within the expected range (less than 0.5 milliseconds). RDMA can also provide low-latency and high-bandwidth communication in large-scale data centers and cloud computing environments, with excellent scalability; through distributed computing and optimized network transmission paths, the system can efficiently utilize computing resources while ensuring low latency. The processing time T of each core core and the communication latency T communication are optimized so that the system can still operate efficiently under high concurrency, ensuring task processing efficiency and data transmission speed.

[0136] This embodiment also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the lossless switching method as described above is implemented.

[0137] This embodiment also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by the processor, the lossless switching method as described above is implemented.

[0138] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in this application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0139] Those skilled in the art can clearly understand that, for the convenience and conciseness of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be allocated to different functional units and modules as needed, that is, the internal structure of the system is divided into different functional units or modules to complete all or part of the functions described above.

[0140] The above-mentioned embodiments of the present invention do not limit the protection scope of the present invention. The implementation manners of the present invention are not limited thereto. All such modifications, substitutions or changes in various other forms made to the above structure of the present invention according to the above content of the present invention, in accordance with the common general knowledge and conventional means in the art, without departing from the above basic technical idea of the present invention, shall fall within the protection scope of the present invention.

Claims

1. A non-destructive switching system, characterized in that, It includes: A resource management and discovery module, configured to actively discover or configure available link resources, form a resource pool, actively identify the status of link resources, and select the best primary link according to link quality; A real-time fault detection and switching module, based on an efficient link fault discovery protocol, selects an appropriate probe packet type according to the traffic load level when a network fault occurs, performs real-time link fault detection, and immediately initiates link switching after detecting a fault; A zero-packet-loss module, used to ensure data integrity through redundancy and packet loss retransmission technologies during link switching; A link status synchronous and asynchronous processing module, used to process synchronous or asynchronous operations of link switching, and provide fast link detection capabilities, supporting synchronous and asynchronous switching of north-south traffic; A low-latency processing module, through a cross-kernel and protocol distributed processing architecture, makes the latency of single-point data processing not exceed 0.5 milliseconds.

2. A method for realizing lossless switching, characterized in that, It includes the following steps: Actively discover and configure available link resources, form a resource pool, and select the best primary link according to link quality; Use an efficient link fault discovery protocol for real-time fault detection, and quickly switch to a standby link when a fault occurs; Achieve zero packet loss during link switching through redundant links and packet loss retransmission; Provide link status synchronous and asynchronous switching capabilities to avoid the spread of network faults caused by single-point failures; Ensure that the latency of network operations is less than or equal to 0.5 milliseconds through cross-kernel and protocol distributed processing.

3. The lossless switching method according to claim 2, wherein The method for actively discovering and configuring link resources and selecting the best primary link is: Identify and configure available link resources in the network through an active discovery mechanism, and select the optimal link based on link quality evaluation. The link quality score Qi is determined by the following parameters: Select the best link as the primary link: Among them, w1, w2, w3, w4 are weight coefficients, Bandwidth i is bandwidth, Latency i Delay Loss Rate i is the packet loss rate, Load i For load.

4. The non-destructive switching method according to claim 2, characterized in that, The method for using an efficient link fault discovery protocol for real-time fault detection and quickly switching to a standby link when a fault occurs is: Set the health Hi of each link, and calculate its real-time health through dynamic parameters such as latency, packet loss rate, and bandwidth: Among them, α, β, γ are weighting coefficients that adjust the evaluation criteria of health. When the link health is lower than the threshold, the link is considered to have failed, that is: H i <Threshold At this time, trigger link switching and select the standby link Backup Link:

5. The lossless switching method according to claim 2, wherein The method for achieving zero packet loss during link switching through redundant links and packet loss retransmission is: When a link failure is detected, cache the currently transmitted data packet and transmit it after switching to the backup link. Let the delay of link switching be Δt and the caching time be t cache , and the amount of cached data packets is N cache , and use the following method to manage packet loss and retransmission: Cache Time=△t and N cache =Throughput·△t Use the TCP protocol for packet loss retransmission to retransmit the packets that are lost during the switch.

6. The non-destructive switching method according to claim 2, characterized in that The method for providing link status synchronous and asynchronous switching capabilities is: Synchronous switching: All switches or routers switch to the standby link at the same time; Asynchronous switching: Make the data flow consistent during the link switching process through a multi-path protocol and routing calculation; During synchronous switching, the device waits for all routers / switches to confirm the switch. During asynchronous switching, each device independently determines the link switch.

7. The non-destructive switching method according to claim 2, characterized in that The method for ensuring that the latency of network operations is less than or equal to 0.5 milliseconds through cross-kernel processing is: There are N computing cores, and the processing time of each core is Tcore. Assuming the tasks are evenly distributed, the total delay T total is expressed as: Among them: T task is the processing time of a single task, and N is the number of cores, T communication is the communication delay between cores.

8. The non-destructive switching method according to claim 7, characterized in that The method for ensuring that the latency of network operations is less than or equal to 0.5 milliseconds through protocol distributed processing is: By directly accessing remote memory without going through the operating system kernel; Using RDMA, network latency T network is expressed as: T hetwork = T NIC + T transport + T protocoloverhead Among them, T NIC is the latency of the network interface card, T transport is the latency of the transmission protocol, T protocol overhead is the overhead of the protocol stack.

9. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the lossless switching method described in claim 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the lossless switching method described in claim 7.

Citation Information

Patent Citations

  • Data processing method and device for link switching process

    CN102238069A

  • Roaming link switching method, mobile terminal, network module and storage medium

    CN107708163A

  • Fault link detection and recovery method based on link quality evaluation model

    CN116455729A

  • Cloud computing-based method for smart link selection in sd-wan network

    WO2022007244A1