A method and system for multi-link testing based on SD-WAN

By employing bidirectional communication and clustering algorithm models for probe and feedback messages, the high cost and complexity issues caused by UCPE boxes in SD-WAN networks are resolved. This enables real-time monitoring and automated management of link quality, improving network testing efficiency and stability.

CN119892693BActive Publication Date: 2025-11-04CHINA TELECOM CLOUD TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411773120.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-04
Publication Date
2025-11-04
Estimated Expiration
2044-12-04

AI Technical Summary

Technical Problem

In SD-WAN networks, existing technologies require the deployment of UCPE boxes at each node, resulting in complex and costly testing environments and making it difficult to effectively manage and maintain the performance of a large number of links.

Method used

By using a two-way communication mechanism of probe and feedback messages, a pre-trained clustering algorithm model is used to cluster link quality data, identify the normal and abnormal states of the link, and monitor link quality in real time through NETPERF and SAR traffic monitoring tools, automatically adjusting scheduling priorities and resource allocation strategies.

Benefits of technology

It simplifies the complexity of the SD-WAN testing environment, reduces costs, improves testing efficiency, can quickly identify abnormal link states, reduces manual intervention, optimizes network configuration, and improves network stability and transmission efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119892693B_ABST
    Figure CN119892693B_ABST
Patent Text Reader

Abstract

The disclosure provides a kind of SD-WAN-based multi-link test method and system, it is related to computer network technical field, the method comprises: by first test equipment to second test equipment sends probe message, and receives the feedback message returned by second test equipment to probe message;According to probe message and feedback message, determine the link quality data of each of multiple test links;Based on the clustering algorithm model obtained by pre-training, the link quality data of each of multiple test links is clustered, to obtain the multiple clusters for representing link quality and the convergence area of each of multiple clusters;Link quality data not belonging to convergence area is determined as abnormal data, and abnormal data is analyzed.This disclosure uses the two-way communication mechanism of probe message and feedback message, effectively simplifies the complexity of SD-WAN test environment, reduces test cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer network technology, specifically to a multi-link testing method and system based on SD-WAN. Background Technology

[0002] In the process of digital transformation of modern enterprises, SD-WAN (Software-Defined Wide Area Network) has been widely used as an important network architecture in scenarios such as branch office interconnection, data center connection, cloud service access, and cross-border office operations. In a typical SD-WAN deployment, the enterprise headquarters or central node needs to establish connections with multiple branch offices, cloud service providers, and data centers, which usually involves probing dozens or even hundreds of links.

[0003] However, each node typically requires a UCPE (Unified Edge Computing Platform) box, resulting in a complex and costly testing environment. As businesses expand, the need for increased link capacity becomes commonplace, posing significant challenges to network management and maintenance. The deployment and management of each UCPE box not only increases hardware costs but also complicates the network architecture, especially when real-time monitoring and performance evaluation of each link are required. This necessitates greater resource and manpower investment in maintenance, leading to a complex and costly testing environment. Summary of the Invention

[0004] This disclosure provides a multi-link testing method and system based on SD-WAN, aiming to solve the problems existing in the background art.

[0005] To solve the above-mentioned technical problems, this disclosure is implemented as follows:

[0006] In a first aspect, embodiments of this disclosure provide a multi-link testing method based on SD-WAN, the method comprising:

[0007] The first device under test (DUT) sends a probe message to the second device under test (DUT) and receives a feedback message from the second device under test in response to the probe message. The first DUT and the second DUT are each configured with multiple links under test.

[0008] Based on the probe messages and the feedback messages, the link quality data of each of the multiple links under test is determined;

[0009] Based on the pre-trained clustering algorithm model, the link quality data of each of the multiple links to be tested are clustered to obtain multiple clusters for characterizing link quality and the convergence regions of each of the multiple clusters.

[0010] Link quality data that does not belong to the convergence region is identified as abnormal data, and anomaly analysis is performed on the abnormal data.

[0011] Optionally, after receiving the feedback message returned by the second device under test in response to the probe message, the method further includes:

[0012] Based on the probe messages and the feedback messages, the actual bandwidth of each of the multiple links under test is determined;

[0013] For any link to be tested, compare the actual bandwidth of the link to be tested with a preset first threshold.

[0014] When the actual bandwidth is greater than or equal to a preset first threshold, a data stream is sent to the first device under test for the link under test using the NETPERF streaming tool of the second device under test.

[0015] The SAR traffic monitoring tool of the first device under test is used to determine the current remaining bandwidth of the link under test, and the current remaining bandwidth is used as the link quality data of the link under test.

[0016] Optionally, the method further includes:

[0017] The CPU utilization and memory usage of each of the multiple test links are determined using the SAR traffic monitoring tool of the first device under test.

[0018] From the plurality of test links, identify the test links that are exclusively used by the CPU with a CPU utilization rate higher than a preset second threshold, and identify the test links that have memory leaks with a memory utilization rate higher than a preset third threshold.

[0019] For the test link that is CPU-exclusive or has memory leaks, the scheduling priority and resource allocation strategy of the test link are adjusted in the QOS yang files of the first and second test devices, respectively.

[0020] Optionally, the method further includes:

[0021] In the configuration files of the first device under test and the second device under test, the link UUID number is increased to a preset link limit, and the link UUID number corresponds one-to-one with the plurality of links under test;

[0022] The SAR traffic monitoring tool of the first device under test is used to monitor the link quality data of each of the multiple links under test. The link quality data includes: latency, jitter and packet loss rate.

[0023] Based on the link quality data of each of the multiple links under test, determine the interface status between the first device under test and the second device under test;

[0024] When the interface status indicates that the interface is blocked, the scheduling priority and resource allocation strategy of the multiple test links are adjusted accordingly in the QoS yang files of the first and second test devices.

[0025] Optionally, the clustering algorithm model can be pre-trained according to the following steps:

[0026] Collect multiple historical link quality data of multiple links under test at different time periods, and input the multiple historical link quality data as training samples into the clustering algorithm model;

[0027] From the multiple historical link quality data, N historical link quality data are randomly selected as the first cluster centroid of each of the N clusters, where N is an integer greater than 1;

[0028] Calculate the distance from each historical link quality data point to the cluster centroid, and assign the multiple historical link quality data points to the nearest cluster;

[0029] Based on the average value of historical link quality data within the same cluster, the cluster centroid is iteratively updated until the difference between the changes of the cluster centroids of two adjacent updates is less than a preset change threshold, thus obtaining N convergent clusters.

[0030] Based on the centroids of the N convergent clusters, the convergence regions of the N clusters are determined.

[0031] Optionally, after clustering the link quality data of each of the multiple links to be tested based on the pre-trained clustering algorithm model, the method further includes:

[0032] Determine the centroid of each of the multiple clusters, and the link quality index corresponding to each centroid, wherein the link quality index is a one-dimensional index;

[0033] The link quality data of each of the multiple links under test is mapped to the cluster centroid of the cluster to obtain the link quality index of each of the multiple links under test. The link quality data includes at least: latency, jitter and packet loss rate.

[0034] The method further includes:

[0035] Monitor whether any of the multiple links under test are faulty;

[0036] In the presence of a faulty link, determine the cluster centroid of the cluster to which the faulty link belongs;

[0037] From the cluster to which the faulty link belongs, determine the backup link that is closest to the abnormal link in terms of link quality index and whose current remaining bandwidth meets the preset bandwidth requirement.

[0038] Switch the faulty link to the backup link.

[0039] Optionally, after clustering the link quality data of each of the multiple links to be tested based on the pre-trained clustering algorithm model, the method further includes:

[0040] Collect historical link quality data from multiple links under test at different time periods to form a historical dataset;

[0041] Based on the historical dataset, determine the link quality evolution trend of each of the multiple links to be tested;

[0042] Based on the clustering results and link quality evolution trends of the link quality data, potential faulty links whose link quality data exceeds the set tolerance range for multiple consecutive time periods are identified, and result-oriented analysis is performed on the potential faulty links.

[0043] Secondly, this disclosure provides a multi-link testing system based on SD-WAN, applied to the steps of the multi-link testing method based on SD-WAN as described above. The system includes: a first device under test (DUT), a second DUT, and a PC. The first DUT and the second DUT are each configured with multiple DUT links. The first DUT is connected to the PC via a LAN port, and the first DUT and the second DUT are connected via a WAN port.

[0044] The first device under test is configured to send a probe message to the second device under test, and receive a feedback message returned by the second device under test in response to the probe message, and report the link quality data of each of the plurality of links under test to the PC device according to the probe message and the feedback message;

[0045] The second device under test is used to receive the probe message and send the feedback message to the first device under test;

[0046] The PC device performs clustering processing on the link quality data of each of the multiple links under test based on a pre-trained clustering algorithm model. The result of the clustering processing represents the link quality characteristics of each of the multiple links under test.

[0047] Optionally, the first device under test (DUT) accesses and obtains its IP address through the PC; the WAN port of the second DUT is configured with an IP address on the same network segment as the first DUT but a different IP address.

[0048] The first device under test has a first routing detail, and the first routing detail accesses the IP address of the second device under test through the WAN port of the second device under test.

[0049] The second device under test has a second routing detail added, and the second routing detail accesses the IP address of the first device under test through the WAN port of the second device under test;

[0050] The system also includes a WAN port verification module, used to verify whether the WAN port of the first device under test can ping the IP address of the WAN port of the second device under test, and to verify whether the WAN port of the second device under test can ping the IP address of the WAN port of the first device under test.

[0051] Optionally, the first device under test is configured with a first configuration file, and the second device under test is configured with a second configuration file;

[0052] The first configuration file is used to create the same number of link UUIDs for the multiple links under test, and to set the end probe IP that is the same as the IP address of the second device under test, wherein the link UUIDs correspond one-to-one with the multiple links under test;

[0053] The second configuration file is used to create a link UUID number that corresponds one-to-one with the link UUID number configured in the first configuration file, and to set an end probe IP that is the same as the IP address of the first device under test.

[0054] The technical solutions provided by the embodiments of this disclosure have at least the following beneficial effects:

[0055] This disclosure employs a bidirectional communication mechanism of probe and feedback messages, effectively simplifying the complexity of the SD-WAN testing environment and reducing testing costs. Through real-time data exchange between the first and second devices under test (DUT), it accurately acquires quality data for each link under test. A pre-trained clustering algorithm model is then used to cluster the link quality data, enabling rapid identification of normal and abnormal link states, reducing manual intervention and complex manual analysis processes. Therefore, this disclosure significantly improves the efficiency of SD-WAN testing, reduces enterprises' investment in network management and maintenance, and provides strong support for enterprise digital transformation. Attached Figure Description

[0056] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0057] Figure 1 This is a schematic diagram of a typical SD-WAN application scenario provided by one embodiment of the present disclosure;

[0058] Figure 2 This is a structural block diagram of a multi-link testing system based on SD-WAN provided in one embodiment of the present disclosure;

[0059] Figure 3 This is a schematic diagram of the network environment configuration process of the first device under test and the second device under test in one embodiment of this disclosure;

[0060] Figure 4 This is a schematic diagram illustrating the steps of a multi-link testing method based on SD-WAN provided in one embodiment of this disclosure;

[0061] Figure 5 This is a schematic diagram of the interaction steps between the first device under test and the second device under test in one embodiment of this disclosure;

[0062] Figure 6 This is a flowchart of clustering model training and application in one embodiment of this disclosure;

[0063] Figure 7 This is a schematic diagram of the data distribution of a clustering result in one embodiment of this disclosure;

[0064] Figure 8 This is a schematic diagram of the data distribution of another clustering result in one embodiment of this disclosure. Detailed Implementation

[0065] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this disclosure. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0066] Figure 1 This is a schematic diagram illustrating a typical SD-WAN application scenario provided by an embodiment of this disclosure. Please refer to [link / reference]. Figure 1In typical SD-WAN application scenarios, the central node acts as the core of the network, responsible for efficient communication and data transmission with multiple branch nodes. These branch nodes can include cloud service providers, data centers, and geographically dispersed branch offices. Through SD-WAN technology, enterprises can achieve flexible network connectivity, optimize data flow, and improve network reliability and performance. In this network architecture, the central node communicates with multiple branch nodes in real time via network connections, ensuring rapid data transmission between different locations. For example, a branch node can be a cloud service platform, providing the computing and storage resources needed by the enterprise; it can also be a data center, responsible for storing and processing large amounts of data; furthermore, branch offices are the front line of the enterprise's daily operations, and through SD-WAN connections, employees can access the central database and applications.

[0067] In related technologies, each node typically needs to be equipped with a UCPE (Unified Edge Computing Platform) box to support SD-WAN functionality. However, each branch node requires a separate UCPE box, which not only increases initial investment costs but also leads to higher subsequent maintenance and upgrade expenses. As the number of branch nodes increases, the number of UCPE devices that enterprises need to manage and maintain also increases, making the network architecture more complex. Each UCPE box requires configuration, monitoring, and troubleshooting, further increasing the complexity of network management.

[0068] To address the shortcomings of the aforementioned technologies, this disclosure aims to reduce the hardware requirements of each node. Through centralized management and intelligent scheduling, effective management and monitoring of branch links or nodes can be achieved without the need for separate deployment of physical equipment. Figure 2 This is a structural block diagram of a multi-link testing system based on SD-WAN, provided in one embodiment of the present disclosure. The system is applied to the testing method described below. The system includes: a first device under test (DUT), a second DUT, and a PC. The first DUT and the second DUT are each configured with multiple DUT links.

[0069] The first device under test (DUT) acts as the master node of the network, responsible for actively initiating probe requests. The second DUT acts as a branch node in the network, primarily responsible for receiving probe packets from the first DUT and returning feedback packets. The link under test refers to the network connection configured between the first and second DUTs; this can be a physical connection or a virtual network connection (such as fiber optic, Ethernet, VPN, etc.), configured on the first and second DUTs respectively. Each link under test represents an independent network path, and performance testing is performed through these links to evaluate their link quality.

[0070] The first device under test (DUT) is connected to the PC via a LAN port, and the first DUT and the second DUT are connected via a WAN port.

[0071] Figure 3 This is a schematic diagram of the network environment configuration process for the first device under test and the second device under test in one embodiment of this disclosure, as shown below. Figure 3 As shown, the first device under test (DUT) is connected to the PC via its LAN (Local Area Network) port. Through the LAN port, the DUT can transmit data with the PC. The PC can access the DUT's management interface via the network to perform operations such as setting up, monitoring traffic, and viewing status. During the test, the PC acts as a control console, sending test commands, collecting data, and analyzing results.

[0072] The first and second devices under test (DUTs) are connected via a WAN (Wide Area Network) port. This WAN connection interconnects the two DUTs, creating a wide area network environment to simulate a real network environment and test the device performance under different network conditions. Through the WAN port, the first and second DUTs can send and receive data packets to each other.

[0073] It should be noted that if users need to control the second device under test via a PC to exchange messages and report data, they can also choose to connect the second device under test to the PC via a USB port.

[0074] The first device under test (DUT) is configured to send a probe message to the second DUT, and receive a feedback message returned by the second DUT in response to the probe message, and report the link quality data of each of the plurality of DUT links to the PC device according to the probe message and the feedback message.

[0075] To assess and monitor link quality, a first device under test (DUT) sends probe messages to a second DUT. The probe message includes the sender's timestamp and message count within the service period. Upon receiving the probe message, the second DUT processes it accordingly and returns a feedback message to the first DUT. The first DUT can then obtain link quality data characterizing the link's performance based on the probe message and the feedback message.

[0076] The second device under test is used to receive the probe message and send the feedback message to the first device under test.

[0077] The feedback message includes the receiver's timestamp and message count within the service cycle.

[0078] The following example illustrates how to report link quality data of the link under test to a PC device based on probe and feedback messages. Assume that Device A, the first device under test, sends a probe message to Device B, the second device under test. The message includes: timestamp (T1): 2023-10-01 10:00:00; message count (Seq): 1 (indicating this is the first probe message). Device A sends the probe message to Device B. After receiving the probe message, Device B records the reception timestamp (T2) and processes it. Device B returns a feedback message to Device A, including: reception timestamp (T2): 2023-10-01 10:00:01; message count (Seq): 1 (indicating this is the feedback corresponding to the probe message). After receiving the feedback message, the first device under test can derive the link quality data based on the content of the probe and feedback messages. In this embodiment, the link quality data includes: latency, jitter, and packet loss rate. Suppose that within a single business cycle, Device A sends 10 probe messages, while Device B only successfully receives 8 feedback messages. Packet loss rate = (Total number of messages sent - Total number of feedback messages received) / Total number of messages sent × 100%. Total number of messages sent = 10; Total number of feedback messages received = 8; Therefore, packet loss rate = (10 - 8) / 10 × 100% = 20%. Now suppose Device A sends 3 probe messages within one business cycle, and the timestamps are recorded as follows:

[0079] T1 = 2023-10-01 10:00:00 (Feedback time T2 = 2023-10-01 10:00:01);

[0080] T3 = 2023-10-01 10:00:02 (Feedback time T4 = 2023-10-01 10:00:03);

[0081] T5 = 2023-10-01 10:00:04 (Feedback time T6 = 2023-10-01 10:00:05); Calculate the latency of each probe message:

[0082] First message delay = 1 second; Second message delay = 1 second; Third message delay = 1 second

[0083] Jitter is the change in latency, and the standard deviation of latency is calculated as follows:

[0084] Average delay = (1 + 1 + 1) / 3 = 1 second;

[0085] Standard deviation = = 0 (because all delays are the same), where N is the number of packets = 3. Therefore, the jitter in this example is 0 seconds.

[0086] Through the above steps, the link quality data obtained are as follows: latency 1 second, jitter 0 seconds, and packet loss rate 20%.

[0087] In this embodiment, to ensure normal testing and data transmission of the link, the first and second devices under test (DUTs) need to consider the impact of Network Address Translation (NAT) settings during configuration. If the first DUT has a NAT configuration, data packets may be modified when passing through the device, thus affecting the link's performance test results. Therefore, in this case, the NAT configuration of the first DUT needs to be removed to ensure that data can be transmitted directly through the device, maintaining the original state of the link. Similarly, if the second DUT also has a NAT configuration, it also needs to be removed.

[0088] The PC device performs clustering processing on the link quality data of each of the multiple links under test based on a pre-trained clustering algorithm model. The result of the clustering processing represents the link quality characteristics of each of the multiple links under test.

[0089] By collecting link quality data from the first and second devices under test, the PC device uses a clustering algorithm to analyze and classify this data. The clustering results effectively characterize the quality features of each link under test, allowing network administrators to intuitively identify performance differences between links, discover potential abnormal links, and make corresponding optimizations and adjustments based on this information. The clustering algorithm and its training process will be described in detail in the testing method section below.

[0090] This disclosure, through the collaborative operation of the first and second devices under test (DUTs) and cluster analysis of the PC device, can comprehensively evaluate the quality characteristics of network links, including link quality data such as latency, jitter, and packet loss rate. It is evident that this disclosure can achieve real-time link quality monitoring using only three devices, and further identifies link performance differences and potential anomalies through data clustering, which helps optimize network configuration, improve network stability and transmission efficiency, thereby meeting the high-performance requirements of complex network environments.

[0091] In one optional implementation, the first device under test (DUT) accesses and obtains its IP address through the PC; the WAN port of the second DUT is configured with an IP address on the same network segment as the first DUT but a different IP address; the first DUT has a first routing detail added to it, and the first routing detail accesses the IP address of the second DUT through its WAN port; the second DUT has a second routing detail added to it, and the second routing detail accesses the IP address of the first DUT through its WAN port.

[0092] like Figure 3 As shown, in this embodiment, the first device under test (DUT) is connected to a PC and obtains its own IP address via its LAN port using a network protocol (such as DHCP or static configuration). The PC acts as a management terminal, accessing the network settings of the first DUT through a command line or graphical interface to obtain its assigned IP address. The WAN port of the second DUT is configured to be on the same network segment as the first DUT, but with a different IP address. That is, the IP addresses of the two devices are within the same subnet, enabling them to recognize and communicate with each other. For example, if the IP address of the first DUT is 192.168.1.2, the WAN port of the second DUT can be set to 192.168.1.3. The first DUT adds a route detail to its routing table, pointing to the IP address of the second DUT. This means the first DUT can send data packets to the WAN port of the second DUT through its WAN port. The first route detail ensures the correctness of the data flow, enabling the first DUT to effectively access the resources of the second DUT or perform link quality testing. Similarly, the second device under test adds a second route detail to its routing table, pointing to the IP address of the first device under test, to ensure that the second device under test can access the resources of the first device under test through its WAN port.

[0093] The system also includes a WAN port verification module, used to verify whether the WAN port of the first device under test can ping the IP address of the WAN port of the second device under test, and to verify whether the WAN port of the second device under test can ping the IP address of the WAN port of the first device under test.

[0094] The WAN port verification module verifies the reachability between the WAN ports of two devices under test by sending PING requests. PING is a network diagnostic tool that tests the validity of a network connection by sending an ICMP (Internet Control Message Protocol) echo request to a target IP address and waiting for an echo response. First, the WAN port verification module sends a PING request from the WAN port of the first device under test to the IP address of the WAN port of the second device under test. Specifically, it determines the WAN port IP address of the second device under test, sends a PING request to that IP address via the ICMP protocol, and waits to receive an echo response from the second device under test. If the WAN port of the second device under test successfully receives the PING request and returns a response, the WAN port verification module records the connection as "reachable". If no response is received, it is recorded as "unreachable". Next, the WAN port verification module performs the reverse operation, sending a PING request from the WAN port of the second device under test to the IP address of the WAN port of the first device under test. Specifically, it determines the WAN port IP address of the first device under test, sends a PING request to that IP address via the ICMP protocol, and waits to receive an echo response from the first device under test. Similarly, if the WAN port of the first device under test can successfully receive the PING request and return a response, the WAN port verification module will record the connection as "reachable". If no response is received, it will be recorded as "unreachable".

[0095] In one optional implementation, the first device under test (DUT) is configured with a first configuration file, and the second DUT is configured with a second configuration file. The first configuration file is used to create the same number of link UUIDs for the plurality of links under test, and to set an end probe IP that is the same as the IP address of the second DUT, wherein the link UUIDs correspond one-to-one with the plurality of links under test. The second configuration file is used to create link UUIDs that correspond one-to-one with the link UUIDs configured in the first configuration file, and to set an end probe IP that is the same as the IP address of the first DUT.

[0096] like Figure 3As shown, in either the first or second configuration file, a unique link UUID is generated for each link under test. A UUID (Universally Unique Identifier) ​​is a standard identifier that ensures each link is unique within the system, with a one-to-one correspondence between the link UUID and the link under test. In the first link configuration file, an end probe IP address is set to be the same as the IP address of the second device under test. That is, the first device under test will use the IP address of the second device under test as the probe target for link quality testing. Similarly, in the second link configuration file, an end probe IP address is set to be the same as the IP address of the first device under test. The second device under test will use the IP address of the first device under test as the probe target.

[0097] In one optional implementation, the QoS yang file of the first device under test is configured with attribute information for each of the plurality of links under test for the link UUID number. The first device under test monitors the configuration of the QoS yang file through a monitoring switch. The attribute information includes link priority and bandwidth limitation. The QoS yang file of the second device under test is configured with attribute information for each of the plurality of links under test for the link UUID number. The second device under test monitors the configuration of the QoS yang file through a monitoring switch.

[0098] like Figure 3 As shown, in this embodiment, the QoS (Quality of Service) yang files of the first and second devices under test are configured for their respective link UUIDs to define the attribute information of multiple links under test. The attribute information includes link priority and bandwidth limits to ensure the reasonable allocation and optimized use of network resources. The first device under test monitors the configuration status of its QoS yang file in real time through a monitoring switch to adjust the link's QoS settings promptly; similarly, the second device under test also monitors its corresponding QoS yang file configuration through its monitoring switch.

[0099] The testing system provided in this embodiment, after being configured in the aforementioned network environment, can be applied to perform the steps described below. Figure 4 This is a schematic diagram illustrating the steps of a multi-link testing method based on SD-WAN provided in one embodiment of this disclosure, as follows: Figure 4 As shown, the method includes:

[0100] Step S101: The first device under test sends a probe message to the second device under test and receives a feedback message returned by the second device under test in response to the probe message. The first device under test and the second device under test are each configured with multiple test links.

[0101] The first device under test (DUT) sends a probe message to the second DUT. The probe message includes the sender's timestamp and message count within the service period. After receiving the probe message, the second DUT processes it accordingly and returns a feedback message to the first DUT.

[0102] Step S102: Determine the link quality data of each of the plurality of links under test based on the probe message and the feedback message.

[0103] As mentioned above, the link quality data is an indicator used to evaluate the performance of the link under test. In this embodiment, the link quality data includes at least latency, jitter, packet loss rate, and current remaining bandwidth.

[0104] Figure 5 This is a schematic diagram of the interaction steps between the first device under test and the second device under test in one embodiment of this disclosure, as shown below. Figure 5 As shown, steps S101 to S102 are executed cyclically in the case of a multi-link scenario in this embodiment until the preset clock cycle condition is met.

[0105] Step S103: Based on the pre-trained clustering algorithm model, cluster the link quality data of each of the multiple links to be tested to obtain multiple clusters for characterizing link quality and the convergence regions of each of the multiple clusters.

[0106] The clustering algorithm described is an unsupervised learning method designed to group objects in a dataset based on their features. The pre-trained clustering algorithm model has learned how to identify and group similar link quality features using historical link quality data.

[0107] The collected link quality data is input into a clustering algorithm model, which analyzes and groups this data. Links with similar quality characteristics are grouped into the same category, while those with different characteristics are assigned to different clusters. Through clustering, multiple clusters are generated, each representing a group of links with similar link quality characteristics.

[0108] Each cluster has its own convergence region, which represents the normal range of link quality data within that cluster. If a data point falls within the convergence region, it indicates that its quality performance is normal; if it falls outside the convergence region, it indicates that there is an anomaly in the link.

[0109] Step S104: Link quality data that does not belong to the convergence area is identified as abnormal data, and abnormal data is analyzed.

[0110] Examine the quality data of all links under test to determine which data points are outside the convergence region of their corresponding clusters. These data points outside the convergence region are marked as anomalous data, indicating that the quality performance of these links differs significantly from normal conditions. For the identified anomalous data, conduct detailed anomaly analysis. Specifically, analyze the causes of the anomalous data, combining information such as network topology, device status, and traffic monitoring to identify the specific factors leading to the degraded link quality. Alternatively, observe the time-series changes of the anomalous data to determine whether it is an occasional event or a persistent problem, analyzing its frequency and patterns. Furthermore, compare anomalous links with normal links to find differences, helping to identify potential root causes of failures.

[0111] This disclosure employs a bidirectional communication mechanism of probe and feedback messages, effectively simplifying the complexity of the SD-WAN testing environment and reducing testing costs. Through real-time data exchange between the first and second devices under test (DUT), it accurately acquires quality data for each link under test. A pre-trained clustering algorithm model is then used to cluster the link quality data, enabling rapid identification of normal and abnormal link states, reducing manual intervention and complex manual analysis processes. Therefore, this disclosure significantly improves the efficiency of SD-WAN testing, reduces enterprises' investment in network management and maintenance, and provides strong support for enterprise digital transformation.

[0112] In one optional implementation, after receiving the feedback message returned by the second device under test in response to the probe message, the method further includes:

[0113] Based on the probe messages and the feedback messages, the actual bandwidth of each of the multiple links under test is determined;

[0114] For any link to be tested, compare the actual bandwidth of the link to be tested with a preset first threshold.

[0115] When the actual bandwidth is greater than or equal to a preset first threshold, a data stream is sent to the first device under test for the link under test using the NETPERF streaming tool of the second device under test.

[0116] The SAR traffic monitoring tool of the first device under test is used to determine the current remaining bandwidth of the link under test, and the current remaining bandwidth is used as the link quality data of the link under test.

[0117] Here, in addition to the timestamp and message count mentioned above, the probe message may also include the packet size. Please see [link / reference]. Figure 5For links with bandwidth requirements, the NETPERF streaming tool on the second device under test (DUT) is used to send a data stream to the first DUT for that link. For example, assuming that Device A and Device B have been successfully configured, for a specific link under test, Device A sends a probe message to Device B. For instance, Device A sends a 1000-byte data packet with timestamp T1. After receiving the probe message, Device B processes it and returns a feedback message containing information such as the timestamp of the received data packet and the reception status. Assuming Device B successfully receives the data packet at timestamp T2 (e.g., T2 = T1 + 0.1 seconds) and returns the feedback message to Device A, Device A calculates the actual bandwidth based on the probe message and the feedback message. Assuming the information in the feedback message indicates that the data packet was successfully transmitted within 0.1 seconds, the actual bandwidth can be calculated as follows:

[0118]

[0119] Assume the preset first threshold is 5 kbps. Device A compares the actual bandwidth of 10 kbps with the preset threshold of 5 kbps. Since 10 kbps is greater than or equal to 5 kbps, it indicates that the link under test has bandwidth requirements. Device B uses the NETPERF streaming tool to send a series of data streams to Device A for this link. For example, Device B sends a data stream to Device A at a rate of 5 kbps for a period of time. Simultaneously, Device A uses the SAR traffic monitoring tool (SystemActivity Reporter) to monitor the link's traffic. Assume Device A detects that the current total bandwidth of the link under test is 20 kbps, and the bandwidth currently in use is 5 kbps (from the data stream from Device B). The current remaining bandwidth is: Current remaining bandwidth = Total bandwidth - Bandwidth currently in use = 20 kbps - 5 kbps = 15 kbps. Device A records the current remaining bandwidth of 15 kbps as the link quality data for the link under test, which will be used for subsequent performance evaluation and monitoring.

[0120] Through the above steps, the first and second devices under test can effectively evaluate and monitor the link performance between them, helping network administrators to understand the bandwidth usage of the link in real time based on the link's bandwidth requirements, so as to take timely measures to ensure the efficient operation of the network.

[0121] In an optional implementation, the method further includes:

[0122] The CPU utilization and memory usage of each of the multiple test links are determined using the SAR traffic monitoring tool of the first device under test.

[0123] From the plurality of test links, identify the test links that are exclusively used by the CPU with a CPU utilization rate higher than a preset second threshold, and identify the test links that have memory leaks with a memory utilization rate higher than a preset third threshold.

[0124] For the test link that is CPU-exclusive or has memory leaks, the scheduling priority and resource allocation strategy of the test link are adjusted in the QOS yang files of the first and second test devices, respectively.

[0125] The SAR traffic monitoring tool on the first device under test (DUT) is used to obtain real-time CPU and memory usage for each DUT link. The SAR traffic monitoring tool periodically monitors the resource consumption of the first DUT, providing real-time CPU and memory consumption for each DUT link. For example:

[0126] Link A under test: CPU utilization 85%, memory utilization 60%.

[0127] Link B under test: CPU utilization 95%, memory utilization 50%.

[0128] Link C under test: CPU utilization 40%, memory utilization 85%.

[0129] Based on a preset second threshold (e.g., 90%), abnormal links are identified. Links with CPU utilization exceeding the preset second threshold are identified as CPU-exclusive links under test. Links with memory utilization exceeding a preset third threshold (e.g., 80%) are identified as memory-leaking links under test. For link B under test, with a CPU utilization of 95%, exceeding the 90% threshold, it is marked as a CPU-exclusive link under test. For link C under test, with a memory utilization of 85%, exceeding the 80% threshold, it is marked as a memory-leaking link under test.

[0130] The aforementioned memory leak and CPU-exclusive links are the bottlenecks causing system performance. Adjustments should be made to the QoS yang files of the first and second devices under test, respectively. For the CPU-exclusive link, reduce its scheduling priority and decrease resource allocation to prevent CPU overload from affecting the normal operation of other links. For the memory leak link, increase its scheduling priority to ensure sufficient resources are allocated to handle memory pressure. If the problem persists, an alarm or restart mechanism can be set up.

[0131] In this example, for link B under test (CPU exclusive), the QoS yang files of both devices are modified to lower the scheduling priority from high to low. The CPU resource usage limit for this link is limited to 80%. For link C under test (memory leak), its scheduling priority is increased to the highest level, the memory allocation strategy is dynamically adjusted, more memory space is reserved, and task interruption is prevented.

[0132] This disclosure automatically adjusts link scheduling strategies by monitoring resource usage in real time. This avoids single points of exhaustion of CPU and memory resources, ensuring the stability and performance of the overall system. At the same time, by adjusting link priorities, it improves the operating efficiency of critical links and reduces link latency and packet loss rate.

[0133] In an optional implementation, the method further includes:

[0134] In the configuration files of the first device under test and the second device under test, the link UUID number is increased to a preset link limit, and the link UUID number corresponds one-to-one with the plurality of links under test;

[0135] The SAR traffic monitoring tool of the first device under test is used to monitor the link quality data of each of the multiple links under test. The link quality data includes: latency, jitter and packet loss rate.

[0136] Based on the link quality data of each of the multiple links under test, determine the interface status between the first device under test and the second device under test;

[0137] When the interface status indicates that the interface is blocked, the scheduling priority and resource allocation strategy of the multiple test links are adjusted accordingly in the QoS yang files of the first and second test devices.

[0138] In the configuration files of both the first and second devices under test (DUTs), the UUIDs (Globally Unique Identifiers) of the links are increased to the preset link limit. Each UUID corresponds one-to-one with a link under test, ensuring that the system can accurately identify and manage each link. For example, assuming the preset link limit is 10, a unique UUID will be assigned to each link, such as uuid1, uuid2, ..., uuid10. The SAR traffic monitoring tool on the first DUT is used to monitor the link quality data of each link under test. The calculation and monitoring process for latency, jitter, and packet loss rate has been explained in detail above and will not be repeated here.

[0139] Based on the monitored link quality data, the interface status between the first and second devices under test is determined. Interface status includes normal, congested, and faulty. If the monitored link quality data (such as high latency, high jitter, or high packet loss rate) indicates interface congestion, the interface status will be identified as congested. Assume the monitoring results are as follows: Link UUID1: Latency = 20ms, Jitter = 5ms, Packet Loss = 0%;

[0140] Link UUID2: Latency = 50ms, Jitter = 10ms, Packet Loss Rate = 2%;

[0141] Link UUID3: Latency = 30ms, Jitter = 3ms, Packet Loss Rate = 1%; It is evident that link UUID2 has a packet loss rate of 2%, and its latency and jitter are also high, indicating the interface status is "congested". Given this congestion status, the scheduling priority and resource allocation strategies for multiple links under test will be adjusted in the QoS yang files of both the first and second devices under test. The scheduling priority of certain critical links will be increased to ensure the smooth transmission of important data flows. More bandwidth resources will be allocated to high-priority links, reducing bandwidth allocation to other low-priority links. For example, the priority of link UUID1 might be increased from normal to high priority, while the bandwidth allocation for link UUID2 is reduced to alleviate congestion.

[0142] By following the steps described above, the quality of multiple links under test can be effectively monitored and managed, interface congestion can be identified in a timely manner, and link priorities and resource allocation strategies can be adjusted according to the actual situation. This dynamic adjustment mechanism can improve the overall performance and stability of the network, ensuring the smooth operation of critical services.

[0143] In one alternative implementation, the clustering algorithm model is pre-trained according to the following steps:

[0144] Collect multiple historical link quality data of multiple links under test at different time periods, and input the multiple historical link quality data as training samples into the clustering algorithm model;

[0145] From the multiple historical link quality data, N historical link quality data are randomly selected as the first cluster centroid of each of the N clusters, where N is an integer greater than 1;

[0146] Calculate the distance from each historical link quality data point to the cluster centroid, and assign the multiple historical link quality data points to the nearest cluster;

[0147] Based on the average value of historical link quality data within the same cluster, the cluster centroid is iteratively updated until the difference between the changes of the cluster centroids of two adjacent updates is less than a preset change threshold, thus obtaining N convergent clusters.

[0148] Based on the centroids of the N convergent clusters, the convergence regions of the N clusters are determined.

[0149] Figure 6 This is a flowchart of clustering model training and application in one embodiment of this disclosure. Please refer to [link / reference]. Figure 6 In this embodiment, the clustering algorithm model is an unsupervised learning-optimized K-means algorithm model. During its pre-training process, historical link quality data of multiple test links are collected at different time periods. The historical link quality data may include latency, jitter, and packet loss rate. For example, suppose the following historical link quality data (each data point includes latency, jitter, and packet loss rate) were collected at different time periods:

[0150] Data 1: Latency = 20ms, Jitter = 2ms, Packet Loss Rate = 0%;

[0151] Data 2: Latency = 30ms, Jitter = 5ms, Packet Loss Rate = 1%;

[0152] Data 3: Latency = 25ms, Jitter = 3ms, Packet Loss Rate = 0%;

[0153] Data 4: Latency = 40ms, Jitter = 7ms, Packet Loss Rate = 2%;

[0154] Data 5: Latency = 35ms, Jitter = 6ms, Packet Loss Rate = 1%;

[0155] Data 6: Latency = 50ms, Jitter = 10ms, Packet Loss Rate = 3%; N data points are randomly selected from the collected historical link quality data as the initial centroids for clustering, where N is an integer greater than 1. Here, N=2 can be chosen. Correspondingly, the randomly selected initial centroids are: Cluster 1 centroid, Data 1 (20ms, 2ms, 0%); Cluster 2 centroid, Data 4 (40ms, 7ms, 2%). The distance from each historical link quality data point to the cluster centroid (Euclidean distance can be used) is calculated, and the data is assigned to the nearest cluster. For example, calculating the distance from Data 1 to the two centroids: the distance to the Cluster 1 centroid (20, 2, 0) is 0 (because it is the centroid). The distance to the Cluster 2 centroid (40, 7, 2) is:

[0156]

[0157] Repeat this process, calculating the distance of other data points to the two centroids, and finally assigning the data to the nearest cluster. Iteratively update the cluster centroids based on the average historical link quality data within the same cluster. Assume that after the first assignment, data 1, data 2, and data 3 are assigned to cluster 1, and data 4, data 5, and data 6 are assigned to cluster 2. The new centroid 1 for cluster 1 and the new centroid 2 for cluster 2 are as follows:

[0158]

[0159] Repeat the steps of calculating distances and assigning clusters until the difference between the changes in the centroids of two consecutive updates is less than a preset threshold (e.g., 0.1). Assuming that after multiple iterations, the centroids no longer change significantly, N convergent clusters are obtained. Based on the centroids of each of the N convergent clusters, the convergence regions of each of the N clusters are determined. The convergence region can be defined by setting a radius, which can be based on the average distance from each cluster's data point to its centroid, or on the standard deviation (the square of the difference between each distance and the average distance, then the average of the squared differences and the square root).

[0160] In an optional implementation, after clustering the link quality data of the plurality of links under test based on the pre-trained clustering algorithm model, the method further includes: determining the centroid of each of the plurality of clusters, and the link quality index corresponding to each centroid, wherein the link quality index is a one-dimensional index; mapping the link quality data of the plurality of links under test to the centroid of the cluster to obtain the link quality index of the plurality of links under test, wherein the link quality data includes at least: latency, jitter, and packet loss rate.

[0161] The following example illustrates how to cluster link quality data of the link under test based on a pre-trained clustering algorithm model, and determine the cluster centroids and their corresponding link quality metrics. Assume that two clusters have been obtained through the clustering algorithm, and the centroid of each cluster has been determined. As shown in Table 1, the following three-dimensional link quality data (latency, jitter, and packet loss rate) of the link under test are available:

[0162]

[0163] Table 1

[0164] Please see Figure 6 After clustering, the following two clusters and their centroids were obtained: Cluster centroid 1 (better performance), (25ms, 3ms, 0.33%); Cluster centroid 2 (poorer performance), (37.5ms, 6.5ms, 1.5%).

[0165] The distance from link 1 to cluster 1 is

[0166] The distance from link 1 to cluster 2 is

[0167] Similarly, calculate the distances of other links to the centroids of the two clusters. Based on the calculated distances, select the cluster with the smaller distance as the link's affiliation. Therefore, link 1 is closer to cluster 1 and belongs to cluster 1. Link 2 is closer to cluster 1 and belongs to cluster 1. Link 3 is closer to cluster 1 and belongs to cluster 1. Link 4 is closer to cluster 2 and belongs to cluster 2. Link 5 is closer to cluster 2 and belongs to cluster 2.

[0168] Figure 7 This is a schematic diagram of the data distribution of a clustering result in one embodiment of this disclosure, such as... Figure 7 As shown, the area enclosed by the dashed circular boundary is the convergence region corresponding to the centroid after convergence. The figure illustrates five different clusters, showing the spatial distribution of their respective three-dimensional link quality data. Figure 7 The link quality data in the leftmost cluster represents a higher packet loss rate, while the link quality data in the rightmost cluster represents a higher latency.

[0169] To map the three-dimensional link quality data of each link to a one-dimensional link quality metric, weights can be assigned to each dimension. Let's assume the weights assigned to latency, jitter, and packet loss rate are respectively... It can be defined as:

[0170] For weighted calculation, the packet loss rate needs to be normalized, resulting in the clustering results shown in Table 2:

[0171]

[0172] Table 2

[0173] Figure 8 This is a schematic diagram of the data distribution of another clustering result in one embodiment of this disclosure, such as... Figure 8 As shown, the three-dimensional link quality data is quantified into one-dimensional link quality, which can more intuitively reflect the link quality of each link under test. For example, it can be seen that the link quality data (probe data) in the rightmost cluster represents a link with worse quality. The results shown in Table 2 are obtained by weighting the actual link quality data. It can be understood that the higher the value of the link quality data, the higher the link quality index, and the worse the link quality. That is to say, link quality data and link quality have an inverse correlation. In order to positively reflect the link quality, the reciprocal of the weighted link quality index results should be taken to more clearly reflect the link quality. For example, the link quality index of link 1 is the lowest, and the value is the highest after taking the reciprocal. Therefore, compared with the other links, link 1 reflects the best link quality and represents the best link performance.

[0174] The method further includes: monitoring whether there is a faulty link among the plurality of links to be tested; if there is a faulty link, determining the cluster centroid of the cluster to which the faulty link belongs; from the cluster to which the faulty link belongs, determining the backup link that is closest to the abnormal link in terms of link quality index and whose current remaining bandwidth meets the preset bandwidth requirement; and switching the faulty link to the backup link.

[0175] Continuously monitor the status of all links under test. Once a link is identified as faulty, find its cluster and determine its centroid. Specifically, extract the faulty link's latency, jitter, and packet loss rate. The distance from the faulty link to each cluster centroid can be calculated using weighted distance (as described above). Assign the faulty link to the nearest cluster and record its centroid. Once the faulty link's cluster and its centroid are determined, find suitable backup links within that cluster. Iterate through all links belonging to that cluster; all links in that cluster are multiple candidate backup links for the faulty link. Calculate the distance from each link to the faulty link's link quality metric. For each candidate backup link, check if its current remaining bandwidth meets the preset bandwidth requirement. From the links that meet the bandwidth requirement, select the link closest to the faulty link's quality metric as the backup link. Finally, perform link switching, transferring the faulty link to the selected backup link.

[0176] By following the steps above, link failures can be effectively monitored, the cluster to which the failed link belongs can be quickly determined, and a suitable backup link can be selected from the cluster for switching, which significantly improves the reliability and stability of the network and ensures the continuity of services.

[0177] In an optional implementation, after clustering the link quality data of the multiple links under test based on a pre-trained clustering algorithm model, the method further includes: collecting multiple historical link quality data of the multiple links under test at different time periods to form a historical dataset; determining the link quality evolution trend of the multiple links under test based on the historical dataset; identifying potential fault links whose link quality data exceeds a set tolerance range for multiple consecutive time periods based on the clustering results and link quality evolution trends of the link quality data, and performing result-oriented analysis on the potential fault links.

[0178] Using the first device under test (DUT), quality data for each link under test should be collected periodically. Data should be collected at different time intervals to create a time series. Please refer to [link to relevant documentation]. Figure 6After collecting sufficient historical link quality data, potential faulty links are identified by combining clustering results and link quality evolution trends. Specifically, the clustering results obtained from the previous clustering algorithm model are reviewed to understand the current status of each link. Tolerance ranges for each link's quality data are set based on business requirements and network performance standards. For example, latency should be less than 50ms, jitter less than 10ms, and packet loss rate less than 1%. The quality data of each link is monitored to determine whether it exceeds the set tolerance ranges over multiple time periods. If a link's quality indicators exceed the tolerance ranges for multiple consecutive time periods, it is marked as a potentially faulty link.

[0179] Once a potentially faulty link is identified, a thorough, results-oriented analysis is conducted. Specifically, historical data for the potentially faulty link is analyzed in detail to identify the causes of link quality degradation. For example, analyzing changes in latency, jitter, and packet loss rate helps identify whether specific time periods or events lead to performance degradation. The relationship between the potentially faulty link and other links can also be examined to see if there are common problems, such as network congestion or equipment failure. Based on the analysis results, improvement suggestions are proposed for the potentially faulty link, such as adjusting routes, increasing bandwidth, or optimizing configurations.

[0180] Regarding step S104, Figure 8 The link quality distribution results shown also clearly show that the location of the anomaly detection data points is outside all convergence ranges, allowing for further analysis of the anomaly data or the link corresponding to it.

[0181] By following the steps outlined above, we can effectively analyze link quality evolution trends and identify potentially faulty links. This method not only helps network administrators identify problems promptly but also provides data support for subsequent troubleshooting and performance optimization, thereby improving network reliability and performance.

[0182] Those skilled in the art will understand that embodiments of this disclosure can be provided as methods, apparatus, electronic devices, and storage media. Therefore, embodiments of this disclosure can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, embodiments of this disclosure can take the form of a computer program product embodied on one or more computer-readable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0183] While preferred embodiments of the present disclosure have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the present disclosure.

[0184] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the term "comprising" or any other variations thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element. The above provides a detailed description of a multi-link testing method and system based on SD-WAN provided by this disclosure. Specific examples have been used to illustrate the principles and implementation methods of this disclosure. The descriptions of the above embodiments are only for the purpose of helping to understand the method and its core ideas; at the same time, for those skilled in the art, there will be changes in specific implementation methods and application scope based on the ideas of this disclosure. Therefore, the content of this specification should not be construed as a limitation of this disclosure.

Claims

1. A multi-link testing method based on SD-WAN, characterized in that, The method includes: The first device under test (DUT) sends a probe message to the second device under test (DUT) and receives a feedback message from the second device under test in response to the probe message. The first DUT and the second DUT are each configured with multiple links under test. Based on the probe messages and the feedback messages, the link quality data of each of the multiple links under test is determined; Based on the pre-trained clustering algorithm model, the link quality data of each of the multiple links to be tested are clustered to obtain multiple clusters for characterizing link quality and the convergence regions of each of the multiple clusters. Link quality data that does not belong to the convergence region is identified as abnormal data, and anomaly analysis is performed on the abnormal data.

2. The method according to claim 1, characterized in that, After receiving the feedback message returned by the second device under test in response to the probe message, the method further includes: Based on the probe messages and the feedback messages, the actual bandwidth of each of the multiple links under test is determined; For any link to be tested, compare the actual bandwidth of the link to be tested with a preset first threshold. When the actual bandwidth is greater than or equal to a preset first threshold, a data stream is sent to the first device under test for the link under test using the NETPERF streaming tool of the second device under test. The SAR traffic monitoring tool of the first device under test is used to determine the current remaining bandwidth of the link under test, and the current remaining bandwidth is used as the link quality data of the link under test.

3. The method according to claim 2, characterized in that, The method further includes: The CPU utilization and memory usage of each of the multiple test links are determined using the SAR traffic monitoring tool of the first device under test. From the plurality of test links, identify the test links that are exclusively used by the CPU with a CPU utilization rate higher than a preset second threshold, and identify the test links that have memory leaks with a memory utilization rate higher than a preset third threshold. For the test link that is CPU-exclusive or has memory leaks, the scheduling priority and resource allocation strategy of the test link are adjusted in the QOS yang files of the first and second test devices, respectively.

4. The method according to claim 1, characterized in that, The method further includes: In the configuration files of the first device under test and the second device under test, the link UUID number is increased to a preset link limit, and the link UUID number corresponds one-to-one with the plurality of links under test; The SAR traffic monitoring tool of the first device under test is used to monitor the link quality data of each of the multiple links under test. The link quality data includes: latency, jitter and packet loss rate. Based on the link quality data of each of the multiple links under test, determine the interface status between the first device under test and the second device under test; When the interface status indicates that the interface is blocked, the scheduling priority and resource allocation strategy of the multiple test links are adjusted accordingly in the QoS yang files of the first and second test devices.

5. The method according to claim 1, characterized in that, The clustering algorithm model is pre-trained according to the following steps: Collect multiple historical link quality data of multiple links under test at different time periods, and input the multiple historical link quality data as training samples into the clustering algorithm model; From the multiple historical link quality data, N historical link quality data are randomly selected as the first cluster centroid of each of the N clusters, where N is an integer greater than 1; Calculate the distance from each historical link quality data point to the cluster centroid, and assign the multiple historical link quality data points to the nearest cluster; Based on the average value of historical link quality data within the same cluster, the cluster centroid is iteratively updated until the difference between the changes of the cluster centroids of two adjacent updates is less than a preset change threshold, thus obtaining N convergent clusters. Based on the centroids of the N convergent clusters, the convergence regions of the N clusters are determined.

6. The method according to claim 2, characterized in that, After clustering the link quality data of the multiple links to be tested based on the pre-trained clustering algorithm model, the method further includes: Determine the centroid of each of the multiple clusters, and the link quality index corresponding to each centroid, wherein the link quality index is a one-dimensional index; The link quality data of each of the multiple links under test is mapped to the cluster centroid of the cluster to obtain the link quality index of each of the multiple links under test. The link quality data includes at least: latency, jitter and packet loss rate. The method further includes: Monitor whether any of the multiple links under test are faulty; In the presence of a faulty link, determine the cluster centroid of the cluster to which the faulty link belongs; From the cluster to which the faulty link belongs, determine the backup link that is closest to the link quality index of the faulty link and whose current remaining bandwidth meets the preset bandwidth requirement. Switch the faulty link to the backup link.

7. The method according to claim 1, characterized in that, After clustering the link quality data of the multiple links to be tested based on the pre-trained clustering algorithm model, the method further includes: Collect historical link quality data from multiple links under test at different time periods to form a historical dataset; Based on the historical dataset, determine the link quality evolution trend of each of the multiple links to be tested; Based on the clustering results and link quality evolution trends of the link quality data, potential faulty links whose link quality data exceeds the set tolerance range for multiple consecutive time periods are identified, and result-oriented analysis is performed on the potential faulty links.

8. A multi-link testing system based on SD-WAN, characterized in that, The system is used to perform the steps of the method as described in any one of claims 1-7, the system comprising: a first device under test (DUT), a second DUT, and a PC, wherein the first DUT and the second DUT are respectively configured with multiple DUT links; the first DUT is connected to the PC via a LAN port, and the first DUT and the second DUT are connected to each other via a WAN port; The first device under test is configured to send a probe message to the second device under test, and receive a feedback message returned by the second device under test in response to the probe message, and report the link quality data of each of the plurality of links under test to the PC device according to the probe message and the feedback message; The second device under test is used to receive the probe message and send the feedback message to the first device under test; The PC device performs clustering processing on the link quality data of each of the multiple links under test based on a pre-trained clustering algorithm model. The result of the clustering processing represents the link quality characteristics of each of the multiple links under test.

9. The system according to claim 8, characterized in that, The first device under test (DUT) accesses and obtains its IP address through the PC device; the WAN port of the second DUT is configured with an IP address that is on the same network segment as the first DUT but different from its IP address. The first device under test has a first routing detail, and the first routing detail accesses the IP address of the second device under test through the WAN port of the second device under test. The second device under test has a second routing detail added, and the second routing detail accesses the IP address of the first device under test through the WAN port of the second device under test; The system also includes a WAN port verification module, used to verify whether the WAN port of the first device under test can ping the IP address of the WAN port of the second device under test, and to verify whether the WAN port of the second device under test can ping the IP address of the WAN port of the first device under test.

10. The system according to claim 8, characterized in that, The first device under test is configured with a first configuration file, and the second device under test is configured with a second configuration file; The first configuration file is used to create the same number of link UUIDs for the multiple links under test, and to set the end probe IP that is the same as the IP address of the second device under test, wherein the link UUIDs correspond one-to-one with the multiple links under test; The second configuration file is used to create a link UUID number that corresponds one-to-one with the link UUID number configured in the first configuration file, and to set an end probe IP that is the same as the IP address of the first device under test.

Citation Information

Patent Citations

  • BFD link detection and evaluation system and method based on SD-WAN scene

    CN108173718A

  • Link quality detection method and device

    CN114615178A