Active and passive cooperative network telemetering method and system based on sparsity prediction and path planning

By using a collaborative active-passive network telemetry system based on sparsity prediction and path planning, the problem of telemetry blind spots in large-scale data center networks has been solved, achieving network monitoring with high visibility and low resource overhead across the entire network, thus improving service quality and anomaly detection efficiency.

CN121967304APending Publication Date: 2026-05-01SOUTHEAST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SOUTHEAST UNIV
Filing Date
2026-03-05
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing network telemetry methods have telemetry blind spots in large-scale data center networks, and cannot monitor low-load or idle links in real time, resulting in monitoring delays and resource waste. Furthermore, they lack the ability to predict sparse links in a forward-looking manner, making it impossible to detect potential risks in a timely manner.

Method used

The active-passive collaborative network telemetry system (HybINT) employs sparse prediction and path planning. It reduces redundancy in the data plane through adaptive probability sampling and label cooling mechanism, and performs prediction in the control plane by combining log-normal distribution transformation and Holt double exponential smoothing algorithm. It constructs a sparse link subgraph and plans the optimal detection path to achieve forward-looking sparse link coverage and reconstruction.

Benefits of technology

It achieves 99% visibility across the entire network, significantly reduces resource consumption, improves network anomaly detection and service quality, and enables low-latency, high-visibility and efficient sparse link monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121967304A_ABST
    Figure CN121967304A_ABST
Patent Text Reader

Abstract

The invention discloses an active and passive cooperative network telemetering method based on sparsity prediction and path planning, and the method comprises the steps: designing an adaptive probability sampling function combining the queue depth and the number of INT labels in a passive sensing module of a data plane; a mark cooling mechanism aiming at a port is combined to reduce extra expenditure brought by a telemetering label, but monitoring of high-value flow state information is not lost; on a control plane, the invention provides an algorithm for converting the packet forwarding rate in an actual environment into a logarithmic form by using logarithmic normal distribution transformation and Holt double-exponential smoothing prediction, so as to carry out smoothing prediction on a future flow condition, thereby finding a sparse risk link of which the packet forwarding rate is smaller than a certain threshold value in advance; the discovery of the detection blind area is changed from passive to active; then, according to the predicted sparse link set and a prospective sparse link coverage and reconstruction (PSCR) mechanism, the topology of the whole network is partitioned, a series of loops are generated, an optimal loop is selected to serve as an active detection path, and a detection task is executed by the nearest host; therefore, the consumed bandwidth is minimized, and potential blind areas can be detected as much as possible; the method is suitable for a large-scale data center network, the blind area of network detection can be eliminated under extremely low overhead, and the visibility of the whole network reaches 99%.
Need to check novelty before this filing date? Find Prior Art

Description

A method and system for active-passive cooperative network telemetry based on sparse prediction and path planning Technical Field

[0001] This invention belongs to the technical field of large-scale data center network monitoring and measurement. More specifically, it is a method and system for active-passive collaborative network telemetry based on logarithmic domain smoothing prediction and graph theory path planning, applicable to ultra-large-scale data center network (DCN) environments, achieving fine-grained, high-visibility network-wide telemetry. This invention is primarily used to ensure Quality of Service (QoS) in scenarios such as cloud computing and big data processing. It overcomes the monitoring failure problem of traditional in-band network telemetry (INT) methods in areas with low traffic. Background Technology

[0002] In recent years, with the development of global cloud computing, hyperscale data centers, and generative artificial intelligence applications, the scale and complexity of data center networks have increased at an unprecedented rate, and network traffic has become increasingly bursty and dynamic. To meet stringent SLAs and ensure a good user experience, it is necessary to accurately perceive the status information of every link and every queue in the entire network at every point in time. Therefore, in-band network telemetry (INT), as a new technology based on a programmable data plane, detects the status of traffic by embedding telemetry metadata at the locations where service data packets pass through. Because INT offers higher accuracy and lower latency compared to traditional SNMP and NetFlow, it has become the mainstream choice for data center monitoring today.

[0003] However, for large-scale deployments, current in-band network telemetry technology has fundamental shortcomings when facing severe traffic imbalances. Because its implementation method dictates that it is only a passive telemetry method, and its measurement is based on the distribution of traffic flows, it suffers from a fatal flaw—visibility is limited to where the traffic flows can reach. In real-world network environments, some links are always under low load or even no load at all. These links, known as "silent links," cannot be monitored in real time due to a lack of sufficient telemetry samples, resulting in serious monitoring blind spots. If these silent links experience performance degradation or hardware connection problems, the monitoring system will be unable to detect these problems in a timely manner because it lacks the corresponding traffic flow information, thus prolonging the time for problem localization and resolution, and severely impacting service quality.

[0004] To alleviate the limitations of passive telemetry due to its limited coverage, current research is exploring active telemetry methods. Under certain conditions, probe data packets are constructed and monitored via segment routing (SR) to target any path requiring observation. While active telemetry offers high flexibility, it also has its drawbacks and has encountered numerous problems during application, primarily resulting in significant resource waste and placing a heavy burden on equipment. Current solutions present a dilemma: one approach is periodic network-wide probing, such as Ping Mesh, which sends a large number of probe packets to the entire network at regular intervals for comprehensive coverage. However, this wastes valuable bandwidth and puts immense pressure on switch hardware and control plane servers. The other approach is simply setting a threshold to trigger a response, probing only when a link's sampling rate decreases. This method is inherently delayed, as there is a time interval between detecting network degradation and initiating probing. Probing typically begins only after the link has ceased operation, making continuous monitoring data unreliable and thus failing to provide accurate QoS analysis results.

[0005] In addition, existing methods lack the forward-looking predictive ability for randomly occurring sparse links in large-scale networks, making it difficult to accurately and promptly compensate for blind links. While existing telemetry prediction technologies primarily employ machine learning methods, these are mainly used for high-load conditions, with little research on sparse links. Furthermore, while machine learning methods are effective, they require significant computational resources, resulting in substantial overhead. Data center traffic is driven by a variety of applications, exhibiting a pronounced heavy-tailed distribution. Therefore, traditional linear prediction models suffer from significant prediction errors due to drastic fluctuations in traffic along links, failing to detect potential risks in a timely manner. Moreover, existing active probing planning is primarily based on a single end-to-end path, neglecting risks present in sub-topologies throughout the network. When dealing with scattered sparse links, this leads to numerous redundant probes, wasting valuable bandwidth resources and increasing the burden on north-south communication.

[0006] In summary, to achieve ultra-high visibility across the entire network, existing network measurement schemes face the following challenges: First, they lack a low-overhead prediction mechanism that can transform telemetry blind spot detection from passive feedback to proactive prediction; second, they lack a low-overhead sampling mechanism in the data plane that can perceive multi-dimensional congestion states and adaptively adjust; and third, they lack graph theory modeling methods capable of optimizing path coverage in predicted sparse risk areas. These problems prevent current methods from selecting suitable nodes for path planning based on the latest network conditions, making it difficult to concentrate probe resources on high-risk links and thus reducing the visibility of active telemetry systems across the entire network. Therefore, developing a collaborative telemetry method that combines the advantages of active and passive methods, possesses proactive risk perception capabilities, and achieves full coverage with minimal resources is a key technical challenge that needs to be overcome to improve the security and performance of large-scale networks. Summary of the Invention

[0007] In large-scale data center networks (DCNs), traditional in-band network telemetry (INT) is inherently a passive measurement method, limiting its visibility to the location of the service flow. This leads to numerous "silent links" under low load or idle conditions, making effective real-time monitoring impossible and significantly complicating monitoring. To overcome these issues, this invention proposes a proactive-passive collaborative network telemetry system, HybINT, based on sparsity prediction and path planning. The core idea is to address the blind zone problem in network telemetry by separating the control plane and data plane. This transforms the blind zone detection process from passive feedback to proactive prediction, while using graph theory modeling to calculate the optimal detection path, thereby maximizing the visibility of the entire network link without excessive resource waste.

[0008] This invention designs an adaptive probability sampling function based on queue depth and the number of INT tags in the passive sensing module of the data plane, and incorporates a tag cooling mechanism for ports. This reduces telemetry tag waste while accurately acquiring status information of important traffic. In the prediction part of the control plane, this invention introduces a preprocessing mechanism of log-normal distribution transformation and Holt's double exponential smoothing algorithm to transform the packet forwarding rate of the physical world into a logarithmic form. Simultaneously, it smooths and models the traffic trend of future periods, thereby detecting sparse risk links with packet forwarding rates below a safety threshold in advance. This transforms the discovery of monitoring blind spots from traditional passive feedback to proactive prediction. Based on the sparse link set generated by the prediction, this invention proposes a prospective sparse link coverage and reconstruction (PSCR) mechanism. This mechanism finds suitable subgraphs in the network topology and constructs corresponding loops. Then, it uses dynamic programming to find an optimal proactive detection path and assigns it to the host physically closest to this path for detection. In this way, the blind spot problem in the entire network is solved at a very low cost. This solution is suitable for large-scale data center networks, eliminating network monitoring blind spots with extremely low resource overhead, achieving a network link visibility rate of 99%, which is beneficial for subsequent network anomaly detection, performance improvement, and service quality assurance.

[0009] To achieve the objectives of this invention, the specific technical steps of this solution are as follows: A telemetry method for active-passive cooperative networks based on sparse prediction and path planning, the method comprising the following steps:

[0010] Step (1) On the telemetry system, the passive sensing module in the data plane uses a distributed switch to adaptively sample the service packets passing through the node. Improved in-band network telemetry (INT) is used to obtain real-time metadata such as link latency, port bandwidth utilization, queue length and current packet forwarding per second. In order to reduce the redundancy of telemetry data, adaptive probability sampling and label cooling methods are used to pre-suppress telemetry traffic.

[0011] Step (2) is to obtain the original packet forwarding rate vector of each monitoring link in the prediction and planning module on the control plane. During the data preprocessing process, the heavy-tailed distribution characteristics of the data center network traffic in the physical domain are taken into account. The data is transformed in the logarithmic domain during the preprocessing process, so that the multiplicative changes in the physical domain are transformed into the additive changes in the logarithmic domain, which facilitates the subsequent data modeling.

[0012] Step (3) uses the Holt double exponential smoothing algorithm to obtain the link traffic change trend for the logarithmic domain observation sequence obtained in step (2), and obtains the horizontal component of packet forwarding rate and the slope component of packet forwarding rate change trend, thereby obtaining a forwarding rate trend prediction model for future sampling times.

[0013] Step (4) Construct a sparsity determination mechanism based on risk perception. After converting the sampling pass rate threshold specified by the system into logarithmic form, compare it with the logarithmic prediction value generated in step (3) to find the set of sparse risk links whose packet forwarding rate will be less than the threshold in the future cycle.

[0014] Step (5) For the sparse risk link set found in step (4), the entire network topology is modeled using the prospective sparse link coverage and reconstruction (PSCR) method on the control plane to obtain a sparse link subgraph and divide the sparse link subgraph into several disjoint subgraph parts. Then, the optimal probe coverage path is calculated for each subgraph.

[0015] Step (6) Based on the path planning scheme obtained in step (5), select the host with the closest physical distance as the starting source node for the probe, encapsulate the active probe message on the basis of the source route and send it into the network to accurately compensate for potential monitoring blind spots and ensure that the links in the entire network have continuous visibility.

[0016] Step (7) Establish a dynamic reconstruction mechanism. The system periodically compares the structure of the latest predicted sparse risk subgraph with the corresponding subgraph of the previous period. Based on the comparison results, it changes the corresponding active detection strategy and compensation path, deletes invalid rules and issues new rules, thereby achieving closed-loop automated management with high visibility of the entire network.

[0017] Furthermore, in step (1), the steps for collecting and preprocessing the basic data are as follows:

[0018] (1.1) The control plane sends intelligent sampling strategies to the data plane switch through the north-south communication interface. The switch abandons the traditional fixed probability sampling mode and instead adopts an adaptive probability function that integrates multi-dimensional network states. ;

[0019] (1.2) When a data plane switch receives a service packet, it first reads the number of INT tags already carried in the packet header. Calculate the baseline sampling probability :

[0020]

[0021] in, The maximum path hop count is preset; to reduce hardware processing overhead, the division operation is approximated in the data plane by a right shift operation. This method inserts INT telemetry data into data packets already carrying INT tags as much as possible, reducing the number of reporting operations and lowering communication overhead.

[0022] (1.3) The switch monitors the depth of its egress queue in real time. And determine the load factor according to the preset water level. The value is 0 when the queue depth is less than 50% of the maximum capacity; 1 when the queue depth is between 50% and 75%; and 2 when the queue depth exceeds 75%. (Final sampling probability) According to the formula Calculations show that by using a proactive reduction method that dynamically increases the number of shift bits, telemetry traffic can be prevented from exacerbating network congestion under high load scenarios.

[0023] (1.4) Implement the forced marking rule at the end of the cycle: When the switch detects that the current sampling cycle is nearing its end, it determines whether a valid telemetry data packet has been obtained in the cycle. If not, it skips the probability calculation step and directly performs a forced marking operation on the next data packet, so that each detection cycle has corresponding underlying metadata sent to the controller to ensure data consistency.

[0024] (1.5) The switch transmits data packets containing node status information (such as latency, bandwidth utilization, throughput, etc.) to the telemetry tail node. At the tail node, hop-by-hop data is packaged and sent to the control plane for noise reduction and standardization. The processed data is stored in time series form for subsequent prediction. This adaptive probability sampling mechanism concentrates telemetry data into a single message for transmission as much as possible to reduce the number of reporting times, greatly reducing the communication overhead from the data plane to the control plane. At the same time, combined with the suppression of sampling frequency by the egress queue depth and the method of forced marking at the end of the period, it ensures the absolute continuity of the underlying monitoring data under extremely sparse traffic conditions while effectively preventing further deterioration of network congestion.

[0025] Furthermore, in step (1), the marking cooling mechanism specifically includes the following sub-steps:

[0026] (1.1) The switch port maintains a last-marked timestamp for each detection packet. ;

[0027] (1.2) When a new service packet arrives at the port, the system obtains the current system time. Calculate the time difference and compare it with the cooling threshold. Compare the data to determine the port cooling status. :

[0028]

[0029] in, The port cooling time threshold is determined by the baseline cooling time. Combined with adjustable coefficient The following was calculated using displacement operations: ;

[0030] The coefficients shown in (1.3) The risk level of the link is dynamically adjusted by the controller: if the controller detects a sparse, risky link, the risk level is increased. This makes the cooling threshold This reduces redundancy, but indirectly increases the number of passive samplings.

[0031] (1.4) If the system considers the port to be cooling (i.e.) If the condition is true, then the current packet should not be marked and should be sent directly; otherwise, the adaptive sampling probability determination step described above should be executed, updating the port's status upon successful marking. This port tagging cooling mechanism utilizes port cooling to prevent continuous re-tags of burst traffic, thereby saving significant bandwidth resources. It also allows the control plane to adjust the port cooling time based on displacement, increasing the sampling rate in sparse link scenarios to ensure a good trade-off between accuracy and resource consumption.

[0032] Furthermore, in step (2), the process of log-normal distribution modeling and data preprocessing is as follows:

[0033] Considering that data center traffic exhibits non-stationary, nonlinear, and heavy-tailed distribution characteristics due to bursts from the application layer within the physical domain, this invention collects the raw packet forwarding rate on the link. (Unit: PPS), and a minimum value is added to prevent logarithmic singularity issues that may occur under low or no traffic conditions. Then, the logarithmically transformed observation sequence is obtained according to the logarithmic transformation formula. :

[0034]

[0035] This preprocessing step transforms the multiplicative changes in the physical domain into additive changes in the logarithmic domain. The purpose is to make the sequence symmetric and homoscedastic, reduce the prediction error caused by the drastic fluctuations in the original PPS values, and provide a good data foundation for the subsequent smoothing algorithm.

[0036] Furthermore, the implementation details of sparsity trend prediction and risk identification in steps (3) and (4) are as follows:

[0037] (3.1) The system uses the Holt double exponential smoothing algorithm to obtain the dynamic momentum characteristics of the flow rate, and uses a smoothing factor. To update the level status , used to represent the basic forwarding rate level after noise reduction; using a smoothing factor To update trend status This is used to obtain the slope of the dynamic evolution of the forwarding rate change; its state update expression is:

[0038]

[0039]

[0040] (3.2) The system uses logarithmic extrapolation to determine the predicted value at the next time t+1. To align with the heterogeneous nature of data center network structures, different weights are assigned to switches at different layers during prediction: for core layer switches, a smaller factor is used to eliminate the fluctuations caused by aggregated flows; while for edge layer switches, a larger factor is used to better detect significant drops in packet forwarding rate due to sudden events. This prediction process proactively detects sparse risk links to prevent detection blind spots, and uses logarithmic transformation and Holt smoothing algorithms to replace high-computing-power models, greatly reducing server load. Based on this, parameters are set hierarchically to eliminate traffic fluctuations and increase sensitivity to drops, thereby obtaining accurate and interference-resistant lightweight predictions.

[0041] (4.1) Define the sampling pass rate Define a security threshold to determine the minimum unit traffic density (e.g., 100 pps) required to reconstruct the network state. It is twice the qualified sampling rate, that is If the link traffic is within the safe threshold, i.e. If so, the link is determined to be a "sparse risk link" with weak ability to resist sudden risks;

[0042] (4.2) The controller sets the safety threshold during the initialization phase. Offline mapping to logarithmic threshold Real-time comparison of predicted values ​​during operation and The size of the predicted value; if the predicted value is lower than this threshold, the link is marked as a potential target for coverage and merged into the sparse link set of the next cycle. In this way, proactive risk warnings can be issued before monitoring continuity is interrupted.

[0043] Furthermore, in step (5), the path planning logic of the forward-looking sparse link coverage and reconstruction (PSCR) mechanism is as follows:

[0044] (5.1) Based on the sparse risk link set predicted from the entire network Construct a real-time sparse link subgraph in the control plane. ;

[0045] (5.2) Based on the connected component analysis algorithm, the subgraph G′ is divided into n mutually disjoint independent subgraph components. For each individual component, the PSCR mechanism uses a loop generation method to obtain the corresponding Eulerian circuit or optimal path cover set. The purpose of this design logic is to ensure that as many predicted risk links as possible are detected in the subgraph, while minimizing the sending of too many probe messages to non-sparse parts.

[0046] (5.3) The algorithm automatically calculates the distance between all paths and the host node, and then selects the host node with the closest physical distance as the probe source node.

[0047] (5.4) Afterwards, the controller generates an active probe command with a source routing label stack and sends the command through the north-south interface to inject probe traffic into the selected probe source host node, thereby achieving high visibility of the entire network link. This path planning process utilizes connected component partitioning and the use of Eulerian circuits within local subgraphs to find probe paths, thus avoiding the high overhead of large-scale network searches and covering all sparse risk links with minimal link duplication, achieving low redundancy and full coverage of monitoring blind spots.

[0048] Furthermore, the implementation details of the dynamic reconfiguration mechanism in step (7) are as follows:

[0049] The controller maintains a set of currently active probe paths. After each telemetry cycle ends, the system compares the currently calculated risk subgraph with the risk subgraph from the previous cycle. If a difference is found, dynamic reconfiguration is triggered: the controller retracts the compensation path rules set on links that have recovered to normal levels in the previous cycle (i.e., traffic expectations have reached a safe range), and uses the latest... Based on the set, the path planning and injection process from step (5) to step (6) is carried out to form a closed loop mechanism, so that the probe resources can adjust themselves in real time and automatically as the business grows.

[0050] Compared with the prior art, the technical solution of the present invention has the following beneficial technical effects:

[0051] (1) This invention provides a forward-looking telemetry risk perception mechanism. By defining and identifying sparse risk links, it solves the problem of "data only when traffic is seen" in the previous passive telemetry system. The algorithm can start detection and compensation work before the link sampling capability reaches the monitoring blind zone, thereby ensuring the continuity of performance data in time and space and achieving ultra-high real-time visibility in huge dynamic network topologies.

[0052] (2) This invention significantly reduces the resource overhead and bandwidth usage of the telemetry system. In the passive sensing mechanism on the data plane, adaptive probability sampling and label cooling are used to reduce a large amount of unnecessary metadata, alleviating the pressure on hardware resources and the storage burden on the control plane server. In the active detection part on the control plane, the PSCR algorithm is used to effectively model the sparse risk subgraph, avoiding the bandwidth waste caused by the traditional method of blindly and comprehensively detecting the entire network, so that each detection path obtains the maximum monitoring benefit, thereby maximizing the detection efficiency.

[0053] (3) This invention uses log-normal distribution transformation and Holt double exponential smoothing algorithm with very low time complexity to solve the problem of huge computational resource consumption caused by complex prediction models. It completes the prediction task with very low computational overhead and has high accuracy. This is a lightweight method that can solve the problems of nonlinear, non-stationary and heavy-tailed data center traffic. It can accurately reflect the horizontal component and trend of packet forwarding rate, and at the same time greatly reduce the probability of missed reports in monitoring blind spots, thereby improving the predictability of sparse risk links.

[0054] (4) This invention realizes fully automated closed-loop telemetry operation and maintenance as well as dynamic reconfiguration. Based on the centralized planning capability of SDN and the perception feedback mechanism of the data plane, the system can automatically adjust the active detection range according to the real-time traffic situation. It can automatically adjust according to the periodic changes of service load without human intervention, with very low operation and maintenance overhead, while ensuring the effectiveness of high-performance detection strategies in the entire network. Attached Figure Description

[0055] Figure 1 is a schematic diagram of the overall architecture of a remote telemetry system based on sparse prediction and path planning, which is a collaborative active-passive network.

[0056] Figure 2 is a schematic diagram of the overall process of the active-passive collaborative network telemetry method based on sparse prediction and path planning. Detailed Implementation

[0057] The technical solutions provided by the present invention will be described in detail below with reference to specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and are not intended to limit the scope of the present invention.

[0058] Example: The present invention proposes an active-passive collaborative network telemetry system (HybINT) based on sparse prediction and path planning, as shown in Figure 1. The system adopts a software-defined network (SDN) architecture with separate control and data planes. The entire system operates in a closed-loop manner, specifically including a passive sensing module to acquire data, a prediction and planning module to establish a sparse risk link prediction model, and dynamic reconstruction of active detection paths.

[0059] This invention provides a telemetry method for active-passive cooperative networks based on sparse prediction and path planning. The overall process is shown in Figure 2, and includes the following steps:

[0060] Step (1) For the passive sensing module in the telemetry system, adaptive probability sampling is performed on the service traffic flowing through the network node (switch) to collect network status metadata such as link delay, port bandwidth utilization, queue length and current packet forwarding rate.

[0061] In one embodiment of the present invention, the data acquisition process of the passive sensing module is as follows:

[0062] (1.1) The control plane sends intelligent sampling strategies to the telemetry source nodes (switches) in the data plane through the north-south communication interface. The strategies define the reference sampling frequency and the maximum path hop count. and marking cooling parameters ;

[0063] (1.2) The data plane switch processes the service traffic. The system abandons the traditional fixed probability sampling mode and instead adopts an adaptive probability function that integrates multi-dimensional network states. ;

[0064] (1.3) The calculation of the sampling probability depends on the number of in-band network telemetry (INT) tags already carried in the message. and the real-time depth of the switch's current egress queue. ;

[0065] (1.4) Baseline Probability From the formula This mechanism ensures that telemetry data is aggregated into messages containing the most telemetry data possible, reducing the number of reporting operations and thus reducing communication overhead between the data plane and the control plane. At the hardware level, to reduce computational overhead, the division operation is approximated through a right shift operation.

[0066] (1.5) The system introduces a load factor. The sampling probability is dynamically suppressed, and the value of the load factor is determined according to the queue level classification: when the queue depth is less than 50% of the maximum capacity, the load factor is 0; when it is between 50% and 75%, the load factor is 1; when it exceeds 75%, the load factor is set to 2.

[0067] (1.6) The final sampling probability is calculated according to the formula. The calculation, through a proactive reduction method that dynamically increases the number of shift bits, ensures that the telemetry strength is automatically reduced in high network load scenarios, thus preventing telemetry packets from becoming a factor that induces network congestion;

[0068] (1.7) To ensure the temporal continuity of monitoring data, the data plane also adds a forced marking mechanism at the end of the cycle. When the sampling cycle is about to end and no valid telemetry sample has been captured in the cycle, the system will skip the probability calculation logic and force the metadata marking operation to be performed on the next arriving business data packet.

[0069] In another embodiment of the present invention, in order to optimize the redundancy of INT tags and reduce processing overhead, a tag cooling mechanism is also implemented in step (1), the specific logic of which is as follows:

[0070] (1.1) Each port of the switch maintains and records the last timestamp of the specific detection packet in real time. ;

[0071] (1.2) When a new service data packet arrives at the port, the system first extracts the number of tags it already carries. and get the current system time. ;

[0072] (1.3) Calculate the current cooling status of the port based on the preset cooling parameters. The judgment formula is as follows:

[0073]

[0074] (1.4) Among them, the port cooling time threshold From the base time With the adjustment coefficient dynamically issued by the controller Determined, that is ;

[0075] (1.5) The adjustment coefficient The value of is typically an integer between 1 and 3; when the control plane identifies a link as being in a sparse risk state, the controller will increase the coefficient. The value of can significantly shorten the cooling cycle, thereby indirectly increasing the sampling frequency;

[0076] (1.6) If the determination result shows that the port is in a cooling state (i.e.) If the condition is true, the system prohibits inserting the INT tag into the data packet, skips subsequent logic, and forwards it directly; if the port is not in a cooldown state, the above adaptive probability judgment process is entered, and the port's condition is updated synchronously after successful tagging. for .

[0077] In step (2), within the prediction and planning module, the control plane server aggregates the telemetry data reported by the passive sensing module and extracts the original packet forwarding rate vectors for each link. Furthermore, a log-normal distribution transformation mechanism was introduced to perform statistical preprocessing on the data;

[0078] In one embodiment of the present invention, the process of transforming data in the number field is described in detail below:

[0079] (2.1) In a real production environment, the packet forwarding rate of the data center network link (Unit: PPS) It has a very large dynamic fluctuation range and exhibits a significant heavy-tailed distribution characteristic driven by application layer burst traffic.

[0080] (2.2) To address the problem of excessively large residuals in traditional prediction models caused by non-stationary and nonlinear distribution characteristics, this invention maps the multiplicative evolution law of the physical domain to the additive evolution law of the logarithmic domain in the preprocessing stage, using the mapping formula... Construct observation sequences;

[0081] (2.3) In the formula For the introduction of a very small amount (e.g.) The aim is to avoid the logarithmic singularity problem in the state of complete link idleness (i.e. zero PPS), thereby endowing the observation sequence with symmetry and homoscedasticity, and providing a stable statistical basis for subsequent numerical modeling.

[0082] Step (3) involves processing the logarithmic field sequence transformed in step (2). The Holt double exponential smoothing algorithm, which is fast and has very low hardware requirements, is used to capture the dynamic momentum of link traffic and achieves forward prediction of the forwarding rate trend in the future by decoupling state variables.

[0083] In one embodiment of the present invention, the specific logic for sparse link trend prediction is as follows:

[0084] (3.1) The system establishes a double exponential smoothing model and utilizes a smoothing factor. The level component, which represents the denoised baseline forwarding rate level, is updated:

[0085]

[0086] (3.2) The system utilizes a smoothing factor The trend component is updated; this trend component is used to capture the dynamic momentum of changes in the forwarding rate, i.e., the evolution slope.

[0087]

[0088] (3.3) By combining the above components, the system calculates the predicted logarithmic domain packet forwarding rate for the next cycle:

[0089]

[0090] (3.4) To adapt to the heterogeneity of data center network topology, the prediction engine dynamically adjusts the weight of the smoothing factor according to the layer to which the link belongs: for core layer links, the system allocates a smaller weight. and The coefficient is used to effectively filter out high-frequency random jitter generated by aggregated traffic; while for the edge layer link, a higher sensitivity coefficient is used to enhance the system's ability to capture the risk of sudden drops in business flow and ensure a high recall rate for potential blind spots.

[0091] Step (4) Establish a sparsity determination model based on risk perception. By comparing the predicted values ​​generated in step (3) with the preset safety threshold, identify the set of sparse risk links that are expected to fall to the unqualified state in the next cycle. ;

[0092] In one embodiment of the present invention, the specific criteria for risk determination are as follows:

[0093] (4.1) The system first defines the sampling pass rate. This represents the minimum packet forwarding rate threshold (e.g., 100 pps) required to ensure the telemetry control plane server accurately reconstructs the link state. If so, the link is called an unqualified link.

[0094] (4.2) To achieve proactive intervention, this invention innovatively proposes the concept of "Dangerous Link": if the current sampling rate of a link meets the qualification standard, but its predicted forwarding rate is lower than the security threshold. If so, it is identified as a risky link;

[0095] (4.3) The safety threshold is set to twice the sampling pass rate, i.e. ;

[0096] (4.4) During system operation, the controller will set a safety threshold. Offline mapping to logarithmic threshold And directly compare in the logarithmic space and The numerical value;

[0097] (4.5) If If so, the link is merged into the predicted sparse link set for this period. This allows the compensation process to be triggered in advance, before the link traffic actually drops to an unmonitorable blind zone.

[0098] Step (5) is based on the sparse risk link set identified in step (4). We utilize the forward-looking sparse link coverage and reconstruction (PSCR) mechanism for graph theory modeling and active path exploration construction.

[0099] In one embodiment of the present invention, the execution steps of the PSCR path planning algorithm are as follows:

[0100] (5.1) The system extracts the predicted risk link set and its corresponding network topology information, and constructs a sparse link subgraph. ;

[0101] (5.2) Use connectivity analysis algorithms to divide the subgraph Decomposed into Independent subgraph components that are not connected to each other ;

[0102] (5.3) For each individual component The system executes the circuit construction logic as follows: it prioritizes searching for and constructing a set of Eulerian circuits that can cover all sparse links within the component. ;

[0103] (5.4) The loop construction logic is designed to ensure that each subgraph component is efficiently covered, while minimizing the redundant transmission of probe messages in non-risk areas (i.e., healthy areas with sufficient traffic sampling).

[0104] (5.5) System calculates the set of coverage paths The starting and ending points of each path are identified, and the edge host with the closest physical topological distance is selected as the probe emission source node. .

[0105] Step (6) Based on the path planning instructions generated in step (5), the designated source node injects active probe messages into the network through source routing technology to achieve accurate coverage compensation for potential monitoring blind spots;

[0106] In one embodiment of the present invention, the injection and feedback process for active detection is as follows:

[0107] (6.1) The controller sends an activation command containing the probe path label stack to the selected source node host. ;

[0108] (6.2) The source node constructs a probe message carrying a specific HybINT identifier. After the switches along the way identify the message identifier, they force the INT metadata to be marked. Finally, the probe reaches the destination node along the planned path and feeds back the path status to the control plane server, achieving ultra-high visibility of the entire network link.

[0109] Step (7) Establish a periodic dynamic reconstruction mechanism. The system monitors the evolution of the sparse risk subgraph in real time and updates the active detection compensation strategy dynamically based on the predicted offset.

[0110] In one embodiment of the present invention, the specific maintenance logic for dynamic reconstruction is as follows:

[0111] (7.1) As data center business traffic migrates dynamically over time, the system continuously tracks the prediction results for each period;

[0112] (7.2) The controller compares the sparse risk link subgraph generated in the current period with the subgraph distribution of the previous historical period. If it finds that the two structures are inconsistent (i.e., the original risk link has been restored or the new link has become risky), the reconstruction process is automatically triggered.

[0113] (7.3) The system will delete the active detection compensation path rules that are no longer applicable in the previous round, and will adjust them according to the newly generated rules. Re-execute the path planning logic in step (5), and distribute the newly generated compensation rules to achieve adaptive adjustment of measurement resources according to the network situation.

[0114] The technical means disclosed in this invention are not limited to those disclosed in the above embodiments, but also include technical solutions composed of any combination of the above technical features. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of this invention, and these improvements and modifications are also considered within the scope of protection of this invention.

Claims

1. A telemetry method for active-passive cooperative networks based on sparse prediction and path planning, characterized in that, The method includes the following steps: Step (1) In a distributed network environment, the passive sensing module of the data plane realizes real-time monitoring of the service flow of each switching node and uses an adaptive sampling mechanism to obtain the metadata of the network data plane; the metadata includes switch ID, ingress / egress port utilization, queue depth, high-precision timestamp and packet forwarding rate of the link, etc.; in order to ensure the accuracy of subsequent prediction, the passive sensing module achieves pre-suppression by marking and cooling the telemetry traffic and using adaptive probability sampling. The collected data is denoised and preprocessed and saved to the control plane server in the form of a multi-dimensional time matrix for subsequent traffic feature extraction and prediction. Step (2) In the prediction and planning module of the control plane, the original packet forwarding rate vector of each monitoring link is obtained. Considering that the data center network traffic has a heavy-tailed distribution, the original values ​​are processed by using logarithmic transformation to make the multiplicative relationship of packet forwarding rate into the additive relationship in the logarithmic domain, so as to obtain homoscedastic data for smooth prediction. Step (3) According to the logarithmic domain sequence after step (2), the Holt double exponential smoothing algorithm is used to extract the trend of link traffic change, separate the horizontal state quantity and trend state quantity of traffic, so as to obtain the extrapolation formula of packet forwarding rate for the next period, and obtain the predicted value at the next sampling time. Step (4) Establish sparse The risk assessment model converts the pre-set sampling pass rate threshold of the system into a logarithmic form, and then compares whether the predicted value obtained in step (3) is less than the corresponding preset security threshold. In this way, it obtains a set of "sparse risk links" whose packet forwarding rate is expected to be lower than the security threshold in the future period, so that the perception of the monitoring blind spot changes from passive response to active early warning. In step (5), for the set of sparse risk links found in step (4), the topology of the entire network is modeled using graph theory on the control plane using the prospective sparse link coverage and reconstruction (PSCR) mechanism to obtain a sparse link subgraph and divide it into several disjoint parts. Then, an optimal link is found for each part. Detection path, step (6) According to the path planning scheme obtained in step (5), select the host with the closest physical distance as the probe starting source node, encapsulate the active probe message in the form of source routing and send it to the network, thereby making up for the monitoring missing area caused by too little link service flow, and ensuring the visibility of all links in the entire network. Step (7) Establish a dynamic reconstruction mechanism, periodically compare the difference between the sparse risk subgraph obtained in this prediction and the topology of the historical sparse risk subgraph in the previous round, and then adjust the active probe rules and compensation path according to this difference, delete invalid rules and send new rules to the entire network, thereby achieving automatic management with high visibility throughout the process.

2. The active-passive cooperative network telemetry method based on sparse prediction and path planning according to claim 1, characterized in that, The process of intelligent sampling performed by the passive sensing module of the data plane in step (1) specifically includes the following sub-steps: (1.1) The control plane sends an adaptive sampling strategy to the data plane exchange node through the north-south communication interface. The strategy includes a preset maximum path hop count. Reference sampling time window and cooling parameters ; (1.2) When a data plane node receives a service message, it first reads the number of tags carried in the message's INT header. Calculate the baseline sampling probability : In order to reduce the hardware computing logic overhead, the division operation is approximated by a right shift operation; (1.3) The switch monitors the depth of its egress queue in real time. And determine the load factor according to the preset water level. When the queue depth is less than 50% of the maximum capacity, the value is 0; when the queue depth is between 50% and 75%, the value is 1; when the queue depth exceeds 75%, the value is 2; (1.4) Final sampling probability The calculation method is as follows: The adaptive probability function is achieved by increasing... This proactively reduces telemetry intensity under high load conditions, preventing telemetry traffic from inducing or exacerbating network congestion.

3. The active-passive cooperative network telemetry method based on sparse prediction and path planning according to claim 1, characterized in that, The intelligent sampling process in step (1) also includes a forced marking mechanism at the end of the cycle. The data plane maintains a message counter and a time window for each sampling cycle. When the sampling cycle is about to end (i.e. the remaining window duration is less than the set threshold) and no valid telemetry sample has been generated in the cycle, the switch will skip the probability calculation process described in step (2) and force the INT tag insertion operation to be performed on the next arriving service message to ensure that there is underlying metadata fed back to the control plane in each sampling cycle and maintain the continuity of monitoring data.

4. The active-passive cooperative network telemetry method based on sparse prediction and path planning according to claim 1, characterized in that, The marking cooling mechanism in step (1) specifically includes the following implementation steps: (1.1) The switch port maintains a last marking timestamp for each detection packet. ; (1.2) When a new data packet arrives at the port, obtain the current system time. Calculate the time difference and compare it with the cooling threshold. The formula for determining the cooling status is as follows: Among them, the cooling time threshold Based on the reference cooling time Combined with adjustable coefficient The calculation shows that: (1.3) coefficient The control plane dynamically adjusts based on the sparsity risk level of the link: if a link is identified as a sparse risk link, then by increasing... Values ​​to reduce cooling threshold This indirectly increases the passive sampling frequency; (1.4) If the system determines that the port is in a cooling state ( If true, the packet is forwarded directly and the insertion of the INT tag is prohibited; if not in a cooling-off state, the subsequent probability sampling logic is entered, and the port is updated after successful marking. 。 5. The active-passive cooperative network telemetry method based on sparse prediction and path planning according to claim 1, characterized in that, The mathematical implementation of the number domain transformation mechanism in step (2) is as follows: Extract the original packet forwarding rate of the physical domain. The unit is packets per second (PPS); a very small amount is introduced. (Its value range is within) to (between) to avoid logarithmic singularity in the state of zero flow; calculate the logarithmic domain observation sequence through mapping formula. : This transformation process converts packet forwarding rate fluctuations from an asymmetric distribution with heavy-tailed characteristics to a symmetric distribution close to a normal distribution, aiming to eliminate abnormal burst noise in the original traffic data and enhance the homoscedasticity of the prediction algorithm; the state update and prediction process of the Holt double exponential smoothing algorithm in step (3) is as follows: (3.1) Given a smoothing factor Smoothing factor used to control the response speed of the horizontal component. Used to control the sensitivity of trend components; (3.2) In At any given time, based on the current logarithmic field observations Update horizontal state variables : in Characterize the baseline forwarding rate level after denoising; (3.3) Update trend state quantity : in The slope momentum characterizing the change in forwarding rate; (3.4) Calculating the future Predicted value at time : (3.5) For heterogeneous topologies, the system assigns differentiated smoothing parameters to switches at different levels: core layer switches use smaller smoothing parameters. To filter high-frequency random jitter; edge layer switches use larger... To enhance the ability to detect sudden drops in business flow.

6. The active-passive cooperative network telemetry method based on sparse prediction and path planning according to claim 1, characterized in that, The sparsity risk assessment model in step (4) specifically includes: (4.1) defining the sampling pass rate. The minimum traffic density required to reconstruct the network state, when The link is determined to be an "unqualified link"; (4.2) Define the security threshold. It is twice the pass rate, that is If the link traffic meets the requirements If the link is identified as a "sparse risk link", it indicates that its ability to resist traffic fluctuations falling to a blind zone is extremely weak; (4.3) Perform logarithmic domain inverse determination: during the initialization phase, Offline mapping to logarithmic threshold During operation, if the predicted value satisfies If the link is not found, it will be directly marked and merged into the predicted sparse link set for the next period. middle.

7. The active-passive cooperative network telemetry method based on sparse prediction and path planning according to claim 1, characterized in that, The path planning logic of the forward-looking sparse link coverage and reconstruction (PSCR) mechanism in step (5) is as follows: (5.1) Based on the current network topology Extract from The edge set constitutes a sparse link subgraph. ; (5.2) Using the connected component analysis algorithm, the subgraph is divided into subgraphs. Decomposed into Independent subgraph components that are not connected to each other ; (5.3) For each component Execution loop construction: If If the graph is an Eulerian graph or a semi-Eulerian graph, then use the Fleury algorithm or the Hierholzer algorithm to construct an Eulerian circuit or Eulerian path that covers all sparse links within the component. ; like If the Euler condition is not met, the probe path sequence can be transformed into an optimal probe path sequence that can cover the target link set by adding auxiliary virtual edges or overlapping paths, so as to maximize the probe gain and minimize redundant probes in non-risk areas. (5.4) Calculate the shortest topological distance between each probe path and its surrounding physical hosts, and select the host with the fewest hops from the starting node of the path as the probe source node.

8. The active-passive cooperative network telemetry method based on sparse prediction and path planning according to claim 1, characterized in that, The active probe packet injection process in step (6) includes: the controller generates an active probe instruction based on the SegmentRouting (SR) tag stack; the tag stack encapsulates the complete coverage path information planned in step (5); the probe packet header carries a special HybINT identifier to distinguish it from ordinary service packets; when the probe source node host sends a probe, the switches along the way forward it according to the direction indicated by the tag stack, and forcibly add node metadata, and finally feed back the complete path status to the control plane server; the dynamic reconstruction mechanism in step (7) includes: (7.1) the controller maintains a set of currently active probe paths. (7.2) After each periodic prediction is completed, compare the newly generated data. The set of links currently being covered; (7.3) If the topology of the predicted subgraph shifts (drift rate exceeds a preset threshold) If the reconstruction instruction is triggered, the flow table rules or label paths corresponding to the invalid links (i.e., the traffic rises back to above the safe range) will be cancelled; the planning and injection process of steps (5) to (6) will be executed for the newly generated sparse risk areas; (7.4) The reconstruction mechanism realizes the dynamic allocation of measurement resources, ensuring that when the network traffic fluctuates drastically, the active detection probe always focuses on the potential monitoring blind area.

9. A system for implementing the method as described in any one of claims 1 to 8, characterized in that, The system comprises: a data plane, including multiple programmable switches, with a passive sensing submodule deployed on each switch; the data plane is responsible for implementing passive sensing functions, using the passive sensing submodule to adaptively sample and label service packets, and adding INT tags to active probe packets according to instructions from the control plane; the control plane server is the logical control and analysis center of the entire system, integrating prediction and planning modules; the control plane server receives and parses passive and active telemetry packets from the data plane to obtain a real-time state map of the entire network; the prediction and planning modules use this real-time state map of the entire network to collect information about network nodes, perform predictions based on the logarithmic domain Holt double exponential smoothing prediction algorithm, identify sparse risk links, and then combine the PSCR algorithm to generate new active probe paths and continuously update and reconstruct them.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, it implements a remote telemetry method for active and passive cooperative networks based on sparse prediction and path planning as described in any one of claims 1 to 8.