A method for screening important nodes in active telemetry scenarios
By introducing a node importance scoring model in the active telemetry scenario and combining static topology and dynamic state characteristics, the problem of unreasonable probe resource allocation is solved, load optimization and anomaly detection are improved, and adaptation to complex network environments is achieved.
Patent Information
- Application Number
- CN202411782642.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-05
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2044-12-05
AI Technical Summary
The unreasonable allocation of probe resources in existing active telemetry scenarios leads to excessive load, unreasonable path planning, and difficulty in accurately selecting important nodes. In addition, existing methods fail to comprehensively consider static topological characteristics and dynamic state characteristics, resulting in inaccurate evaluation and lack of real-time adjustment capabilities.
By introducing the concept of node importance and combining static topology information with dynamic state characteristics, weighted TOPSIS, PCA, Hadamard product and other methods are used to construct an importance scoring model to screen out high-priority nodes, and the screening results are optimized through a dynamic feedback mechanism.
It achieves load optimization in active telemetry scenarios, improves probe resource utilization efficiency, enhances anomaly detection capabilities, adapts to dynamic network changes, and ensures the accuracy and real-time nature of screening results.
Smart Images

Figure CN119652605B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of cyberspace security, and in particular to a method for screening important nodes in active telemetry scenarios. Background Art
[0002] In recent years, with the diversification of business applications and the continued growth of user base, networks have gradually become characterized by high speed, large scale, multi-access, and unpredictable characteristics, creating an increasingly urgent need for efficient network measurement. As a core technology for network perception and management, network measurement is directly related to the implementation of key tasks such as network anomaly detection, performance optimization, and resource scheduling. In-band network telemetry (INT), a network measurement method developed in recent years based on programmable data plane technology, dynamically attaches telemetry metadata to data packets at switching nodes along the path, enabling real-time collection and hop-by-hop monitoring of network status. Compared to traditional network measurement solutions, in-band telemetry transcends the traditional approach of treating network switching devices as intermediate black boxes. Instead, it leverages the data plane to directly drive the network measurement process, offering enhanced real-time and fine-grained performance monitoring capabilities. It has gradually become a key direction in modern network measurement.
[0003] From an implementation perspective, network telemetry is primarily categorized into passive and active telemetry. Passive network telemetry uses existing service flows to carry telemetry metadata, efficiently collecting various network status information along target paths. Because passive telemetry relies on service traffic to cover the path, it struggles to achieve flexible and comprehensive monitoring of certain paths not traversed by traffic, limiting its network-wide applicability. In contrast, active telemetry, a key in-band telemetry technology, has gained widespread attention due to its flexibility and coverage. Unlike passive telemetry, which relies on service traffic to cover the path, the recent maturity of technologies such as Segment Routing (SR) has enabled probe-based active network telemetry to easily cover the entire network without being constrained by service traffic, and has gradually gained widespread attention. Active telemetry constructs specific probe packets to cover any target path and precisely customizes the path using source routing labels. This feature enables active telemetry to flexibly adapt to diverse network structures and dynamically changing service scenarios, providing unprecedented flexibility and adaptability for network status monitoring.
[0004] Despite the potential of active telemetry, its development in practice is still hampered by data redundancy and resource waste. Many probes collect relatively static information, lacking the anomalous events that upper-layer applications would notice. These ineffective probes not only consume valuable bandwidth but also significantly increase the load on north-south network traffic, further exacerbating communication pressure between the data plane and telemetry servers. This resource waste directly leads to reduced system performance and negatively impacts telemetry efficiency.
[0005] To address these issues, existing research has primarily focused on optimizing passive telemetry systems, aiming to improve the effectiveness of telemetry data by reducing system overhead. For example, Flexible Sampling introduces a dynamic sampling mechanism that dynamically adjusts the frequency of probe insertion based on changes in network traffic, thereby reducing the amount of meaningless telemetry data. Selective In-band Network Telemetry optimizes the allocation of probe resources by reducing the telemetry frequency when the network is stable. These methods have demonstrated excellent performance in reducing the load on passive telemetry systems and significantly improving their efficiency. However, because passive telemetry relies on existing traffic coverage paths, its optimization strategies are difficult to adapt to the needs of active telemetry. In particular, in active telemetry scenarios, the issue of how to rationally allocate probe resources to key nodes remains unresolved.
[0006] Furthermore, active telemetry is limited by its underlying framework, resulting in significant technical bottlenecks in practical applications. A large amount of information in the telemetry data is smooth and lacks value, with only a small portion containing abnormal events or critical states of interest to upper-layer applications. This redundant data not only wastes a significant amount of computing time and storage space on the telemetry server but also severely impacts overall system performance. Compared to passive telemetry, the additional INT probes introduced by active telemetry further increase network load, especially in the face of invalid probes. The resulting network bandwidth usage and communication pressure between the data plane and the telemetry server are even more significant.
[0007] In summary, existing research on reducing telemetry overhead has achieved certain results, but in active telemetry scenarios, the following problems exist, which make it impossible to accurately select important telemetry nodes according to specific network conditions: (1) Existing research results are mostly concentrated on passive telemetry systems, and lack targeted optimization of active telemetry scenarios. Passive telemetry relies on existing business traffic for path coverage, while active telemetry requires active construction of probes to complete network monitoring. This essential difference makes it difficult for the optimization strategy of passive telemetry to be directly applied to active telemetry scenarios, especially in the probe allocation and resource scheduling of important nodes, which cannot meet the flexibility requirements of active telemetry. (2) Current methods are usually based on a single network feature (such as degree centrality) and fail to comprehensively consider static topological features and dynamic state features, resulting in inaccurate node importance assessment and lack of comprehensiveness and reliability in the judgment of node importance. (3) Existing probe allocation strategies do not fully focus on important nodes. The data collected by some probes remain stable for a long time, and cannot effectively capture abnormal events that upper-layer applications are concerned about, resulting in a waste of probe resources. (4) The network state has the characteristics of dynamic change, but the existing methods are too static in evaluating node importance. The node screening mechanism lacks real-time adjustment capabilities and is difficult to adapt to the needs of dynamic network changes, thus affecting the effectiveness of node selection.
[0008] The important node selection method for active telemetry provided by the present invention effectively solves the technical difficulties caused by excessive additional load and unreasonable path planning of active telemetry in the current network security field. Summary of the Invention
[0009] To address these issues, this paper proposes a versatile probe frequency planning method for active telemetry, aiming to optimize the active telemetry load across the entire network. The innovation of this method lies in the first incorporation of the concept of telemetry node importance into probe frequency planning. Specifically, across the entire network, we consider that the importance of each node in the telemetry process varies. Some nodes are more important and require higher probe frequencies to ensure the real-time availability of their status information. For less important nodes, the probe frequency can be appropriately reduced, effectively reducing the additional load introduced by active telemetry.
[0010] To achieve the purpose of the present invention, the specific technical steps of this solution are as follows: A method for screening important nodes for active telemetry scenarios, the method comprising the following steps:
[0011] Step (1) In the telemetry system, distributed probes are used to collect network data plane data, including network node latency information, bandwidth occupancy, throughput information, and other node status information. To ensure data quality, noise filtering and missing value filling algorithms are used to clean the data. The cleaned data is stored in the form of a time series matrix for subsequent feature extraction and analysis.
[0012] Step (2) extracts the topological information of the nodes based on the basic data collected in step (1), analyzes the spatial factors affecting the nodes, and uses the weighted TOPSIS method to integrate multiple centrality measurements to obtain the spatial factor of one of the important node evaluation indicators;
[0013] Step (3) extracts important features that affect the node state based on the basic data collected in step (1), including dynamic features such as significant change frequency and state entropy. The significant change frequency is obtained by counting the number of significant changes in the node's periodic historical telemetry data, and the state entropy is calculated by calculating the state fluctuations in the node's periodic historical telemetry data. Based on the calculation results, a node state feature matrix is formed, and then high-dimensional state features are extracted. In order to eliminate feature redundancy and multicollinearity, principal component analysis (PCA) is used for dimensionality reduction.
[0014] In the actual deployment of step (4), many telemetry tasks are carried out simultaneously, that is, the probe message needs to collect a variety of network status data at each hop node. In the case of multiple tasks, each task data can reflect the status characteristics of the telemetry node. However, the correlation that may exist between different telemetry tasks will cause the final generated state factor to be inaccurate. Therefore, based on the dimensionality reduction results of step (3), the present invention processes the multi-source data features generated by multiple telemetry tasks, and fuses the different task features after standardization. The fusion result is used as a unified comprehensive state factor for the importance scoring model;
[0015] Step (5) Given a network topology with n nodes, construct the spatial importance and comprehensive state importance vectors of the entire network based on the spatial factor and state factor of each node obtained in steps (2) and (4);
[0016] Step (6) Based on the spatial importance vector and the optimal state importance vector obtained in step (5), the Hadamard product is used to construct an importance scoring model for the telemetry nodes. By calculating the scores, the importance of the nodes is arranged in descending order, and the high-priority nodes are marked;
[0017] Step (7) combines the historical telemetry data collected in step (1) and the node importance score result calculated in step (6) to set the score threshold σ. Nodes with a node importance score exceeding σ are screened as priority telemetry objects and serve as key targets for probe resource allocation. After the initial threshold is set, during the periodic operation of the telemetry system, the latest information on network topology and traffic status is collected in real time. The scoring model is continuously optimized and the threshold σ is updated through a dynamic feedback mechanism to ensure the real-time and accuracy of the screening results.
[0018] Step (8) To ensure the accuracy and reliability of the important node screening, the nodes screened in step (7) are applied to the anomaly detection and active telemetry probe allocation optimization experiment. Specifically, it includes verifying the effectiveness of the screening results, evaluating its performance by calculating the node coverage rate and the improvement ratio of anomaly detection capability, and comparing the high-scoring node coverage rate, anomaly detection capability and other indicators with the existing methods. Figure 2 The working framework for verification and application of important nodes in active telemetry systems is demonstrated.
[0019] Furthermore, in step (1), the steps of collecting and preprocessing basic data are as follows:
[0020] (1.1) In the network telemetry system, the control plane sends the probe configuration policy to the telemetry source node of the data plane through the north-south communication interface. The configuration includes the probe path, probe identification header information, and probe collection frequency.
[0021] (1.2) The telemetry probe is injected into the network according to the configuration policy in step (1.1). The data plane node processes the data packets matching the probe identifier header hop by hop, adds node status information (such as latency, bandwidth utilization, throughput information, etc.) hop by hop, and transmits it along the path to the telemetry tail node;
[0022] (1.3) The telemetry tail node packages the hop-by-hop data from the starting node to the end node and uploads it to the telemetry server. The telemetry server filters noise, fills in missing values, and normalizes the collected raw data to ensure data quality and consistency. The cleaned data is stored in a structured format for subsequent feature extraction and analysis.
[0023] Furthermore, in step (2), based on the basic data collected in step (1), the topological information of the nodes is extracted, and the spatial factors affecting the nodes are analyzed. The calculation method uses the weighted TOPSIS method to integrate multiple centrality measurement values to obtain the spatial factor of one of the important node evaluation indicators:
[0024]
[0025] Factor spatial =R=TOPSIS(M,W)
[0026] in:
[0027] DC i Degree Centrality (DC) is the number of nodes directly connected to node i;
[0028] CC iCloseness Centrality (CC) is the average path length between node i and other nodes;
[0029] EC i Eigenvector Centrality (EC) is a linear combination of the importance of neighboring nodes of node i;
[0030] Factor spatial It is a spatial factor and one of the important indicators for evaluating important nodes;
[0031] Furthermore, in step (3), the steps of node state feature extraction and feature dimension reduction are as follows:
[0032] (3.1) By performing time series analysis on the node traffic and connection status data collected in step (1), the number of significant changes in the node's historical telemetry data at certain stages (e.g., within a fixed time window of 10 minutes) is counted. This indicator is the significant change frequency, which reflects the degree of drastic changes in the node status. A larger value indicates that the node status changes more frequently.
[0033] (3.2) Based on the state fluctuations in the periodic historical telemetry data, the node state entropy is calculated to evaluate its state stability. A lower state entropy indicates a more stable node state, while a higher state entropy indicates a larger fluctuation in the surrounding traffic.
[0034] (3.3) To eliminate the differences between different feature dimensions, the present invention performs standardization on the state feature data. After standardization, the mean of all feature data is 0 and the variance is 1, ensuring that the weights of different features are balanced during the dimensionality reduction process.
[0035] (3.4) Calculate the covariance matrix C for the feature data matrix X after standardization in step (3.3) to evaluate the correlation between the features. The covariance matrix used in the present invention is a symmetric matrix that reveals the strength of the correlation between the features.
[0036] (3.5) In order to ensure that the key information of high-dimensional feature data is retained as much as possible during the dimensionality reduction process, the present invention performs eigenvalue decomposition on the covariance matrix in step (3.4), arranges the eigenvalues in descending order, calculates the cumulative contribution rate, and selects the top k principal components with a cumulative contribution rate of 95% to construct the eigenvector after dimensionality reduction.
[0037] Furthermore, in step (4), the steps of normalizing telemetry task features and calculating task feature weights are as follows:
[0038] (4.1) Multi-task telemetry requires collecting n sets of indicators (delay, bandwidth, etc.), and the telemetry result set of each set of indicators is Xi ={x′ i1 ,x′ i2 ,···,x′ ik Considering that the values of various indicators may not be of the same order of magnitude, and the positive and negative attributes of various telemetry tasks as network health evaluation indicators are different: for positive indicators, the larger the corresponding collected information, the better the network condition (such as the available bandwidth of the port); for negative indicators, the smaller the corresponding collected information, the better the network condition (such as the number of packet losses). Therefore, for telemetry result sets of different projects, it is necessary to pre-calculate x i Do standardized processing:
[0039] standardization: Where k is the number of probe telemetry information collected periodically, x′ ik is a single telemetry result;
[0040] (4.2) After normalizing the indicators, the task characteristics are weighted:
[0041] Weight w i calculate:
[0042] where x′ i ={x′ i1 ,x′ i2 ,···,x′ ik}, pc ij represents the Pearson correlation coefficient between telemetry indicators;
[0043] Among them CS i Represents the association strength and is used to assign weights to different telemetry tasks;
[0044] where w i Represents the weight of the telemetry task.
[0045] (4.3)The formula for calculating the state factor is as follows:
[0046] Factor state =w1×F1+w2×F2+…+w n ×F n , where Factor state Represents the comprehensive state factor of the node, w i represents the weight of the telemetry task, F i Represents the state features after dimensionality reduction obtained in the telemetry task.
[0047] Step (5) Given a network topology with n nodes, construct the spatial importance and comprehensive state importance vectors of the entire network based on the spatial factor and comprehensive state factor of each node obtained in steps (2) and (4);
[0048] In one embodiment of the present invention, the specific implementation steps of the important node scoring model are as follows:
[0049] (5.1) First, based on the spatial factor of each telemetry node, the spatial importance vector of the entire network is obtained:
[0050] Where V spatial represents the spatial importance vector, is the spatial factor found in step (2).
[0051] (5.2) According to the telemetry state cycle, the highest value among the comprehensive state factors generated by the node ports is selected as the optimal comprehensive state factor and used to calculate the state importance vector of the network:
[0052] in represents the calculation formula in step (4.3) to find the periodic comprehensive state factor of port j of node i, is the optimal comprehensive state factor state of the node;
[0053] Where V state Represents the state importance vector of the node, is the optimal comprehensive state factor obtained in step (5.2).
[0054] Furthermore, in step (6), based on the spatial importance vector and the optimal state importance vector obtained in step (5), the Hadamard product is used to construct an importance scoring model for the telemetry nodes. By calculating the scores, the importance of the nodes is arranged in descending order, and the high-priority nodes are marked:
[0055] Where ⊙ represents Hadamard product, Tele imp Represents the telemetry node importance vector, V spatial Represents the spatial importance vector, V state Represents the state importance vector of the node, and are the spatial factor and comprehensive state factor obtained in steps (2) and (5.2) respectively.
[0056] Furthermore, in step (7), the specific implementation steps of the dynamic adjustment mechanism and node screening are as follows:
[0057] (7.1) According to the telemetry feedback data and a large amount of prior data, the scoring threshold σ is set;
[0058] (7.2) After multiple rounds of telemetry, all nodes are re-evaluated with the new spatial factors and state factors, and the list of important nodes is updated to ensure that the screening results match the real-time network status.
[0059] Furthermore, in step (8), the specific implementation steps for verifying the important node screening results are as follows:
[0060] Step (8.1) simulates abnormal events in the test network environment, including but not limited to traffic surges, increased packet loss rates, and latency jitter. Compare these events with the locations of the selected important nodes and record the proportion of abnormal events captured by the important nodes (capture rate);
[0061] Step (8.2) calculates the coverage based on the records obtained in step (8.1) to evaluate the effectiveness of the filtered nodes in anomaly detection compared to the probe configuration with full network coverage.
[0062] An electronic device includes a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the program, it implements the important node screening method for active telemetry scenarios.
[0063] Compared with the prior art, the technical solution of the present invention has the following beneficial technical effects:
[0064] (1) The present invention proposes a method for screening important nodes suitable for active telemetry scenarios, focusing on improving screening efficiency and accuracy through collaborative analysis of key dimensions of state characteristics in static topology information and historical telemetry information. The method is based on the structural status of the node in the network, comprehensively considers its static position contribution in the topology, and combines characteristics such as significant change frequency and state entropy to comprehensively capture potential anomalies and network events in real-time traffic. Compared with traditional methods, the present invention can flexibly adapt to the diversity of complex network environments, achieve accurate and efficient node selection by dynamically adjusting screening criteria, and provide reliable support for the efficient allocation of active telemetry probes and network anomaly detection.
[0065] (2) In the process of processing node state features, the present invention introduces the principal component analysis (PCA) method to reduce the dimensionality of multiple features, aiming to reduce redundant information and extract key features. In view of the possible correlation between node state features such as significant change frequency and state entropy, the present invention optimizes through linear combination, which not only retains the effective information in the original features, but also effectively eliminates the redundant interference between features. This dimensionality reduction method can fully reflect the real state changes of nodes, significantly improve the accuracy of important node screening, and optimize the utilization efficiency of computing resources.
[0066] (3) The present invention is aimed at multi-telemetry task scenarios, where each task data may reflect different state characteristics of the node. However, since there may be certain correlations between tasks, direct merging may lead to information redundancy or result deviation. By analyzing the correlations between telemetry tasks, a reasonable weight is assigned to each task, reducing information overlap while highlighting the contribution of important tasks. When allocating weights, indicators such as the Pearson correlation coefficient are used to quantify the correlation between tasks to ensure that the comprehensive state characteristics can objectively reflect the importance of the node. This method can flexibly adapt to different telemetry needs in a multi-task environment and provide strong support for the identification of key nodes in complex networks.
[0067] (4) The present invention proposes an importance assessment model that systematically evaluates the importance of network nodes by fusing static network topology features with dynamic traffic state features. The model uses a weighted fusion approach to combine the static features of the topology structure (such as degree centrality, closeness centrality, and eigenvector centrality) with the dynamic features of the node state (such as significant change frequency and state entropy), and generates the final importance score vector through the Hadamard product. This model can not only dynamically adjust the importance score threshold of the node in different telemetry tasks, but also highlight the role of high-priority nodes in resource scheduling, ensuring the pertinence and effectiveness of the screening results in subsequent telemetry tasks, and overall optimizing the utilization efficiency of probe resources and improving the overall performance of the telemetry system. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] Figure 1 A framework diagram of the important node screening method for active telemetry scenarios;
[0069] Figure 2 Assist active telemetry probe optimization process diagram for important node telemetry information feedback; DETAILED DESCRIPTION
[0070] The technical solutions provided by the present invention will be described in detail below with reference to specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention.
[0071] Example: The present invention provides an important node screening method for active telemetry scenarios, the overall system structure of which is as follows: Figure 1 As shown, the following steps are included:
[0072] Step (1) In the telemetry system, distributed probes are used to collect network data plane data, including network node latency information, bandwidth occupancy, throughput information, and other node status information. To ensure data quality, noise filtering and missing value filling algorithms are used to clean the data. The cleaned data is stored in the form of a time series matrix for subsequent feature extraction and analysis.
[0073] In one embodiment of the present invention, the steps of collecting and preprocessing basic data are as follows:
[0074] (1.1) In the network telemetry system, the control plane sends the probe configuration policy to the telemetry source node of the data plane through the north-south communication interface. The configuration includes the probe path, probe identification header information, and probe collection frequency.
[0075] (1.2) The telemetry probe is injected into the network according to the configuration policy in step (1.1). The data plane node processes the data packets matching the probe identifier header hop by hop, adds node status information (such as latency, bandwidth utilization, throughput information, etc.) hop by hop, and transmits it along the path to the telemetry tail node;
[0076] (1.3) The telemetry tail node packages the hop-by-hop data from the starting node to the end node and uploads it to the telemetry server. The telemetry server filters noise, fills in missing values, and normalizes the collected raw data to ensure data quality and consistency. The cleaned data is stored in a structured format for subsequent feature extraction and analysis.
[0077] Step (2) extracts the topological information of the nodes based on the basic data collected in step (1) and analyzes the spatial factors that affect the nodes. The calculation method uses the weighted TOPSIS method to integrate multiple centrality measurements to obtain the spatial factor, which is one of the important node evaluation indicators:
[0078]
[0079] Factor spatial =R=TOPSIS(M,W)
[0080] in:
[0081] DC i Degree Centrality (DC) is the number of nodes directly connected to node i;
[0082] CC iCloseness Centrality (CC) is the average path length between node i and other nodes;
[0083] EC i Eigenvector Centrality (EC) is a linear combination of the importance of neighboring nodes of node i;
[0084] Factor spatial It is a spatial factor and one of the important indicators for evaluating important nodes;
[0085] Step (3) extracts important features that affect the node state based on the basic data collected in step (1), including dynamic features such as significant change frequency and state entropy. The significant change frequency is obtained by counting the number of significant changes in the node's periodic historical telemetry data, and the state entropy is calculated by calculating the state fluctuations in the node's periodic historical telemetry data. Based on the calculation results, a node state feature matrix is formed, and then high-dimensional state features are extracted. In order to eliminate feature redundancy and multicollinearity, principal component analysis (PCA) is used for dimensionality reduction.
[0086] In one embodiment of the present invention, the steps of dimensionality reduction and optimization between features are as follows:
[0087] (3.1) By performing time series analysis on the node traffic and connection status data collected in step (1), the number of significant changes in the node's historical telemetry data at certain stages (e.g., within a fixed time window of 10 minutes) is counted. This indicator is the significant change frequency, which reflects the degree of drastic changes in the node status. A larger value indicates that the node status changes more frequently.
[0088] (3.2) Based on the state fluctuations in the periodic historical telemetry data, the node state entropy is calculated to evaluate the stability of its state.
[0089] (3.3) To eliminate the differences between different feature dimensions, the present invention performs standardization on the state feature data. After standardization, the mean of all feature data is 0 and the variance is 1, ensuring that the weights of different features are balanced during the dimensionality reduction process.
[0090] (3.4) Calculate the covariance matrix C for the feature data matrix X after standardization in step (3.3) to evaluate the correlation between the features. The covariance matrix used in the present invention is a symmetric matrix that reveals the strength of the correlation between the features.
[0091] (3.5) In order to ensure that the key information of high-dimensional feature data is retained as much as possible during the dimensionality reduction process, the present invention performs eigenvalue decomposition on the covariance matrix in step (3.4), arranges the eigenvalues in descending order, calculates the cumulative contribution rate, and selects the top k principal components with a cumulative contribution rate of 95% to construct the eigenvector after dimensionality reduction.
[0092] In the actual deployment of step (4), many telemetry tasks are carried out simultaneously, that is, the probe message needs to collect a variety of network status data at each hop node. In the case of multiple tasks, each task data can reflect the status characteristics of the telemetry node. However, the correlation that may exist between different telemetry tasks will cause the final generated state factor to be inaccurate. Therefore, based on the dimensionality reduction results of step (3), the present invention processes the multi-source data features generated by multiple telemetry tasks, and fuses the different task features after standardization. The fusion result is used as a unified comprehensive state factor for the importance scoring model;
[0093] In one embodiment of the present invention, the steps of normalizing telemetry task features and calculating task feature weights are as follows:
[0094] (4.1) Multi-task telemetry requires collecting n sets of indicators (delay, bandwidth, etc.), and the telemetry result set of each set of indicators is i = {x′ i1 ,x′ i2 ,···,x′ ik Considering that the values of various indicators may not be of the same order of magnitude, and the positive and negative attributes of various telemetry tasks as network health evaluation indicators are different: for positive indicators, the larger the corresponding collected information, the better the network condition (such as the available bandwidth of the port); for negative indicators, the smaller the corresponding collected information, the better the network condition (such as the number of packet losses). Therefore, for telemetry result sets of different projects, it is necessary to pre-calculate x i Do standardized processing:
[0095] standardization: Where k is the number of probe telemetry information collected periodically, x′ ik is a single telemetry result;
[0096] (4.2) After normalizing the indicators, the task characteristics are weighted:
[0097] Weight w i calculate:
[0098] where X′ i ={x′ i1 ,x′ i2 ,···,x′ ik}, pcij represents the Pearson correlation coefficient between telemetry indicators;
[0099] Among them CS i Represents the association strength and is used to assign weights to different telemetry tasks;
[0100] where w i Represents the weight of the telemetry task.
[0101] (4.3)The formula for calculating the state factor is as follows:
[0102] Factor state =w1×F1+w2×F2+…+w n ×F n , where Factor state Represents the comprehensive state factor of the node, w i represents the weight of the telemetry task, F i Represents the state features after dimensionality reduction obtained in the telemetry task.
[0103] Step (5) Given a network topology with n nodes, construct the spatial importance and comprehensive state importance vectors of the entire network based on the spatial factor and comprehensive state factor of each node obtained in steps (2) and (4);
[0104] In one embodiment of the present invention, the specific implementation steps of the important node scoring model are as follows:
[0105] (5.1) First, based on the spatial factor of each telemetry node, the spatial importance vector of the entire network is obtained:
[0106] Where V spatial represents the spatial importance vector, is the spatial factor found in step (2).
[0107] (5.2) According to the telemetry state cycle, the highest value among the comprehensive state factors generated by the node ports is selected as the optimal comprehensive state factor and used to calculate the state importance vector of the network:
[0108] in represents the calculation formula in step (4.3) to find the periodic comprehensive state factor of port j of node i, is the optimal comprehensive state factor state of the node;
[0109] Where V state Represents the state importance vector of the node, is the optimal comprehensive state factor obtained in step (5.2).
[0110] Step (6) Based on the spatial importance vector and the optimal state importance vector obtained in step (5), the Hadamard product is used to construct an importance scoring model for the telemetry nodes. By calculating the scores, the importance of the nodes is arranged in descending order, and the high-priority nodes are marked;
[0111] In one embodiment of the present invention, the comprehensive eigenvalue is used as input, and the Hadamard product operation is performed to multiply the corresponding elements of two matrices of the same size, namely the space vector and the state vector, one by one, thereby constructing an importance scoring model. The scoring formula is as follows:
[0112] Where ⊙ represents Hadamard product, Tele imp Represents the telemetry node importance vector, V spatial Represents the spatial importance vector, V state Represents the state importance vector of the node, and are the spatial factor and comprehensive state factor obtained in steps (2) and (5.2) respectively.
[0113] Step (7) combines the historical telemetry data collected in step (1) and the node importance score result calculated in step (6) to set the score threshold σ. Nodes with a node importance score exceeding σ are screened as priority telemetry objects and serve as key targets for probe resource allocation. After the initial threshold is set, during the periodic operation of the telemetry system, the latest information on network topology and traffic status is collected in real time. The scoring model is continuously optimized and the threshold σ is updated through a dynamic feedback mechanism to ensure the real-time and accuracy of the screening results.
[0114] In one embodiment of the present invention, the specific implementation steps of the dynamic adjustment mechanism and node screening are as follows:
[0115] (7.1) According to the telemetry feedback data and a large amount of prior data, the scoring threshold σ is set;
[0116] (7.2) After multiple rounds of telemetry, all nodes are re-evaluated with the new spatial factors and state factors, and the list of important nodes is updated to ensure that the screening results match the real-time network status.
[0117] Step (8) To ensure the accuracy and reliability of the important node screening, the nodes screened in step (7) are applied to the anomaly detection and active telemetry probe allocation optimization experiment. Specifically, it includes verifying the effectiveness of the screening results, evaluating its performance by calculating the node coverage rate and the improvement ratio of anomaly detection capability, and comparing the high-scoring node coverage rate, anomaly detection capability and other indicators with the existing methods. Figure 2 The working framework for verification and application of important nodes in active telemetry systems is demonstrated.
[0118] In one embodiment of the present invention, the specific implementation steps for verifying the important node screening results are as follows:
[0119] Step (8.1) simulates abnormal events in the test network environment, including but not limited to traffic surges, increased packet loss rates, and latency jitter. Compare these events with the locations of the selected important nodes and record the proportion of abnormal events captured by the important nodes (capture rate);
[0120] Step (8.2) calculates the coverage based on the records obtained in step (8.1) to evaluate the effectiveness of the filtered nodes in anomaly detection compared to the probe configuration with full network coverage.
[0121] The technical means disclosed in the solutions of the present invention are not limited to those disclosed in the above-mentioned embodiments, but also include technical solutions composed of any combination of the above-mentioned technical features. It should be noted that those skilled in the art may make various improvements and modifications without departing from the principles of the present invention, and such improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. A method for screening important nodes in active telemetry scenarios, characterized in that: The method comprises the following steps: Step (1) In the telemetry system, network data plane data is collected through distributed probes. The data includes network node delay information, bandwidth occupancy, throughput information, and other node status. To ensure data quality, noise filtering and missing value filling algorithms are used to clean the data. The cleaned data is stored in the form of a time series matrix for subsequent feature extraction and analysis. Step (2) extracts the topological information of the nodes based on the basic data collected in step (1), analyzes the spatial factors affecting the nodes, and uses the weighted TOPSIS method to integrate multiple centrality measurements to obtain the spatial factor, which is one of the important node evaluation indicators; Step (3) extracts important features that affect the node state based on the basic data collected in step (1), including significant change frequency and state entropy dynamic features. The significant change frequency is obtained by counting the number of significant changes in the node's periodic historical telemetry data. The state entropy is calculated by calculating the state fluctuation in the node's periodic historical telemetry data. The node state feature matrix is formed based on the calculation results, and then high-dimensional state features are extracted. In order to eliminate feature redundancy and multicollinearity, principal component analysis is used for dimensionality reduction. In the actual deployment of step (4), many telemetry tasks are carried out simultaneously, that is, the probe message needs to collect a variety of network status data at each hop node. In the case of multiple tasks, each task data can reflect the status characteristics of the telemetry node. Based on the dimensionality reduction results in step (3), the multi-source data features generated by multiple telemetry tasks are processed, and the different task features are fused after standardization. The fusion result is used as a unified comprehensive state factor for the importance scoring model; Step (5) Given a Based on the network topology of each node, the spatial importance and comprehensive state importance vectors of the entire network are constructed according to the spatial factor and comprehensive state factor of each node obtained in steps (2) and (4); Step (6) Based on the spatial importance vector and the comprehensive state importance vector obtained in step (5), the Hadamard product is used to construct an importance scoring model for the telemetry nodes. By calculating the scores, the importance of the nodes is arranged in descending order, and high-priority nodes are marked. Step (7) combines the historical telemetry data collected in step (1) and the node importance score calculated in step (6) to set the score threshold , the node importance score exceeds The nodes are selected as priority telemetry objects and serve as the key targets for probe resource allocation. After the initial threshold is set, the latest information on network topology and traffic status is collected in real time during the periodic operation of the telemetry system. The scoring model is continuously optimized and the threshold is updated through a dynamic feedback mechanism. , ensuring the real-time and accuracy of screening results; In order to ensure the accuracy and reliability of the important node screening, step (8) applies the nodes screened out in step (7) to the allocation optimization experiment of anomaly detection and active telemetry probes. Specifically, it verifies the effectiveness of the screening results, evaluates its performance by calculating the improvement ratio of node coverage and anomaly detection capability, and compares the high-scoring node coverage and anomaly detection capability indicators with existing methods.
2. The important node screening method for active telemetry scenarios according to claim 1 is characterized in that: The step (1) specifically includes the following sub-steps: (1.1) In the network telemetry system, the control plane sends the probe configuration policy to the telemetry source node of the data plane through the north-south communication interface. The configuration includes the probe path, probe identification header information, and probe collection frequency. (1.2) The telemetry probe is injected into the network according to the configuration policy in step (1.1). The data plane node processes the data packet matching the probe identifier header hop by hop, adds node status information hop by hop, and transmits it along the path to the telemetry tail node; (1.3) The telemetry tail node packages the hop-by-hop data set from the starting node to the terminal node and uploads it to the telemetry server. The telemetry server performs noise filtering, missing value filling, and normalization on the collected raw data to ensure data quality and consistency. The cleaned data is stored in a structured format for subsequent feature extraction and analysis.
3. The important node screening method for active telemetry scenarios according to claim 1 is characterized in that: Based on the basic data collected in step (1), the topological information of the nodes is extracted and the spatial factors affecting the nodes are analyzed. The calculation method uses the weighted TOPSIS method to integrate multiple centrality measurements to obtain the spatial factor, which is one of the important node evaluation indicators: in: Degree centrality, for nodes The number of nodes directly connected to this node; Closeness centrality, for nodes Average path length between nodes; Eigenvector centrality, for nodes A linear combination of the importance of adjacent nodes; It is a spatial factor and one of the important indicators for evaluating important nodes.
4. The important node screening method for active telemetry scenarios according to claim 1 is characterized in that: The step (3) specifically includes the following sub-steps: (3.1) By performing time series analysis on the node traffic and connection status data collected in step (1), the number of significant changes in the node's periodic historical telemetry data is counted; (3.2) Based on the state fluctuations in the periodic historical telemetry data, calculate the node state entropy to evaluate the stability of its state; (3.3) Standardize the state feature data. After standardization, the mean of all feature data is 0 and the variance is 1, ensuring that the weights of different features are balanced during the dimensionality reduction process. (3.4) After the feature data matrix is standardized in step (3.3), Calculate the covariance matrix , used to evaluate the correlation between features. The covariance matrix used is a symmetric matrix, which reveals the strength of the correlation between features; (3.5) Perform eigenvalue decomposition on the covariance matrix in step (3.4), sort them in descending order according to the size of the eigenvalues, calculate the cumulative contribution rate and select the top ones with a cumulative contribution rate of 95%. principal components and construct the eigenvector after dimensionality reduction.
5. The important node screening method for active telemetry scenarios according to claim 1 is characterized in that: The step (4) specifically includes the following sub-steps: (4.1) Multi-task telemetry needs to be collected Group indicators, the telemetry result set of each group of indicators is , considering that the values of various indicators may not belong to the same order of magnitude, and the positive and negative attributes of various telemetry tasks as network health evaluation indicators are different: for positive indicators, the larger the corresponding collected information, the better the network condition; for negative indicators, the smaller the corresponding information, the better the network condition. Therefore, for telemetry result sets of different projects, it is necessary to pre- Do standardized processing: standardization: ,in, is the number of probe telemetry information collected periodically, is a single telemetry result; (4.2) After normalizing the indicators, weight the task characteristics: Weight calculate: ,in , represents the Pearson correlation coefficient between telemetry indicators; ,in Represents the association strength and is used to assign weights to different telemetry tasks; ,in Represents the weight of the telemetry task, (4.3)The formula for calculating the state factor is as follows: ,in Represents the comprehensive state factor of the node, Represents the weight of the telemetry task, Represents the state features after dimensionality reduction obtained in the telemetry task.
6. The important node screening method for active telemetry scenarios according to claim 1 is characterized in that: The step (5) specifically includes the following sub-steps: Step (5) Given a Based on the network topology of each node, the spatial importance and comprehensive state importance vectors of the entire network are constructed according to the spatial factor and comprehensive state factor of each node obtained in steps (2) and (4); The specific implementation steps of the important node scoring model are as follows: (5.1) First, based on the spatial factor of each telemetry node, the spatial importance vector of the entire network is obtained: ,in represents the spatial importance vector, is the spatial factor found in step (2), (5.2) According to the telemetry state cycle, the highest value among the comprehensive state factors generated by the node ports is selected as the optimal comprehensive state factor and used to calculate the state importance vector of the network: ,in Represents the use of the calculation formula in step (4.3) to find the node Port The periodic comprehensive state factor, is the optimal comprehensive state factor state of the node; ,in Represents the state importance vector of the node, is the optimal comprehensive state factor obtained in step (5.2).
7. The important node screening method for active telemetry scenarios according to claim 1 is characterized in that: In step (6), the Hadamard product is used to construct an importance scoring model for telemetry nodes. By calculating the scores, the importance of the nodes is arranged in descending order, and high-priority nodes are marked. The scoring formula is as follows: ,in Representing Hadamard, represents the telemetry node importance vector, represents the spatial importance vector, Represents the state importance vector of the node, and are the spatial factor and comprehensive state factor obtained in steps (2) and (5.2), respectively.
8. The important node screening method for active telemetry scenarios according to claim 1 is characterized in that: The step (7) specifically includes the following sub-steps: (7.1) Set scoring thresholds based on telemetry feedback data and a large amount of prior data ; (7.2) After multiple rounds of telemetry, all nodes are re-evaluated with the new spatial factors and state factors, and the list of important nodes is updated to ensure that the screening results match the real-time network status.
9. The important node screening method for active telemetry scenarios according to claim 1 is characterized in that: The step (8) specifically includes the following sub-steps: Step (8.1) simulates abnormal events in the test network environment, including traffic surges, increased packet loss rates, and latency jitter. Compare these events with the locations of the selected important nodes and record the proportion of abnormal events captured by the important nodes. Step (8.2) calculates the coverage based on the records obtained in step (8.1) to evaluate the effectiveness of the filtered nodes in anomaly detection compared to the probe configuration with full network coverage.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method for screening important nodes for active telemetry scenarios as described in any one of claims 1 to 9 above is implemented.