A data path identification method, device, and electronic device

By standardizing data packet information and performing privacy-preserving feature space analysis, combined with adaptive thresholding and subspace recognition, a path search algorithm is used to capture micro-data transmission anomalies in real time, solving the need for rapid response in high-speed network environments and realizing real-time detection of anomalies.

CN122457520APending Publication Date: 2026-07-24ZUNYI BRANCH OF CHINA MOBILE GRP GUIZHOU COMPANY +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZUNYI BRANCH OF CHINA MOBILE GRP GUIZHOU COMPANY
Filing Date
2026-04-01
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing technologies are insufficient to capture micro-data transmission anomalies in real time in high-speed network environments, thus failing to meet the need for rapid response.

Method used

After acquiring data packet information and performing standardized processing, anomaly detection is performed using a preset hybrid model in the privacy-preserving feature space. Combined with adaptive threshold strategy and subspace recognition, anomalies are captured in real time using a path search algorithm.

Benefits of technology

It enables real-time capture of micro-data transmission anomalies in high-speed network environments, avoiding the inability to meet rapid response requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122457520A_ABST
    Figure CN122457520A_ABST
Patent Text Reader

Abstract

The application discloses a data path identification method and device and electronic equipment. First, data packet information is acquired, and then the data packet information is standardized to obtain standardized data. The standardized data and a preset hybrid model are used for processing in a privacy protection feature space to obtain abnormality discrimination information. The adaptive threshold strategy is used for detecting the abnormality discrimination information to obtain a first detection result. The first detection result is identified based on a subspace to obtain position information. The position information is processed based on a path search algorithm to obtain a first path. The first path is processed, so that micro data transmission abnormalities can be captured in real time, and the demand for rapid response in a high-speed network environment can be met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to artificial intelligence, and to, but is not limited to, a method, apparatus, and electronic device for identifying data paths. Background Technology

[0002] With social development and progress, data path identification technology is receiving increasing attention. Currently, the more advanced data transmission path monitoring technology usually adopts a combination of passive monitoring and active detection. It collects data packet transmission information by deploying probes at key network nodes and analyzes the data packet header information to determine the transmission path status.

[0003] Existing technologies and path monitoring methods struggle to capture anomalies in micro-data transmission in real time, failing to meet the rapid response requirements of high-speed network environments. Summary of the Invention

[0004] In view of this, embodiments of this application provide a method, apparatus, and electronic device for identifying data paths.

[0005] The technical solution of this application embodiment is implemented as follows: This application provides a data path identification method, the method comprising: acquiring data packet information; standardizing the data packet information to obtain standardized data; processing the standardized data and a preset hybrid model in a privacy-preserving feature space to obtain anomaly discrimination information; detecting the anomaly discrimination information based on an adaptive threshold strategy to obtain a first detection result; identifying the first detection result based on a subspace to obtain location information; and processing the location information based on a path search algorithm to obtain a first path.

[0006] Optionally, the data packet information is standardized to obtain standardized data, including: extracting the data packet information based on the time window segmentation method to obtain time-series features; and performing normalization processing based on the time-series features to obtain the standardized data.

[0007] Optionally, processing is performed on the standardized data and the preset hybrid model in a privacy-preserving feature space to obtain anomaly detection information, including: performing sensitivity calculation on the standardized data to obtain a first sensitivity; performing HSIC independence calculation with noise compensation based on the first sensitivity to obtain a first calculation result; and processing is performed on the first calculation result and the preset hybrid model in the privacy-preserving feature space to obtain the anomaly detection information.

[0008] Optionally, the anomaly detection information is obtained by processing the first calculation result and the preset hybrid model in a privacy-preserving feature space, including: calculating the first calculation result based on a dynamic feature filtering strategy to obtain a second calculation result; and processing the second calculation result and the preset hybrid model in a privacy-preserving feature space to obtain the anomaly detection information.

[0009] Optionally, the step of detecting the anomaly discrimination information based on an adaptive threshold strategy to obtain a first detection result includes: performing multi-index time series prediction on the anomaly discrimination information to obtain a first performance index; performing dynamic threshold-based calculation on the first performance index to obtain a first threshold; and processing the first threshold based on an adaptive weight adjustment mechanism to obtain the first detection result.

[0010] Optionally, before identifying the first detection result based on the subspace to obtain location information, the method includes: acquiring historical abnormal patterns; performing feature activation intensity vector fusion based on the historical abnormal patterns to obtain the subspace; and identifying the first detection result based on the subspace to obtain the location information.

[0011] Optionally, the step of identifying the first detection result based on the subspace to obtain location information includes: identifying the first detection result based on the subspace to obtain a first abnormal sample; and locating the first abnormal sample based on the abnormal link to obtain the location information.

[0012] Optionally, the step of processing the location information based on the path search algorithm to obtain the first path includes: obtaining the path search algorithm based on a genetic algorithm and a simulated annealing strategy; and processing the location information based on the path search algorithm to obtain the first path.

[0013] An identification device includes: an acquisition unit, an analysis unit, and a processing unit; the acquisition unit is used to acquire data packet information; the analysis unit is used to standardize the data packet information to obtain standardized data; process the standardized data and a preset hybrid model in a privacy-preserving feature space to obtain anomaly discrimination information; detect the anomaly discrimination information based on an adaptive threshold strategy to obtain a first detection result; identify the first detection result based on a subspace to obtain location information; and the processing unit is used to process the location information based on a path search algorithm to obtain a first path.

[0014] An electronic device includes: a memory for storing at least one set of instructions; a processor for acquiring data packet information; standardizing the data packet information to obtain standardized data; processing the standardized data and a preset hybrid model in a privacy-preserving feature space to obtain anomaly detection information; detecting the anomaly detection information based on an adaptive threshold strategy to obtain a first detection result; identifying the first detection result based on a subspace to obtain location information; and processing the location information based on a path search algorithm to obtain a first path.

[0015] This application provides a data path identification method, apparatus, and electronic device. First, data packet information is acquired. Then, the data packet information is standardized to obtain standardized data. Next, based on the standardized data and a preset hybrid model, processing is performed in a privacy-preserving feature space to obtain anomaly detection information. Then, based on an adaptive threshold strategy, the anomaly detection information is detected to obtain a first detection result. Next, the first detection result is identified based on a subspace to obtain location information. Then, based on a path search algorithm, the location information is processed to obtain a first path. Processing is performed based on the first path to capture micro-data transmission anomalies in real time, avoiding the inability to quickly respond to demands in a high-speed network environment. Attached Figure Description

[0016] Figure 1 A flowchart illustrating the data path-based identification method provided in this application embodiment; Figure 2 Another flowchart of the data path-based identification method provided in the embodiments of this application; Figure 3 Another flowchart of the data path-based identification method provided in the embodiments of this application; Figure 4 Another flowchart of the data path-based identification method provided in the embodiments of this application; Figure 5 Another flowchart of the data path-based identification method provided in the embodiments of this application; Figure 6 Another flowchart of the data path-based identification method provided in the embodiments of this application; Figure 7 Another flowchart of the data path-based identification method provided in the embodiments of this application; Figure 8 Another flowchart of the data path-based identification method provided in the embodiments of this application; Figure 9 A schematic diagram of the identification device provided in the embodiments of this application; Figure 10 This is a schematic diagram of the structural composition of the electronic device provided in the embodiments of this application. Detailed Implementation

[0017] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0018] Please refer to Figure 1 ,in, Figure 1 A flowchart illustrating an implementation of a data path-based identification method provided in this application embodiment may include: Step S101: Obtain data packet information; Step S102: Standardize the data packet information to obtain standardized data; Step S103: Process the data in the privacy-preserving feature space based on standardized data and a preset hybrid model to obtain anomaly detection information; Step S104: Detect anomaly discrimination information based on an adaptive threshold strategy to obtain the first detection result; Step S105: Identify the first detection result based on the subspace to obtain location information; Step S106: Process the location information based on the path search algorithm to obtain the first path.

[0019] Among them, the data packet information is obtained by deploying multi-layer network probes at key network nodes and capturing network data packet information in real time through mirroring ports and bypass acquisition technology; the standardized processing includes at least: data packet timing feature extraction, transmission path feature modeling, inter-packet delay pattern analysis, feature vector standardization and integration; the preset hybrid model is an improved Gaussian mixture model (GMM); and the location information is the specific location where the anomaly occurred.

[0020] Specifically, step 1: Construct a deep network packet inspection and analysis framework. 1.1 Network Probe Deployment and Data Acquisition Multi-layered network probes are deployed at key network nodes to capture network data packet information in real time through port mirroring and bypass acquisition techniques. The probes employ a combination of flow sampling and packet sampling to ensure the comprehensiveness and representativeness of the data collected.

[0021] 1.2 Data Packet Temporal Feature Extraction By using time window segmentation and statistical analysis methods, the temporal characteristics of data packets are extracted, including inter-packet delay distribution, burst index, periodic patterns, etc., and a temporal characteristic descriptor is generated.

[0022] 1.3 Transmission Path Feature Modeling By analyzing packet header information (TTL value changes, IP options, DSCP markers, etc.) and path tracing techniques, a packet transmission path model is constructed, and path stability indicators and topology change characteristics are calculated.

[0023] 1.4 Analysis of Private Room Delay Patterns We use statistical learning methods to analyze the delay patterns of data packet transmission and establish a baseline model of delay distribution under normal transmission conditions.

[0024] 1.5 Feature Vector Standardization and Integration The extracted features are normalized, and a standardized set of data packet feature vectors is generated through a feature fusion algorithm.

[0025] Step 2: Hilbert-Schmidt Independence Modeling Based on Rényi Differential Privacy 2.1 Adaptive Rényi Differential Privacy (RDP) Mechanism for Real-Time Streaming Data To meet microsecond-level response requirements, this solution improves the standard RDP mechanism with real-time and adaptive features, enabling it to dynamically adapt to changes in network traffic while ensuring strict privacy protection (providing quantifiable...). Under the premise of protection, minimize the impact of noise on the accuracy of subsequent feature analysis.

[0026] Dynamic global sensitivity Calculate: Global sensitivity in traditional dynamic programming It is a fixed worst-case upper bound, which is often too conservative in real network traffic, leading to the addition of excessive noise. This solution calculates the sensitivity of the feature dimension in real time.

[0027] Implementation: The system maintains a sliding statistics table of the feature value within a short time window (e.g., the past 100 milliseconds). For numerical features (e.g., packet size), the sensitivity is... It is approximated as the mean absolute deviation (MAD) or scaled range of the observations within the window. For categorical features (such as protocol type), its sensitivity is defined as the upper bound of the change in class distribution entropy within the window. This dynamic calculation allows... It more closely reflects the actual fluctuation range of current traffic, thus allowing for the addition of more accurate and less noise.

[0028] Privacy Budget Adaptive scheduling: fixed The current value cannot adapt to the balance between security and utility requirements at different times of the network. This solution designs a privacy budget scheduler.

[0029] Implementation method: The system dynamically adjusts based on the real-time network threat level (input from an external threat intelligence system) and the current link load. During periods of high attack activity or on critical links under high load, adopt more stringent (smaller) measures. Values ​​should be set to enhance privacy protection; restrictions should be appropriately relaxed on trusted networks or during periods of low load. To improve the utility of feature data and support more refined anomaly detection.

[0030] Noise injection for feature correlation awareness: To avoid the potential disruption of intrinsic correlations between features by adding independent and identically distributed noise to all feature dimensions (which is crucial for HSIC analysis), this approach introduces noise in the covariance structure.

[0031] Implementation: First, estimate the covariance matrix of the eigenvectors on clean data (or a historical reference window). Then, multivariate Gaussian noise with a similar correlation structure to the original data is generated: noise vector ,in Based on the current RDP formula and dynamic Calculation

[0032] In this way, the added noise preserves the correlation profile between features, making the noisy data more conducive to subsequent independence analysis.

[0033] 2.2 HSIC Independence Calculation and Adaptive Kernel Selection for Noise Compensation Applying standard HSIC directly to noisy data can underestimate the true dependencies between features due to noise interference. This solution proposes a noise-compensated HSIC computation framework and integrates adaptive kernel function selection to improve the accuracy and efficiency of real-time computation.

[0034] Kernel matrix calculation for noise compensation: For continuous features, noise variance is explicitly considered when calculating the Gaussian kernel.

[0035] Implementation method: The modified Gaussian kernel function is as follows: ,in It is the bandwidth parameter of the feature itself. This is the noise variance added by the RDP mechanism on this feature dimension. This correction offsets the reduction in "spurious" similarity between samples caused by noise, making the HSIC value closer to the result calculated on noise-free data. Adaptive kernel function and parameter selection: To balance real-time performance and accuracy, the system adopts a lightweight online strategy instead of time-consuming cross-validation.

[0036] Implementation method: 1. Kernel type selection: For continuous features, the modified Gaussian kernel is used by default; for discrete / classification features, a dedicated kernel function based on Hamming distance is used.

[0037] 2. Bandwidth parameters Online estimation: For Gaussian kernels, a robust estimate of the median absolute deviation (MAD) of eigenvalues ​​within the current time window is used as the bandwidth. The initial value is set, and fine-tuning is allowed within a small range based on the stability of the calculated HSIC value.

[0038] Parallelized HSIC computation: To meet real-time requirements, the computation of the HSIC matrix is ​​decomposed into multiple kernel function computation blocks that can be executed in parallel. By utilizing CPU vectorized instructions or GPU acceleration, the independence scores between feature pairs are computed synchronously.

[0039] 2.3 Dynamic Feature Selection and Low-Dimensional Space Construction Based on HSIC Scores The goal of this step is to quickly select a subset of low-dimensional features from the original high-dimensional features that can retain anomaly detection information to the greatest extent while meeting the requirements of low-latency processing.

[0040] Dynamic feature selection strategy: 1. Independence Score Ranking: Calculate the HSIC values ​​between all feature pairs and convert them into the "average independence score" or "minimum dependency score" for each feature.

[0041] 2. Multi-objective threshold decision-making: It's not simply about setting a fixed threshold. The screening decision considers multiple factors simultaneously: HSIC Independence Ranking: Features with high independence (i.e., low redundancy with other features) are prioritized. Real-time Cost of Feature Extraction: A measured extraction time cost is associated with each feature. When independence is similar, features with faster computation are prioritized.

[0042] Privacy Budget Consumption Memory: Features that consume less privacy budget during the RDP phase (i.e., have less noise added) are given priority in selection.

[0043] Real-time subset generation: Based on the comprehensive score above, the top K features are dynamically selected (e.g., 5 dimensions are selected from the original 10 dimensions). The value of K can be fine-tuned according to the current system load; when the load is high, K can be appropriately reduced to reduce the amount of computation.

[0044] Privacy-preserving low-dimensional space projection: Use Random Projection instead of PCA. Since the Random Projection (such as Johnson-Lindenstrauss transformation) itself has the differential privacy property, it can form combined privacy protection with the preposed RDP mechanism, and its calculation speed is extremely fast, making it suitable for real-time scenarios.

[0045] The projection matrix W is randomly generated and fixed. The selected K-dimensional feature vector x is projected into a lower M-dimensional space (M < K) through , where is a small additional noise that may be added to meet the final output privacy requirements. This low-dimensional space z is used as the input for subsequent pattern learning.

[0046] 2.4 Real-time normal pattern learning and initial anomaly judgment based on Gaussian Mixture Model (GMM) In the constructed low-dimensional privacy-protected feature space z, the system learns the typical patterns of normal data packet transmission and realizes a preliminary and rapid judgment of the existence of anomalies.

[0047] Lightweight online GMM learning: Use historical normal traffic data to offline train a GMM with C components through the Expectation-Maximization (EM) algorithm to describe the multivariate distribution of normal traffic:

[0048] During online operation, the parameters of this GMM model are fixed and not updated online to ensure the judgment speed.

[0049] Real-time likelihood calculation and threshold judgment: For each arriving data packet, after obtaining its low-dimensional representation through the above steps, the system immediately calculates its log-likelihood value under this GMM:

[0050] The system maintains a dynamic likelihood threshold , which is dynamically updated based on the quantile (such as the 5% quantile) of the likelihood values of normal traffic in the recent period (such as the past 1 second).

[0051] Real-time judgment: If , then this data packet is preliminarily marked as "suspicious anomaly", and its features and context information are quickly transmitted to the subsequent "Step 4: Activate out-of-subspace distribution detection" for precise analysis and positioning. If the likelihood value is normal, the process ends quickly, greatly reducing the overhead of subsequent complex calculations.

[0052] Within the constructed 3D privacy-preserving feature space, the system uses an improved Gaussian Mixture Model (GMM) to learn typical patterns of normal data packet transmission. This application makes three targeted improvements to the traditional GMM to enhance model adaptability, expressive power, and detection accuracy: 2.4.1 Online Incremental Learning Mechanism The system adopts online expected maximization (OPM) The algorithm supports dynamic updates of model parameters. When a new data packet stream arrives, the system does not retrain the entire model, but instead updates it incrementally based on the data within a sliding window. Parameters. In the specific implementation, the weights, mean, and covariance matrix of each Gaussian distribution are progressively adjusted according to the probability of new samples, ensuring that the model can adapt to the time-varying characteristics of network traffic, while avoiding model lag caused by too much historical data.

[0053] 2.4.2 Hierarchical Hybrid Model Structure To enhance the model's expressive power, the system introduces a hierarchical GMM structure: the first layer is a coarse-grained clustering layer that identifies the transmission patterns of major categories (such as web browsing, video streaming, and file downloads); the second layer is a fine-grained sub-model that further decomposes each major category into multiple sub-patterns (such as real-time conferencing and on-demand streaming media in video streaming). This structure can capture both macro-level patterns and distinguish micro-level differences, improving the ability to identify complex and abnormal patterns.

[0054] 2.4.3 Probability calibration of outlier scores Traditional Geometric Analyses (GMMs) rely on likelihood probability thresholds to identify anomalies, which are susceptible to data distribution shifts. This application introduces a probability calibration mechanism, recalibrating the likelihood probability using historical anomaly samples and converting it into a calibrated anomaly confidence level. Specifically, it employs... The proposed method uses a logistic regression model to map the likelihood probability output by the GMM, generating anomaly scores with clear probabilistic interpretations, which significantly reduces the false alarm rate.

[0055] Step 3: Design and Implementation of Adaptive Threshold Detection Mechanism 3.1 Construction of a Multi-Tag Ranking Framework A multi-label ranking framework is constructed to stratify and sort network anomaly types according to their severity and scope of impact.

[0056] 3.2 Implicit Class Significance Modeling We design an implicit class saliency model to extract salient features of anomalous torsion patterns by learning intra-class similarity and inter-class differences.

[0057] 3.3 Dynamic Threshold Calculation Strategy By combining network traffic fluctuation characteristics and historical anomaly detection results, an adaptive threshold calculation strategy is designed, and a sensitivity feedback mechanism based on anomaly priority is used to achieve dynamic adjustment of the threshold.

[0058] Steps 3.1-3.3 are illustrated with examples: Step 3.1 Building a Multi-Tag Ranking Framework The system constructs a multi-layered labeling system for network anomalies based on their impact and urgency. The first layer categorizes anomalies by severity: minor anomalies (Level-1), moderate anomalies (Level-2), severe anomalies (Level-3), and critical anomalies (Level-4). The second layer categorizes anomaly types: path deviation anomalies, latency spikes, throughput drops, and security threats. For example, a "severe path deviation anomaly" would be labeled as [Level-3, Path Deviation]. The system uses a pairwise ranking method to train a ranking model, learning the relative importance of different anomaly labels. Specifically, if historical data shows that "critical security threat anomalies" have a higher processing priority than "moderate latency spikes," the system will learn the ranking relationship R(Level-4 security threat) > R(Level-2 latency spike). This ranking framework enables the system to process and respond to multiple anomalies in order of importance when they are detected.

[0059] 3.1.1 Feedback and Application of Ranking Results The anomaly priority list output by the multi-label ranking model is a key basis for adjusting system resource scheduling and detection strategies. Feedback to the threshold module (step 3.3): As mentioned above, this directly affects the sensitivity settings of various abnormal dynamic thresholds. .

[0060] Feedback to the feature space (step 4.1): High-priority anomaly types will receive more "memory" retrieval weights and may trigger the construction or enhancement of a dedicated activation subspace for them.

[0061] Guidance on evidence collection and reconstruction resource allocation: In steps 4.3 (abnormal link location) and 5 (path reconstruction), the system will prioritize handling high-level anomalies to ensure that critical issues receive an immediate response.

[0062] Step 3.2 Implicit Class Significance Modeling For the specific anomaly type of data twist link, the system needs to learn its saliency in the feature space. Assuming that normal transmission data forms tight clusters in the feature space, while twist link anomalies deviate from these normal clusters, the system uses Support Vector Data Description (SVDD) to build boundary descriptions of normal data and calculates the distance of each observed sample to the boundary as an anomaly score. For example, the feature vectors of normal web traffic might cluster in... Nearby, with a radius of Within a spherical region. When a feature vector is detected. At that time, its distance from the center of the normal region is The radius is much larger than the normal radius r, thus achieving a high anomaly significance score. The system further learns intra-class similarity, i.e., the feature similarity between anomaly samples of the same type, and inter-class dissimilarity, i.e., the feature distinguishability between different anomaly types, thereby accurately identifying the unique pattern of torsional link anomalies.

[0063] Traditional SVDD is a static model, which is difficult to adapt to the time-varying nature of network traffic patterns and conceptual drift.

[0064] (1) Incremental boundary update: The online learning algorithm based on KKT conditions is adopted. When new normal samples arrive, only some support vectors are added, deleted and adjusted to achieve incremental update of the model boundary without full retraining.

[0065] (2) Drift-aware slack variables: An "age" attribute is introduced for each support vector, and its influence decays over time. Simultaneously, the slack variable penalty coefficient in the objective function is dynamically adjusted. Reduce during periods of network stability To increase model compactness and improve performance during periods of volatility To enhance the robustness of the model.

[0066] (3) Linkage with the activation subspace: The SVDD model is trained and inferred in the current activation subspace constructed in step 4.1. This subspace has been optimized for recent abnormal patterns, thereby significantly improving the "saliency" modeling ability of SVDD in distinguishing specific abnormal traffic from normal traffic.

[0067] The goal of this step is to establish an implicit class saliency model that can adapt to dynamic changes in network traffic, in order to accurately identify the saliency of data twisting link anomalies in the feature space. Traditional SVDD models are static models and are difficult to adapt to the time-varying nature of network traffic and concept drift. To address this, this solution makes three targeted improvements to traditional SVDD and forms a complete modeling and inference process.

[0068] 3.2.1 Model Initialization and Historical Memory Loading When the system starts up or detects a new anomaly, it initializes the implicit class saliency model.

[0069] Input: The set of historical normal samples in the current active subspace (from step 4.1), and historical anomalous samples of the same type retrieved from the "Anomalous Pattern Memory" (if they exist).

[0070] Initial Model Training: Using the aforementioned normal sample set, an initial SVDD model is trained. Its goal is to find a minimum hypersphere in the feature space that covers most of the normal data. The objective function is:

[0071] in, The radius of the hypersphere, For the center of the ball, For slack variables This is the penalty coefficient.

[0072] Historical knowledge injection: If there are historical anomalies, the system will analyze the distribution of these samples outside the initial hypersphere boundary and make preliminary adjustments to the sphere's center. Position or radius This allows the model to possess certain prior knowledge of anomaly differentiation during the initialization phase.

[0073] 3.2.2 Online Model Update Based on Incremental Learning To adapt to real-time changes in network traffic, the model is updated using an online incremental learning algorithm based on KKT conditions, rather than periodic full retraining.

[0074] 1. New Sample Arrival Processing: When new normal data samples arrive... Upon arrival, the system determines its relationship with the current hypersphere.

[0075] 2. Dynamic maintenance of support vector sets: If If a vector is located inside the sphere and is not a support vector, only its "age" attribute is updated; the model parameters remain unchanged. If a new sample becomes a new boundary support vector (i.e., satisfies a specific constraint in the KKT conditions), it is added to the support vector set. If an existing support vector no longer satisfies the support vector conditions due to the addition of a new sample, it is removed from the support vector set.

[0076] 3. Incremental adjustment of model parameters: Based on the updated support vector set, the center of the hypersphere is incrementally adjusted through analytical update formulas or fast optimization steps. and radius This enables the smooth evolution of the model boundary.

[0077] 3.2.3 Relaxation Strategies and Saliency Calculation for Drift Perception To address concept drift and quantify anomalous significance, the system introduces a dynamic relaxation strategy and a distance-based significance measure.

[0078] 1. Dynamic penalty coefficient (C) adjustment: penalty coefficient No longer fixed. The system monitors the proportion of recent normal samples that "go out of bounds." During periods of network stability, the proportion is reduced. This value makes the model more compact and more sensitive to slight deviations; it also improves performance during periods of network fluctuation or when potential distribution drift is detected. The value allows for greater relaxation, enhances model robustness, and avoids misclassifying normal fluctuations as anomalies.

[0079] 2. Support Vector Influence Decay: Each support vector is associated with an "age" attribute, whose influence decays exponentially with age. This ensures that the model better reflects recent traffic patterns rather than being dominated by outdated historical patterns.

[0080] 3. Calculation of abnormal significance score: For the sample to be tested Its implicit class saliency score Defined as the distance from the sample to the center of the current hypersphere. The normalized ratio of the squared distance to the square of the current sphere radius: The larger the value, the deeper the deviation of the sample from the normal pattern, and the stronger its significance as an anomaly (especially the data twisting link anomaly).

[0081] 3.2.4 Collaboration with other system modules This implicit class saliency model does not operate in isolation but rather collaborates deeply with other key modules of the system. It is linked to the activation subspace: the model is always trained and inferred within the current activation subspace constructed in step 4.1. This subspace has been feature-optimized for recent traffic patterns (especially suspicious anomaly patterns), significantly improving the SVDD model's "focusing" ability and discrimination efficiency in distinguishing specific anomalies from normal traffic. It provides input for multi-label ranking: the calculated saliency score. It is one of the important input features of the multi-label ranking framework in step 3.1, used to help determine the severity level and type confidence of anomalies.

[0082] Feedback to dynamic threshold: The trend of abnormal pattern changes detected by the model (such as the overall increase in the significance score of a certain type of abnormality) can be used as a feedback signal to affect the sensitivity coefficient of relevant indicators in the dynamic threshold calculation strategy in step 3.3.

[0083] Process Summary: Implicit class saliency modeling is a continuous closed-loop process of "initialization -> online incremental update -> dynamic saliency calculation -> collaborative feedback". Through an improved, online-updable SVDD model, it dynamically learns the boundaries of normal patterns in an optimized feature subspace and quantifies the degree to which each sample deviates from these boundaries. This provides accurate and adaptive saliency measures for anomalies such as data reversal links, supporting subsequent prioritization and precise handling.

[0084] Step 3.3 Dynamic Threshold Calculation Strategy Traditional fixed-threshold detection methods are prone to false positives and false negatives when faced with dynamic changes in network traffic. This system employs an adaptive threshold strategy based on time series prediction and multi-indicator fusion, which can dynamically adjust the detection sensitivity according to the network status, thereby improving the accuracy of anomaly identification.

[0085] 3.3.1 Multi-indicator time series forecasting The system simultaneously monitors multiple key network performance indicators, such as inter-packet latency, throughput, packet loss rate, and path hop count. An independent time-series forecasting model is established for each indicator, and an Autoregressive Integral Moving Average (ARIMA) model is used to predict future values, obtaining the predicted values ​​for each indicator at the next time step. and its forecast uncertainty .

[0086] 3.3.2 Basic Calculation of Dynamic Threshold For the Each indicator's dynamic threshold consists of three parts: the prediction benchmark, the volatility tolerance, and the trend correction term.

[0087] in: The baseline value for the index predicted by the ARIMA model; The confidence level coefficient is usually set based on the quantiles of the normal distribution (e.g., 95% confidence level corresponds to...). ; The standard deviation is the forecast, reflecting the uncertainty of the forecast. For trend correction items, This is a trend sensitivity coefficient used to capture accelerated changes in indicators.

[0088] 3.3.3 Adaptive Weight Adjustment Mechanism (Highlighting Key Innovations) The importance of different network metrics in indicating link anomalies is not fixed and changes dynamically with network status. Therefore, this system introduces an adaptive weight adjustment mechanism based on historical detection feedback to dynamically optimize the weight of each metric in the comprehensive decision-making process.

[0089] 1. Evaluation of indicator effectiveness The system maintains a sliding time window (e.g., the past hour) and statistically analyzes the performance of each metric within that window.

[0090] in: Number of true positive results (correct alert) Number of false positives (false alarms) Number of false negatives (underreporting) Smoothing constant to prevent the denominator from being zero. Reflects the first The effectiveness of an indicator in the current network environment is measured by its value; a higher value indicates that the indicator is more reliable.

[0091] 2. Dynamic weight calculation Based on the effectiveness scores of each indicator, their normalized weights are calculated using the softmax function:

[0092] in The temperature parameter controls the sharpness of the weight distribution. Indicators with high effectiveness will receive higher weights, thus carrying greater weight in the overall decision-making process.

[0093] 3. Comprehensive dynamic threshold The final overall threshold used to trigger an anomaly alarm is determined by the weighted deviation of the thresholds for each indicator:

[0094] in This indicates that only the portion of the indicator value exceeding its individual threshold is considered. Exceeding the preset global threshold When this occurs, the system determines it to be in an abnormal state and triggers the subsequent torsion link identification process.

[0095] 3.3.4 Mechanism Advantages This adaptive weight adjustment mechanism enables the system to: Environmental Adaptation: Based on the actual network operating status and detection history, automatically increase the weight of reliability indicators and reduce the impact of noise or failure indicators.

[0096] Reduce false alarms and false negatives: By dynamically integrating information from multiple indicators, misjudgments caused by fluctuations in a single indicator can be avoided.

[0097] Continuous optimization: The weights are continuously updated based on feedback within the sliding window, enabling the system to learn and adapt over a long period of time.

[0098] Step 4: Activate subspace distribution external detection and accurate identification of abnormal links 4.1 Activation of Subspace Construction For the detected abnormal data packet stream, a feature activation subspace is constructed, and the combination of feature dimensions most sensitive to abnormal patterns is extracted.

[0099] When the system detects abnormal data packet flows, it needs to construct a feature activation subspace specifically for these abnormal samples for more precise analysis. The core idea of ​​the activation subspace is to extract the combination of feature dimensions that are most sensitive and discriminative to the current abnormal pattern from the original high-dimensional feature space. In specific implementation, the system first evaluates the feature importance of the detected abnormal data packet flows, determining the activation level by calculating the contribution of each feature dimension in distinguishing between normal and abnormal traffic. For example, when an abnormality is detected in the transmission path of a data flow, the system will find that the features "path hop count change" and "abnormal TTL value fluctuation" show abnormally high activation levels in this abnormal sample, while the activation level of the "data packet size" feature is relatively low.

[0100] The system employs an attention mechanism to construct the activation subspace, assigning different weight coefficients to each feature dimension. Higher weights indicate greater sensitivity of the feature to current anomaly patterns. During the actual construction process, the system creates a low-dimensional subspace projection, projecting data from the original feature space onto this subspace most sensitive to anomaly patterns. This projection not only preserves the most crucial information for anomaly detection but also significantly reduces the computational complexity of subsequent processing. The dimension of the activation subspace is typically much smaller than that of the original feature space, yet it contains the core information needed to identify specific anomaly patterns.

[0101] 4.1.1 Dynamic Optimization Based on Anomaly Priority and Pattern Memory The system maintains an "abnormal pattern memory bank" that stores various abnormal samples that have been detected and confirmed in history, along with their feature activation intensity vectors A.

[0102] 1. Initialization guidance: When it is necessary to build an activation subspace for a new round of initial screening anomalies, the system first retrieves historical anomaly patterns of the same type or higher priority type from the memory.

[0103] 2. Feature weight warm-up: The feature activation intensity vectors of these historical patterns are fused together and used as the initial attention weights to guide the Fisher discriminant analysis to focus more on the feature dimensions related to historical anomalies.

[0104] 3. Subspace Iterative Update: Once a new detection result is confirmed, its pattern will be abstracted and updated to the memory. When the same anomaly recurs, the system can quickly call or fine-tune the existing dedicated subspace, rather than building it from scratch, greatly improving response speed and accuracy.

[0105] 4.2 Identification of out-of-distribution samples We design an out-of-distribution sample identification method based on local density estimation to accurately locate out-of-distribution abnormal data points.

[0106] Within the constructed activation subspace, the system needs to accurately identify anomalous samples that deviate from the normal distribution pattern. The core of out-of-distribution detection technology is to establish an accurate description of the normal data distribution and then measure the degree of deviation of newly observed samples from this normal distribution. The system employs a method based on local density estimation to achieve this goal. This method can capture the local characteristics of the data distribution and is more sensitive to anomalous samples.

[0107] In practice, the system first trains a density estimation model in the activation subspace using historical normal data. This model describes the distribution pattern of normal data in the activation subspace. When a new data sample arrives, the system calculates the local density value of the sample at its current location and compares it with the density threshold of the normal distribution. If the density value of a region containing a data sample is significantly lower than the normal density threshold, it is identified as an out-of-distribution sample, i.e., a potential anomalous sample.

[0108] The system also incorporates multi-scale density estimation techniques to calculate sample density values ​​at different spatial scales, thereby improving detection robustness. For each observed sample, the system considers not only its density within its local neighborhood but also its density distribution over a larger area. This multi-scale analysis method effectively distinguishes between genuine anomalous samples and accidental deviations caused by data noise, improving the accuracy and reliability of anomaly detection.

[0109] 4.3 Abnormal Link Feature Extraction and Localization By combining network topology information, the specific network links and nodes through which abnormal data packets pass can be located, and the network link segments where the torsion occurred can be accurately identified.

[0110] Once the system identifies out-of-distribution anomalous data samples, the next step is to extract specific anomalous link features from these samples and combine them with network topology information to pinpoint the exact location of the anomaly. This process involves multi-level feature analysis and topology mapping techniques.

[0111] First, the system performs in-depth feature analysis on the identified abnormal samples, extracting their path and timing features. Path features include the sequence of network nodes the data packet traverses, the processing latency of each node, and the link characteristics between nodes. The system reconstructs the data packet's transmission path by analyzing changes in the TTL value, IP option fields, and timestamp information in the data packet header. Timing features include the time interval between data packets arriving at each node, changes in transmission rate, and time deviations from normal transmission patterns.

[0112] The system establishes a comprehensive network topology database containing detailed information about each node and link in the network, such as bandwidth capacity, historical performance data, geographical location, and device type. By matching and correlating the extracted abnormal sample path features with the topology database, the system can accurately locate the specific network links and nodes traversed by abnormal data packets. When the system detects a significant deviation between the transmission path of a data packet and the expected optimal path, it further analyzes the specific location and cause of the deviation.

[0113] During the construction of the abnormal transmission path graph, the system not only records the abnormal paths themselves but also compares and analyzes the normal transmission paths between the same source-destination pairs. Through difference analysis, it accurately identifies network link segments and critical nodes that have experienced tortuosity. The system uses graph theory algorithms to analyze the structural characteristics of the path graph, identify critical nodes and bottleneck links in abnormal paths, and provide accurate location information for subsequent path reconstruction.

[0114] Step 4.4 Collaborative Closed Loop of Detection, Ranking, and Feedback Step 3 (adaptive threshold detection), Step 4 (activation subspace fine-tuning), and the multi-label ranking framework together constitute a dynamic, self-optimizing anomaly detection system. Their collaborative workflow is as follows: 1. Data Flow and Triggers: Real-time traffic data undergoes initial screening using dynamic thresholds in step 3.3, generating a set of "suspicious abnormal events". .

[0115] 2. Refined Analysis and Space Optimization: E_suspect triggers step 4.1, where the system combines the "abnormal mode memory" and current data to construct or optimize a highly targeted activation subspace S_active.

[0116] 3. Precise localization and classification: In the S_active subspace, out-of-distribution detection is performed using the improved incremental SVDD (step 3.2) and density estimation method (step 4.2) to accurately authenticate events in E_suspect, remove false alarms, and output the confirmed anomaly set E_confirmed and its preliminary type label.

[0117] 4. Priority Decision: E_confirmed, its labels, and contextual features are fed into the multi-label ranking framework in step 3.1 to calculate a processing priority sequence P_rank that integrates severity, urgency, and resource impact.

[0118] 5. Closed-loop feedback: The priority sequence P_rank and the new patterns learned in this detection will generate two types of feedback: Short-term parameter adjustments: immediately affect the dynamic threshold sensitivity for different anomaly types in step 3.3.

[0119] Long-term knowledge accumulation: The abstracted abnormal pattern features, activation subspace configuration parameters and ranking relationships are stored in the "abnormal pattern memory" to optimize future detection loops (step 4.1).

[0120] This closed-loop mechanism ensures that the system can learn from historical experience, making the coarse screening threshold more accurate, the fine inspection feature space more discriminative, and the resource allocation more efficient, thus achieving a continuous improvement in detection accuracy and system adaptability.

[0121] Step 5: Design and Optimization of Intelligent Path Reconstruction Algorithm 5.1 Multi-constraint path optimization model A mathematical optimization model is established, taking into account multi-dimensional constraints such as bandwidth requirements, latency sensitivity, and reliability requirements.

[0122] Based on the identified tortuous link locations and characteristics from the preceding steps, the system needs to establish a path optimization mathematical model that comprehensively considers multiple constraints. This model is not a simple shortest path problem, but a complex multi-objective optimization problem that requires simultaneous consideration of constraints from multiple dimensions, including bandwidth requirements, latency sensitivity, reliability requirements, and cost control.

[0123] Regarding bandwidth constraints, the system needs to ensure that the selected new path can meet the bandwidth requirements of the current data stream without causing congestion on other links in the network. The system maintains real-time network link status information, including the current utilization, available bandwidth, and predicted future load for each link. Latency constraints consider the latency sensitivity of different applications; for example, real-time video calls have much higher latency requirements than file downloads. The system sets corresponding latency upper limits based on the application type of the data stream.

[0124] Reliability constraints consider the stability and fault tolerance of network paths. The system analyzes factors such as the historical failure rate, equipment redundancy, and maintenance plans of each node and link in the candidate paths to ensure that the selected paths have sufficient reliability. Cost constraints consider the cost differences in using different network resources, including bandwidth leasing fees, equipment usage costs, and operation and maintenance costs. The system integrates these multi-dimensional constraints into a unified optimization objective function, balancing the importance of different constraints through weight adjustments.

[0125] 5.2 Heuristic Pathfinding Algorithm An improved A* search algorithm was designed, combining genetic algorithms and simulated annealing strategies to efficiently search for a set of feasible alternative paths.

[0126] Since multi-constraint path optimization is an NP-hard problem, traditional exact algorithms often fail to find the optimal solution in a reasonable time in large-scale networks. Therefore, this system employs an improved heuristic search algorithm to efficiently solve this optimization problem. This algorithm combines the goal-oriented nature of the A* search algorithm, the global search capability of the genetic algorithm, and the local optimization characteristics of the simulated annealing strategy.

[0127] The improved A* algorithm provides the basic framework and direction of the search. The system designs a heuristic function for each network node, which estimates the optimal path cost from the current node to the target node. This heuristic function considers not only geographical distance and hop count, but also bandwidth availability, historical performance data, and current network load. During the search, the algorithm prioritizes exploring nodes with smaller heuristic function values, thus converging quickly towards the target.

[0128] The genetic algorithm component is responsible for maintaining and evolving a population of candidate paths. Each candidate path is encoded as a gene sequence, representing the sequence of network nodes traversed from the source node to the target node. The system employs specialized crossover and mutation operations to generate new candidate paths. Crossover combines the characteristics of two excellent paths, while mutation introduces randomness to explore new path possibilities. A fitness function comprehensively evaluates the performance of each path across multiple constraints; paths with higher fitness have a greater probability of being selected for the next generation of the population.

[0129] Simulated annealing is integrated into the entire search process to control the balance between exploration and development. In the early stages of the search, a higher "temperature" parameter is used, allowing for some temporarily poor solutions to avoid getting trapped in local optima. As the search progresses, the temperature gradually decreases, and the algorithm focuses more on a refined search around the current optimal solution. This strategy ensures that the algorithm can perform global exploration while converging to a high-quality solution within a finite time.

[0130] 5.3 Dynamic Path Deployment and Verification The optimized path is distributed to relevant network devices through the network control interface, and the performance of the new path is monitored.

[0131] Once the path search algorithm finds the optimal alternative path, the system needs to actually deploy this new path into the network and establish a robust verification mechanism to ensure that the path reconstruction effect meets expectations. This process involves multiple levels of operation and verification.

[0132] During the path deployment phase, the system communicates with relevant network devices, including routers, switches, and SDN controllers, through the network control interface. For traditional distributed routing networks, the system issues new routing table entries or policy routing rules to the relevant routers, guiding data packets to be forwarded along the new path. For SDN networks, the system communicates directly with the SDN controller, issuing new flow table entries to achieve path switching. During deployment, the system adopts a gradual switching strategy, first directing a small amount of test traffic to the new path to verify its basic availability, then gradually increasing the traffic proportion, and finally completing a full switchover.

[0133] The verification mechanism includes multiple layers of performance monitoring and quality assessment. First, there's basic connectivity verification: the system sends probe packets along the new path to confirm end-to-end connectivity and basic forwarding functionality. Next, there's performance metric verification: the system continuously monitors key performance indicators of the new path, including end-to-end latency, packet loss rate, jitter, and throughput, and compares them with expected targets. If performance metrics are found to be unsatisfactory, the system automatically triggers a rollback mechanism, either reverting to the original path or initiating a new round of path searching.

[0134] The system also establishes a long-term path performance tracking mechanism to continuously monitor the stability and reliability of deployed paths. By collecting and analyzing historical path performance data, the system can continuously optimize path selection strategies and constraint parameters, improving the quality and efficiency of future path reconstruction. When the network environment changes or new anomalies are detected, the system will reassess the applicability of the current path and, if necessary, initiate a new path optimization process to ensure that network transmission always remains in optimal condition.

[0135] Step 6: Implementation of Zero-Copy Memory Management and Parallel Processing Architecture 6.1 Zero-copy memory architecture design By using direct memory access (DMA) and memory mapping techniques, the copying of data packets between kernel space and user space is reduced.

[0136] 6.2 Packet Processing Pipeline Construction This application proposes a multi-stage pipelined processing architecture that decomposes packet processing into independent stages, enabling parallel execution of each stage. The architecture integrates core steps such as Rényi differential privacy noise injection, HSIC feature selection, and GMM anomaly detection into pipeline stages, achieving high-speed packet processing and real-time anomaly response.

[0137] 6.2.1 Pipeline Stage Division and Parallel Execution Mechanism The system divides the data processing flow into the following stages, which are connected by a lock-free circular buffer to achieve parallel execution: 1. Data Acquisition and Preprocessing Stage: Real-time capture of data packets and extraction of raw features.

[0138] 2. RDP noise injection stage: Add noise that satisfies Rényi differential privacy to sensitive features.

[0139] 3. HSIC Feature Selection Stage: Calculate feature independence on privacy-preserving data and screen key features.

[0140] 4. GMM Anomaly Detection Phase: Use an improved GMM for pattern matching and anomaly scoring.

[0141] 5. Decision and Response Phase: Trigger alarms or path reconstruction based on anomaly confidence levels.

[0142] Each stage runs on an independent CPU core or hardware acceleration unit, achieving high throughput through pipelined parallelism.

[0143] 6.2.2 Inter-stage data flow optimization and cache design To reduce data copying and synchronization overhead, the system employs a zero-copy data transfer mechanism, sharing a memory-mapped data buffer between each stage. The output of each stage is written to the input buffer of the next stage. The buffer is designed as a multi-producer-single-consumer queue, supporting batch processing and asynchronous notifications to ensure low latency and high reliability of data flow.

[0144] 6.2.3 Rapid Feedback Mechanism for Anomaly Detection Results The pipeline architecture integrates a closed-loop feedback control mechanism: when an anomaly is detected in the GMM stage, the decision stage not only triggers an alarm, but also feeds back the abnormal sample and context information to the RDP and HSIC stages in real time, dynamically adjusting the noise parameters and feature selection weights to achieve adaptive optimization of the detection strategy.

[0145] For example: This pipeline is implemented on an FPGA or smart network interface card (NIC), with each stage corresponding to a hardware processing unit. After data packets enter the NIC, they are processed sequentially through each stage without CPU intervention, achieving microsecond-level anomaly detection and response.

[0146] 6.3 NUMA-aware task scheduling Design a NUMA-aware task scheduling strategy to optimize the affinity between processing cores and memory nodes.

[0147] Step 7: Building a Hardware-Accelerated Pattern Matching Engine 7.1 Selection of Programmable Hardware Acceleration Platform Choose a suitable hardware acceleration platform (such as FPGA, smart network card, etc.) and design a hardware acceleration architecture.

[0148] 7.2 Parallel Matching Algorithm Implementation Implement a highly parallel pattern matching algorithm on a hardware platform to support high-speed pattern matching of multiple data streams simultaneously.

[0149] 7.3 Hardware-Software Co-processing Mechanism Design a collaborative mechanism between the hardware acceleration unit and the software processing system to achieve the best performance balance.

[0150] Please refer to Figure 2 The method in this embodiment may include: standardizing data packet information to obtain standardized data, including: Step S201: Extract data packet information based on the time window segmentation method to obtain time sequence features; Step S202: Perform normalization processing based on time series characteristics to obtain standardized data.

[0151] Specifically, time-window segmentation and statistical analysis methods are used to extract the temporal characteristics of data packets, including inter-packet delay distribution, burstiness indicators, and periodic patterns, generating temporal feature descriptors. The extracted features are then normalized, and a standardized set of data packet feature vectors is generated using a feature fusion algorithm.

[0152] Please refer to Figure 3 The method in this embodiment may include: processing standardized data and a preset hybrid model in a privacy-preserving feature space to obtain anomaly detection information, including: Step S301: Calculate the sensitivity of the standardized data to obtain the first sensitivity; Step S302: Calculate the HSIC independence based on the first sensitivity for noise compensation, and obtain the first calculation result; Step S303: Based on the first calculation result and the preset hybrid model, process the data in the privacy-preserving feature space to obtain anomaly detection information.

[0153] Specifically, dynamic global sensitivity Calculate: Global sensitivity in traditional dynamic programming The fixed worst-case upper bound is often too conservative in real-world network traffic, leading to excessive noise. This solution calculates the sensitivity of feature dimensions in real time. Directly applying standard HSIC to noisy data underestimates the true dependencies between features due to noise interference. This solution proposes a noise-compensated HSIC calculation framework and integrates adaptive kernel function selection to improve the accuracy and efficiency of real-time calculation.

[0154] Please refer to Figure 4 The method in this embodiment may include: processing the first calculation result and a preset hybrid model in a privacy-preserving feature space to obtain anomaly detection information, including: Step S401: Calculate the first calculation result based on the dynamic feature filtering strategy to obtain the second calculation result; Step S402: Based on the second calculation result and the preset hybrid model, process the data in the privacy-preserving feature space to obtain anomaly detection information.

[0155] Specifically, dynamic feature selection strategy: 1. Independence Score Ranking: Calculate the HSIC values ​​between all feature pairs and convert them into the "average independence score" or "minimum dependency score" for each feature.

[0156] 2. Multi-objective threshold decision-making: It's not simply about setting a fixed threshold. The screening decision considers multiple factors simultaneously: HSIC Independence Ranking: Features with high independence (i.e., low redundancy with other features) are prioritized. Real-time Cost of Feature Extraction: A measured extraction time cost is associated with each feature. When independence is similar, features with faster computation are prioritized.

[0157] Privacy Budget Consumption Memory: Features that consume less privacy budget during the RDP phase (i.e., have less noise added) are given priority in selection.

[0158] 3. Real-time subset generation: Based on the comprehensive score above, dynamically select the top K features (e.g., select 5 dimensions from the original 10 dimensions). The value of K can be fine-tuned according to the current system load; when the load is high, K can be appropriately reduced to reduce the amount of computation.

[0159] Privacy-preserving low-dimensional space projection: Use Random Projection instead of PCA. Since the differential privacy property inherent in Random Projection (such as Johnson-Lindenstrauss transformation) can form combined privacy protection with the pre-existing RDP mechanism, and its computing speed is extremely fast, it is suitable for real-time scenarios.

[0160] The projection matrix W is randomly generated and fixed. The selected K-dimensional feature vector x is projected into a lower-dimensional M-dimensional space (M < K) through , where is a tiny additional noise that may be added to meet the final output privacy requirements. This low-dimensional space z is used as the input for subsequent pattern learning.

[0161] Please refer to Figure 5 , the method of this embodiment may include: detecting the abnormal discrimination information based on an adaptive threshold strategy to obtain a first detection result, including: Step S501: Perform multi-index time series prediction on the abnormal discrimination information to obtain a first performance index; Step S502: Perform dynamic threshold base calculation on the first performance index to obtain a first threshold; Step S503: Process the first threshold based on an adaptive weight adjustment mechanism to obtain a first detection result.

[0162] Specifically, for multi-index time series prediction: The system simultaneously monitors multiple key network performance indicators, such as packet delay, throughput, packet loss rate, path hop count, etc. An independent time series prediction model is established for each indicator, and the autoregressive integrated moving average model (ARIMA) is used to predict future values, obtaining the predicted values of each indicator at the next moment and its prediction uncertainty .

[0163] Dynamic threshold base calculation: For the th indicator, its dynamic threshold consists of three parts: the predicted reference value, the fluctuation tolerance, and the trend correction term:

[0164] Where: is the indicator reference value predicted by the ARIMA model; is the confidence coefficient, usually set according to the normal distribution quantile (such as the 95% confidence level corresponding to ; is the predicted standard deviation, reflecting the prediction uncertainty; <00​​This is a trend sensitivity coefficient used to capture accelerated changes in indicators.

[0165] Adaptive Weight Adjustment Mechanism: The importance of different network indicators in indicating link anomalies is not fixed and changes dynamically with the network status. Therefore, this system introduces an adaptive weight adjustment mechanism based on historical detection feedback to dynamically optimize the weight of each indicator in the comprehensive decision-making process.

[0166] Please refer to Figure 6 The method in this embodiment may include: identifying the first detection result based on the subspace; before obtaining the location information, the method includes: Step S601: Obtain historical anomaly patterns; Step S602: Perform feature activation intensity vector fusion based on historical anomaly patterns to obtain a subspace; Step S603: Identify the first detection result based on the subspace to obtain location information.

[0167] Specifically, initialization guidance: When it is necessary to build an activation subspace for a new round of initial screening anomalies, the system first retrieves historical anomaly patterns of the same or higher priority type from the memory.

[0168] Feature weight warm-up: The feature activation intensity vectors of these historical patterns are fused together and used as initial attention weights to guide the Fisher discriminant analysis to focus more on feature dimensions related to historical anomalies.

[0169] Subspace Iterative Update: Once a new detection result is confirmed, its pattern is abstracted and updated in the memory. When the same anomaly recurs, the system can quickly call or fine-tune the existing dedicated subspace, rather than building it from scratch, greatly improving response speed and accuracy.

[0170] Please refer to Figure 7 The method in this embodiment may include: identifying the first detection result based on a subspace to obtain location information, including: Step S701: Identify the first detection result based on the subspace to obtain the first abnormal sample; Step S702: Locate the first abnormal sample based on the abnormal link to obtain location information.

[0171] Please refer to Figure 8 The method in this embodiment may include: processing location information based on a path search algorithm to obtain a first path, including: Step S801: Obtain the path search algorithm based on the genetic algorithm and simulated annealing strategy; Step S802: Process the location information based on the path search algorithm to obtain the first path.

[0172] Specifically, an improved A* search algorithm was designed, combining genetic algorithms and simulated annealing strategies to efficiently search for a set of feasible alternative paths.

[0173] Since multi-constraint path optimization is an NP-hard problem, traditional exact algorithms often fail to find the optimal solution in a reasonable time in large-scale networks. Therefore, this system employs an improved heuristic search algorithm to efficiently solve this optimization problem. This algorithm combines the goal-oriented nature of the A* search algorithm, the global search capability of the genetic algorithm, and the local optimization characteristics of the simulated annealing strategy.

[0174] The improved A* algorithm provides the basic framework and direction of the search. The system designs a heuristic function for each network node, which estimates the optimal path cost from the current node to the target node. This heuristic function considers not only geographical distance and hop count, but also bandwidth availability, historical performance data, and current network load. During the search, the algorithm prioritizes exploring nodes with smaller heuristic function values, thus converging quickly towards the target.

[0175] The genetic algorithm component is responsible for maintaining and evolving a population of candidate paths. Each candidate path is encoded as a gene sequence, representing the sequence of network nodes traversed from the source node to the target node. The system employs specialized crossover and mutation operations to generate new candidate paths. Crossover combines the characteristics of two excellent paths, while mutation introduces randomness to explore new path possibilities. A fitness function comprehensively evaluates the performance of each path across multiple constraints; paths with higher fitness have a greater probability of being selected for the next generation of the population.

[0176] Simulated annealing is integrated into the entire search process to control the balance between exploration and development. In the early stages of the search, a higher "temperature" parameter is used, allowing for some temporarily poor solutions to avoid getting trapped in local optima. As the search progresses, the temperature gradually decreases, and the algorithm focuses more on a refined search around the current optimal solution. This strategy ensures that the algorithm can perform global exploration while converging to a high-quality solution within a finite time.

[0177] The method in this embodiment may include: The advantages of this solution compared to the prior art are: Significantly improves real-time performance: Through a hardware-accelerated pattern matching engine and zero-copy memory management, it achieves microsecond-level anomaly detection response, which is several orders of magnitude better than traditional methods.

[0178] Resolving the conflict between privacy protection and detection accuracy: Innovatively combining Rényi differential privacy with the HSIC independence criterion, high detection accuracy is maintained while strictly protecting data privacy.

[0179] Significantly reduced computational overhead: Through parallel processing architecture, lock-free data structures, and NUMA-aware scheduling, the system's computational overhead is significantly reduced, and processing efficiency is improved.

[0180] Improving anomaly location accuracy: Activating subspace distributed external detection technology can accurately locate the position of anomaly links and describe their torsion characteristics, providing an accurate basis for path reconstruction.

[0181] Enhanced system adaptability: The adaptive threshold detection mechanism can dynamically adjust the detection parameters according to changes in the network environment, thereby improving the system's environmental adaptability.

[0182] Significant improvements to the GMM model: Through online learning, hierarchical structure, and probability calibration, the traditional GMM has been able to overcome its poor adaptability and high false alarm rate in dynamic network environments.

[0183] Systematic optimization of pipeline architecture: Through stage parallelization, data flow optimization and feedback mechanisms, efficient collaboration between privacy protection, feature selection and anomaly detection is achieved, overcoming the shortcomings of high latency and poor scalability of traditional serial processing architecture.

[0184] The system achieves adaptive evolution: through a closed-loop feedback mechanism, the system can accumulate experience and automatically optimize the front-end coarse screening threshold and fine detection feature space, reducing reliance on expert rules and manual parameter tuning.

[0185] Improved the timeliness of handling high-value anomalies: Through multi-label ranking and resource feedback mechanisms, it ensures that serious threats can be identified more sensitively and analytical resources can be prioritized, thus achieving refined risk management.

[0186] Please refer to Figure 9 The apparatus of this embodiment may include the following structure: Acquisition unit 901 is used to acquire data packet information; Analysis unit 902 is used to standardize data packet information to obtain standardized data; Based on standardized data and a pre-defined hybrid model, anomaly detection information is obtained by processing the data in a privacy-preserving feature space. An adaptive threshold strategy is then used to detect the anomaly detection information to obtain a first detection result. Finally, the first detection result is identified based on a subspace to obtain location information. The processing unit 903 is used to process the location information based on the path search algorithm to obtain the first path.

[0187] Please refer to Figure 10This embodiment of the present application also discloses an electronic device, which includes at least one processor 1001, at least one memory 1002 and a bus 1003 connected to the processor 1001; wherein the processor 1001 and the memory 1002 communicate with each other through the bus 1003; the processor 1001 is used to call program instructions in the memory 1002 to execute the above-mentioned data path-based identification method.

[0188] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for identifying data paths, characterized in that, The method includes: Obtain data packet information; The data packet information is standardized to obtain standardized data; Based on the standardized data and the preset hybrid model, the data is processed in the privacy-preserving feature space to obtain anomaly detection information; The anomaly detection information is detected based on an adaptive threshold strategy to obtain a first detection result; The location information is obtained by identifying the first detection result based on the subspace; The location information is processed based on a path search algorithm to obtain a first path.

2. The method according to claim 1, characterized in that, The data packet information is standardized to obtain standardized data, including: The data packet information is extracted based on the time window segmentation method to obtain time sequence features; The standardized data is obtained by normalizing the time-series characteristics.

3. The method according to claim 2, characterized in that, Based on the standardized data and the preset hybrid model, processing is performed in the privacy-preserving feature space to obtain anomaly detection information, including: Sensitivity calculation is performed on the standardized data to obtain a first sensitivity. Based on the first sensitivity, the HSIC independence calculation for noise compensation is performed to obtain the first calculation result; Based on the first calculation result and the preset hybrid model, the anomaly discrimination information is obtained by processing in the privacy protection feature space.

4. The method according to claim 3, characterized in that, Based on the first calculation result and the preset hybrid model, the anomaly detection information is obtained by processing in the privacy-preserving feature space, including: The second calculation result is obtained by calculating the first calculation result based on the dynamic feature filtering strategy; Based on the second calculation result and the preset hybrid model, the anomaly discrimination information is obtained by processing in the privacy protection feature space.

5. The method according to claim 4, characterized in that, The step of detecting the anomaly discrimination information based on an adaptive threshold strategy to obtain a first detection result includes: The anomaly detection information is used to perform multi-index time series prediction to obtain the first performance index; Perform dynamic threshold calculations on the first performance index to obtain the first threshold. The first detection result is obtained by processing the first threshold based on an adaptive weight adjustment mechanism.

6. The method according to claim 5, characterized in that, Before identifying the first detection result based on the subspace to obtain location information, the process includes: Retrieve historical anomaly patterns; The subspace is obtained by fusing feature activation intensity vectors based on the historical anomaly patterns. The location information is obtained by identifying the first detection result based on the subspace.

7. The method according to claim 6, characterized in that, The step of identifying the first detection result based on subspace to obtain location information includes: The first abnormal sample is obtained by identifying the first detection result based on the subspace; The location information is obtained by locating the first abnormal sample based on the abnormal link.

8. The method according to claim 7, characterized in that, The process of processing the location information using the path search algorithm to obtain the first path includes: The path search algorithm is obtained based on genetic algorithm and simulated annealing strategy; The location information is processed based on the path search algorithm to obtain the first path.

9. An identification device, characterized in that, The device includes: an acquisition unit, an analysis unit, and a processing unit. The acquisition unit is used to acquire data packet information; The analysis unit is used to standardize the data packet information to obtain standardized data; Based on the standardized data and the preset hybrid model, the data is processed in the privacy-preserving feature space to obtain anomaly detection information; the anomaly detection information is detected based on the adaptive threshold strategy to obtain a first detection result; and the first detection result is identified based on the subspace to obtain location information. The processing unit is used to process the location information based on a path search algorithm to obtain a first path.

10. An electronic device, characterized in that, include: Memory, used to store at least one set of instructions; The processor is used to acquire data packet information; The data packet information is standardized to obtain standardized data; Based on the standardized data and the preset hybrid model, the data is processed in the privacy-preserving feature space to obtain anomaly detection information; The anomaly detection information is detected based on an adaptive threshold strategy to obtain a first detection result; The location information is obtained by identifying the first detection result based on the subspace; The location information is processed based on a path search algorithm to obtain a first path.