Multi-dimensional network traffic anomaly detection method based on ensemble learning

By employing ensemble learning and manifold learning techniques, high-dimensional network traffic data is mapped to a low-dimensional manifold space. Combined with a multi-model fusion strategy, this approach addresses the shortcomings of traditional methods in terms of detection accuracy and reliability, enabling efficient anomaly detection for complex network traffic.

CN120750676BActive Publication Date: 2025-11-25上海思来氏信息咨询有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511263169.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-05
Publication Date
2025-11-25
Estimated Expiration
2045-09-05

AI Technical Summary

Technical Problem

Traditional network traffic anomaly detection methods struggle to adapt to dynamic changes in network traffic when dealing with complex and ever-changing network environments, resulting in high rates of false negatives or missed positives. Furthermore, they are unable to effectively uncover deep features and correlations within high-dimensional, non-linear network traffic data, leading to insufficient detection accuracy and reliability.

Method used

A multi-dimensional network traffic anomaly detection method based on ensemble learning is adopted. By combining ensemble learning models with manifold learning technology, high-dimensional network traffic data is mapped to a low-dimensional manifold space. Multiple manifold-based learning models are used for feature extraction and learning. Anomaly scores of data points are calculated by combining a voting mechanism and a weighted average method. Combined with distributed data acquisition and an improved manifold mapping algorithm, the model weights are dynamically adjusted to improve detection accuracy and reliability.

Benefits of technology

It significantly improves the accuracy and reliability of anomaly detection, effectively identifies complex network traffic anomalies, enhances the system's adaptability and real-time performance to dynamic network environments, and ensures the timeliness and effectiveness of anomaly detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120750676B_ABST
    Figure CN120750676B_ABST
Patent Text Reader

Abstract

The application discloses a multi-dimensional network traffic anomaly detection method based on ensemble learning, and relates to the field of network security.The method comprises the following components: a data acquisition and preprocessing step, a manifold mapping step, an ensemble model construction step, a feature extraction step, and an anomaly detection step.Through the ensemble learning model, the high-dimensional network traffic data is mapped to a low-dimensional manifold space in combination with manifold learning, the feature information of the data is mined from multiple angles, feature extraction and learning are carried out by using multiple manifold-based learning models, and the anomaly scores of the data points are comprehensively calculated by using a voting mechanism and a weighted average method.This multi-model fusion strategy significantly improves the accuracy and reliability of anomaly detection, and can effectively identify complex network traffic anomalies.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of network security, in particular to a multi-dimensional network traffic anomaly detection method based on ensemble learning. BACKGROUND

[0002] With the rapid development of Internet technology, the network scale is continuously expanding, and the network traffic data is showing an explosive growth trend. In network traffic, not only normal user access behavior is contained, but also various anomalies and malicious activities such as network attacks, data breaches, and malicious software propagation may be hidden. These abnormal traffic not only affects the normal operation of network services, but also may cause serious threats to user privacy and security.

[0003] Traditional network traffic anomaly detection methods mainly rely on threshold setting, rule matching or simple statistical analysis. These methods are not capable of dealing with complex and variable network environments. First, threshold setting methods are often based on experience or historical data, which is difficult to adapt to the dynamic changes of network traffic, resulting in high false negative or false positive rates. Second, rule matching methods can identify known attack patterns, but are helpless against unknown attack methods. In addition, simple statistical analysis methods often cannot effectively mine deep features and association relationships in high-dimensional, nonlinear network traffic data, thereby limiting the accuracy and reliability of anomaly detection.

[0004] In view of the deficiencies of traditional network traffic anomaly detection methods, it is particularly important to develop a multi-dimensional network traffic anomaly detection method based on ensemble learning. SUMMARY

[0005] The purpose of the present application is to overcome the deficiencies of the prior art and provide a multi-dimensional network traffic anomaly detection method based on ensemble learning. The method can map high-dimensional network traffic data to low-dimensional manifold space through an ensemble learning model combined with manifold learning technology, mine feature information from multiple angles, extract and learn features using multiple manifold-based learning models, and calculate the anomaly score of data points through a voting mechanism and weighted average method, thereby significantly improving the accuracy and reliability of anomaly detection.

[0006] To solve the above technical problems, the present application provides the following technical solution: a multi-dimensional network traffic anomaly detection method based on ensemble learning, the specific steps of which are as follows:

[0007] Data acquisition and preprocessing step: collect network traffic data, clean the collected data to remove duplicate, erroneous and incomplete data records; normalize the data to map it to the [0, 1] or [-1, 1] interval, eliminate the dimensional differences between different dimensions of the data, and obtain the preprocessed network traffic data;

[0008] manifold mapping step: using manifold learning method, the pretreated high-dimensional network traffic data is mapped to low-dimensional manifold space, through calculating the geodesic distance between data points, local linear reconstruction coefficient or constructing graph laplacian matrix, the internal geometric structure of network traffic data is revealed, and the mapped low-dimensional manifold data is obtained;

[0009] integrated model construction step: an integrated learning model is constructed based on manifold support vector, each model uses manifold-based neural network to extract and learn features from the mapped traffic data in the low-dimensional manifold space, and the feature information of the data is mined from different angles;

[0010] feature extraction step: the integrated learning model is used to extract features from the mapped low-dimensional manifold data, each manifold-based learning model extracts different feature vectors from the low-dimensional manifold data, and the feature vectors are fused by using the feature fusion algorithm based on entropy weight method to obtain a more comprehensive and representative feature set for describing the features of network traffic data;

[0011] abnormality detection step: based on the extracted feature set, the integrated learning model is used for abnormality detection, the abnormality score of each data point is calculated to determine whether it is an abnormal point, a voting mechanism is used, that is, multiple manifold-based learning models judge the data points, and the judgment result of the majority model is used to determine whether the data point is abnormal; the weighted average method is used to give different weights according to the performance of each model in the training process, and the abnormality score of the data point is calculated, and the distribution rule of the abnormal point in the low-dimensional manifold space is analyzed to further improve the accuracy and reliability of the abnormality detection.

[0012] Further, in the data collection and preprocessing step, a distributed collection method is used, and collection devices are deployed at multiple key nodes in the network to collect network traffic data in real time. Each collection device performs preliminary processing on the collected data, including data format conversion and simple statistical calculation, and then transmits the processed data to the central server. In the transmission process, an encryption transmission protocol is used to ensure the security and integrity of the data. The distributed collection method can fully cover different areas of the network and obtain more comprehensive network traffic data to support subsequent anomaly detection. Encryption transmission ensures that the data cannot be stolen and tampered with during transmission.

[0013] Further, in the data collection and preprocessing step, a density-based outlier detection algorithm is used to remove incorrect and incomplete data records. Specifically, for each data point in the data set , the neighbor distance , ,in , For the first in the dataset , Data points, To exclude the index of the data point itself, Representing data points and The Euclidean distance between them is then calculated, and the data points are then analyzed. Local density , ,in The truncation distance is determined using an adaptive method, initially set to all values ​​in the dataset. The mean of the nearest neighbor distance is used, and then dynamically adjusted according to the data distribution. When the number of data points is large and the distribution is relatively scattered, the mean distance is increased appropriately. Conversely, it decreases. Set a local density threshold ,like Then determine the data points This algorithm effectively identifies and removes noise and abnormal records from data by recording and deleting erroneous or incomplete data, ensuring the quality of subsequent data processing and laying the foundation for accurate anomaly detection. Compared with traditional data cleaning methods, it can more accurately handle abnormal records in complex network traffic data.

[0014] Furthermore, in the manifold mapping step, the manifold learning method employs an improved isometric mapping algorithm to map high-dimensional network traffic data to a low-dimensional manifold space, introducing time-series information for high-dimensional data points. and In calculating geodesic distance When considering the timestamps of data points and The new distance metric formula is: ,in , For data points , timestamp, Let α be the absolute value of the time difference between two data points, and let α be the time weight parameter determined through cross-validation. The dataset is divided into training and validation sets. Under different values ​​of α, the reconstruction error between data points after mapping the training set to the low-dimensional space is calculated. ,in This refers to the reconstruction error after mapping the data to a low-dimensional space. Given the total number of data points, select the one that minimizes the reconstruction error of the validation set. Using the value as the final time weight parameter, this improved algorithm can better capture the characteristics of network traffic data changing over time, more accurately reveal the inherent geometric structure of the data, and make the mapped low-dimensional manifold data better reflect the true state of network traffic.

[0015] Furthermore, in the ensemble model construction step, when constructing the ensemble learning model, the manifold-based support vector machine employs an improved kernel function, which is: ,in , For data points in a low-dimensional manifold space, for and The square of the Euclidean distance π (pi) and angle Representing data points and The included angle in a low-dimensional manifold space, For kernel function bandwidth parameters, For angle adjustment parameters, and Optimization is performed using a particle swarm optimization algorithm, with the particle swarm size set to [value missing]. The position of each particle represents a set and The objective function is the model's accuracy Acc on the validation set. The goal is to find the value of Acc by continuously updating the particle positions. and By combining these features, the improved kernel function can better adapt to the distribution characteristics of data in low-dimensional manifold spaces, improve the support vector machine's ability to extract network traffic data features, and thus enhance the overall performance of the ensemble learning model.

[0016] Furthermore, the manifold-based neural network employs a dynamic weight adjustment mechanism during feature extraction, specifically for hidden layer neurons within the neural network. Its input weights During training, the system is dynamically adjusted based on the importance of data features, and an importance score for each feature dimension of each data point is calculated. ,in For loss function, Indicates the first The first data point Each feature dimension For the number of data points, For loss function pairs The partial derivatives, The number of feature dimensions is used, and then the weights are adjusted based on the importance score. ,in For the adjusted number The first feature to the first The weights of each hidden layer neuron. The adjustment coefficients were determined using a grid search method. The dynamic weight adjustment mechanism, which is the mean of the importance scores of all feature dimensions, enables the neural network to focus more on important network traffic data features, improves the efficiency and accuracy of feature extraction, and enhances the model's ability to identify abnormal traffic.

[0017] Furthermore, in the feature extraction step, a feature fusion algorithm based on entropy weighting is used for... There are 1 data point, each data point is composed of 1 data point. The feature vectors extracted by each model are composed of, denoted as... , , First calculate the first Information entropy of each feature dimension ,in For the first The data point at the th th Probability distribution along each feature dimension It is a constant. For the first The data point is from the first Feature vector values ​​extracted by the model The number of data points is then used to calculate the information utility value. Thus, the feature weights are obtained. ,in For the first Information utility value of each feature dimension The number of models, the fused feature vector ,in No. The feature vector obtained by fusing data points is determined by the information entropy of the data itself, avoiding the subjectivity of manually setting weights. This allows for a more reasonable fusion of features extracted from various models, making the fused feature set more comprehensively reflect the characteristics of network traffic data and providing strong support for accurate anomaly detection.

[0018] Furthermore, in the anomaly detection step, when calculating the anomaly score of data points using a weighted average method, the weights of each model are determined using an optimization method based on a genetic algorithm, with the population size set to [value missing]. Each individual represents a set of model weights. ,in Let represent the number of manifold-based learning models in the ensemble learning model, and let represent the F1 score of the model on the validation set as the objective function. Through selection, crossover, and mutation genetic operations, the population continuously evolves to find ways to make it more efficient. The method can search the optimal model weight combination from the global, give full play to the advantages of each model, and more accurately calculate the anomaly score of the data point and improve the accuracy of anomaly detection compared with the traditional weight determination method.

[0019] Further, after analyzing the distribution rule of the abnormal points in the low-dimensional manifold space, an abnormal region prediction algorithm based on density peak clustering is used to redefine the local density calculation method of the data points in the low-dimensional manifold space, calculate the neighbor distance of the data points , , wherein d represents the Euclidean distance between the data points , , , , , , , , , , , , , , , , ,

[0020] Compared with the prior art, the multi-dimensional network traffic anomaly detection method based on ensemble learning has the following beneficial effects:

[0021] I. The method maps high-dimensional network traffic data to a low-dimensional manifold space by combining manifold learning with an ensemble learning model, mines feature information of the data from multiple angles, uses multiple manifold-based learning models for feature extraction and learning, and comprehensively calculates the anomaly score of the data point by using a voting mechanism and weighted average. This multi-model fusion strategy significantly improves the accuracy and reliability of anomaly detection, and can effectively identify complex network traffic anomalies.

[0022] II. The method adopts a distributed data collection method, and deployment of collection equipment at multiple key nodes in the network to collect and process network traffic data in real time. In the step of manifold mapping, time series information is introduced to improve the isometric mapping algorithm, making it better capture the characteristics of network traffic data changing over time. In addition, the dynamic weight adjustment mechanism and the model weight optimization method based on genetic algorithm enable the model to dynamically adjust according to real-time data characteristics, enhancing the adaptability and real-time performance of the system to dynamic network environment, ensuring the timeliness and effectiveness of anomaly detection.

[0023] Other advantages, objects, and features of the application will be set forth in part in the description and in part will become apparent to those skilled in the art upon examination of the following or can be learned from practice of the application. The objects and advantages of the application can be realized and attained by means of the instrumentalities and combinations particularly pointed out in the appended claims. BRIEF DESCRIPTION OF DRAWINGS

[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiment or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and those skilled in the art can also obtain other drawings according to these drawings without creative labor.

[0025] Figure 1 Flow chart of the multi-dimensional network traffic anomaly detection method based on ensemble learning;

[0026] Figure 2 Flow chart of the multi-dimensional network traffic anomaly detection method based on ensemble learning. DETAILED DESCRIPTION

[0027] In order to further illustrate the technical means and effects adopted by the present application to achieve the predetermined object of the application, the specific embodiments, structures, features and effects according to the present application will be described in detail below with reference to the drawings and preferred embodiments.

[0028] Embodiment one

[0029] Data collection and preprocessing: Deploy multiple collection devices at key nodes such as core switches of the enterprise intranet, department access switches, server connection ports, etc. Collect network traffic data in real time through distributed collection, including employee computer communication records with external networks, data interaction information between internal servers, and networking behavior data of department office equipment. The collection devices first convert the data format, unify different formats generated by different devices into a format convenient for processing, and perform simple statistical calculations such as the total flow sum and connection times of each device per unit time. Then, use an encrypted transmission protocol to transmit the processed data to the central server, which can effectively prevent data from being stolen or tampered with during transmission and ensure data security. After receiving the data, the central server performs data cleaning and uses a density-based outlier detection algorithm to calculate the near neighbor distance of each data point in the data set , , where is the th data point in the data set, is the index excluding the data point itself, represents the Euclidean distance between data points and . Then calculate the local density of data point , where is the cutoff distance, set a local density threshold , if , then determine that data point is an error or incomplete data record and delete it. Remove duplicate data records and error or incomplete data records to reduce the interference of invalid data on subsequent detection and ensure data quality. Then normalize the cleaned data to make different types and magnitudes of flow characteristics in the same order of magnitude, facilitating subsequent model learning and analysis.

[0030] Manifold mapping: Use the improved isometric mapping algorithm in manifold learning to map the preprocessed high-dimensional network traffic data to a low-dimensional manifold space. In the mapping process, calculate the geodesic distance between data points and introduce timestamp information. For high-dimensional data points and , when calculating the geodesic distance , consider the timestamps and of the data points. The new distance measurement formula is: where , For data points , timestamp, Let α be the absolute value of the time difference between two data points, and let α be the time weight parameter determined through cross-validation. The dataset is divided into training and validation sets. Under different values ​​of α, the reconstruction error between data points after mapping the training set to the low-dimensional space is calculated. ,in This refers to the reconstruction error after mapping the data to a low-dimensional space. Given the total number of data points, select the one that minimizes the reconstruction error of the validation set. The value is used as the final time weight parameter to reveal the inherent geometric structure of network traffic data over time. This approach can better preserve the important features and relationships between data in high-dimensional data, simplifying complex high-dimensional data into more easily processed low-dimensional manifold data, laying the foundation for subsequent feature extraction and anomaly detection.

[0031] Ensemble Model Construction: An ensemble learning model is constructed, which includes multiple manifold-based learning models, such as manifold-based support vector machines and manifold-based neural networks. The manifold-based support vector machine employs an improved kernel function, which is as follows: ,in , For data points in a low-dimensional manifold space, for and The square of the Euclidean distance Pi Representing data points and The included angle in a low-dimensional manifold space, For kernel function bandwidth parameters, Using angle adjustment parameters, the distribution characteristics of data in low-dimensional manifold space can be captured more effectively; manifold-based neural networks employ a dynamic weight adjustment mechanism for hidden layer neurons in the neural network. Its input weights During training, the system is dynamically adjusted based on the importance of data features, and an importance score for each feature dimension of each data point is calculated. ,in For loss function, Indicates the first The first data point Each feature dimension For the number of data points, For loss function pairs The partial derivatives, is the number of feature dimensions, and then the weight is adjusted according to the importance score wherein is the adjustment coefficient, is the mean of the importance scores of all feature dimensions, and the weight can be automatically adjusted according to the importance of the data features during the training process, so that the model pays more attention to the features that are more critical to anomaly detection. By constructing such an ensemble model, feature information in low-dimensional manifold data can be mined from different perspectives, and the learning ability of the model for complex network traffic patterns can be improved.

[0032] Feature extraction step: using the constructed ensemble learning model, the mapped low-dimensional manifold data is subjected to feature extraction, and each manifold-based learning model extracts different feature vectors from the low-dimensional manifold data, for example, a manifold-based support vector machine may extract features related to the data classification boundary, a manifold-based neural network may extract nonlinear correlation features between data, etc. Then, a feature fusion algorithm based on entropy weight method is used, for each data point , each data point is composed of feature vectors extracted by models, denoted as , , , first calculate the information entropy of the th feature dimension , wherein is the probability distribution of the th data point in the th feature dimension, is a constant, is the feature vector value extracted by the th model for the th data point, the number of data points, then calculate the information utility value , and then obtain the feature weight , and the fused feature vector , wherein is the fused feature vector of the th data point. These different feature vectors are fused to obtain a comprehensive feature set. This fusion method can integrate the effective information extracted by each model, reduce the limitations of single features, and improve the representativeness of the features.

[0033] Anomaly detection step: based on the extracted feature set, an integrated learning model is used for anomaly detection. First, the anomaly score of each data point is calculated. A voting mechanism is used, that is, multiple manifold-based learning models judge the data points respectively, and the judgment result of the majority model is used to preliminarily determine whether the data point is abnormal. At the same time, the anomaly score of the data point is calculated by weighted average, wherein the weight of each model is determined according to its performance in the training process. The better the performance of the model, the higher the weight. In addition, by analyzing the distribution rule of the abnormal points in the low-dimensional manifold space, an abnormal area prediction algorithm based on density peak clustering is used to redefine the local density calculation method of the data points in the low-dimensional manifold space , calculate the neighbor distance of the data point , , wherein represents the Euclidean distance between the data points and , then calculate the local density of the data point , , wherein is the cutoff distance, then calculate the distance , set the density threshold and the distance threshold , if and , then determine that the data point is the cluster center, identify the area where the abnormal flow is concentrated, and further improve the accuracy and reliability of the anomaly detection. For example, it can timely discover abnormal behaviors such as malicious attacks and unauthorized access in the enterprise internal network.

[0034] Example two

[0035] Data collection and preprocessing: Deploy collection devices at key locations such as server cluster access points, user login verification nodes, payment transaction interfaces, and product browsing data interaction nodes on e-commerce platforms. Collect network traffic data in real time through distributed collection methods, including user browsing records, adding products to shopping carts, order data, payment transaction details, and page dwell time. The collection devices first convert the collected data into a standard format, perform simple statistical calculations, and then use an encrypted transmission protocol to transmit the processed data to the central server. This prevents data from being leaked or tampered with during transmission, ensuring the security of user information and transaction data. The central server cleans the received data using a density-based outlier detection algorithm to remove duplicate user access records, incorrect data due to system failures (such as incorrectly formatted transaction information), and incomplete data (such as order records missing key fields). This step reduces the interference of invalid data on subsequent analysis and ensures data accuracy. After cleaning, the data is normalized to adjust different magnitudes of features (such as transaction amounts and access frequencies) to the same comparison dimension, facilitating subsequent model learning and analysis.

[0036] Manifold mapping: Use the improved isometric mapping algorithm in manifold learning to map preprocessed high-dimensional network traffic data (including user ID, product ID, transaction amount, access duration, payment method, and other dimensions) to a low-dimensional manifold space. In the mapping process, calculate the geodesic distance between data points and introduce timestamp information (such as user access time and transaction completion time) to reveal the internal geometric structure of e-commerce platform network traffic over time (such as the difference in traffic during promotional activities and daily periods). This approach simplifies data dimensions while preserving important feature associations in high-dimensional data (such as the association between user browsing behavior and subsequent transactions), converting complex high-dimensional data into more manageable low-dimensional manifold data, providing a clearer data foundation for subsequent feature extraction and anomaly detection.

[0037] Integrated model construction: Build an integrated learning model containing multiple manifold-based learning models, such as manifold-based support vector machines and manifold-based neural networks. The manifold-based support vector machine uses an improved kernel function that combines distance and angle information in the low-dimensional manifold space to more accurately capture data distribution characteristics and improve the ability to identify abnormal transaction patterns. The manifold-based neural network uses a dynamic weight adjustment mechanism to dynamically adjust the input weights of hidden layer neurons based on the importance of different features (such as transaction amounts and product categories) to detection results during training, allowing the model to focus more on features that play a key role in anomaly detection, thereby improving the model's learning effect on complex e-commerce traffic patterns.

[0038] Feature extraction: using the constructed ensemble learning model, the low-dimensional manifold data after mapping is subjected to feature extraction, each manifold-based learning model extracts feature vectors from different angles, for example, a manifold-based support vector machine may extract features related to the classification boundary of the data (such as the boundary features of normal transactions and abnormal transactions), a manifold-based neural network may extract non-linear correlation features between data (such as the correlation features of multiple abnormal browsing of a user and malicious ordering), and then a feature fusion algorithm based on entropy weight method is used to fuse these different feature vectors to obtain a comprehensive feature set. This fusion method can integrate the effective information extracted by each model, avoid the limitations of single features, enhance the representativeness of the features for e-commerce traffic anomaly patterns, and provide more comprehensive basis for subsequent anomaly detection.

[0039] Anomaly detection: based on the extracted feature set, an ensemble learning model is used for anomaly detection. First, the anomaly score of each data point is calculated, and a voting mechanism is used, that is, multiple manifold-based learning models judge the data points respectively, and according to the judgment results of the majority of the models, it is preliminarily determined whether the data points are abnormal; at the same time, the anomaly score of the data points is calculated by weighted average, and the weight of each model is determined according to its performance in the training process (such as the model with higher detection accuracy has higher weight), which can more reasonably integrate the detection results of each model. In addition, by analyzing the distribution rule of the abnormal points in the low-dimensional manifold space, an anomaly region prediction algorithm based on density peak clustering is used to identify the region where the abnormal traffic is concentrated (such as frequent abnormal ordering of a certain IP segment, a large number of false payments in a certain time period), which further improves the accuracy and reliability of anomaly detection, helps e-commerce platforms to discover malicious brushing, payment fraud, false transactions and other abnormal behaviors in time, and protects the normal operation of the platform and the property safety of users.

[0040] The above is only a preferred embodiment of the present application, and does not limit the present application in any form. Although the preferred embodiment of the present application has been disclosed as above, it is not intended to limit the present application. Any person skilled in the art can make some changes or modifications to the above disclosed technical content without departing from the scope of the technical solution of the present application, and any equivalent embodiments with equivalent changes are equivalent to the above embodiments. Any modification, change and modification of the above embodiments according to the technical essence of the present application are still within the scope of the technical solution of the present application.

Claims

1. A multi-dimensional network traffic anomaly detection method based on ensemble learning, characterized in that, The specific steps of the method are: Data acquisition and preprocessing step: collect network traffic data, clean the collected data to remove duplicate, error and incomplete data records; normalize the data; Manifold mapping step: use manifold learning method to map the preprocessed high-dimensional network traffic data to low-dimensional manifold space, reveal the internal geometric structure of network traffic data by calculating the geodesic distance between data points, local linear reconstruction coefficients or constructing Laplacian matrix of graph, and get the mapped low-dimensional manifold data; Integrated model construction step: an integrated learning model is constructed based on manifold support vector machine, each model uses manifold-based neural network to extract and learn the features of the mapped traffic data in the low-dimensional manifold space, and mines the feature information of the data from different angles; Feature extraction step: use the integrated learning model to extract features from the mapped low-dimensional manifold data, each manifold-based learning model extracts different feature vectors from the low-dimensional manifold data, and fuses these feature vectors using an entropy weight-based feature fusion algorithm; The improved kernel function is: wherein , is a data point in the low-dimensional manifold space, is the Euclidean distance square of , is the ratio of the circumference of a circle to its diameter, and angle represents the angle between the data points and in the low-dimensional manifold space, is a kernel function bandwidth parameter, is an angle adjustment parameter. Abnormality detection step: based on the extracted feature set, use the integrated learning model to detect anomalies, calculate the anomaly score of each data point to determine whether it is an abnormal point, use the voting mechanism, that is, multiple manifold-based learning models judge the data points, and determine whether the data points are abnormal according to the judgment results of the majority models; through weighted average, give different weights according to the performance of each model in the training process, and comprehensively calculate the anomaly score of the data points to analyze the distribution rule of the abnormal points in the low-dimensional manifold space, and use the abnormal area prediction algorithm based on density peak clustering to obtain the abnormal area.

2. The multi-dimensional network traffic anomaly detection method based on ensemble learning according to claim 1, characterized in that, In the data acquisition and preprocessing step, a distributed acquisition method is used, that is, acquisition devices are deployed at multiple key nodes in the network to collect network traffic data in real time, each acquisition device performs preliminary processing on the collected data, including data format conversion and simple statistical calculation, and then transmits the processed data to the central server. In the transmission process, an encrypted transmission protocol is used. 3.The ensemble learning based multi-dimensional network traffic anomaly detection method of claim 1, wherein, In the data cleaning in the data collection and preprocessing step, the density-based outlier detection algorithm is used to remove error and incomplete data records, specifically: for each data point in the data set , calculate its nearest neighbor distance , , where , is the , data point in the data set, is the index excluding the data point itself, represents the Euclidean distance between the data point and , then calculate the local density of the data point , , where is the cutoff distance, set the local density threshold , if , then determine that the data point is an error or incomplete data record and is deleted.

4. The multi-dimensional network traffic anomaly detection method based on ensemble learning according to claim 1, characterized in that, In the manifold mapping step, the manifold learning method employs an improved isometric mapping algorithm to map high-dimensional network traffic data to a low-dimensional manifold space, incorporating time-series information for high-dimensional data points. and In calculating geodesic distance When considering the timestamps of data points and The new distance metric formula is: ,in , For data points , timestamp, Let α be the absolute value of the time difference between two data points, and let α be the time weight parameter determined through cross-validation. The dataset is divided into training and validation sets. Under different values ​​of α, the reconstruction error between data points after mapping the training set to the low-dimensional space is calculated. ,in This refers to the reconstruction error after mapping the data to a low-dimensional space. Given the total number of data points, select the one that minimizes the reconstruction error of the validation set. The value serves as the final time weight parameter.

5. The multi-dimensional network traffic anomaly detection method based on ensemble learning according to claim 1, characterized in that, The manifold-based neural network employs a dynamic weight adjustment mechanism during feature extraction, adjusting the weights for hidden layer neurons within the neural network. Its input weights During training, the system is dynamically adjusted based on the importance of data features, and an importance score for each feature dimension of each data point is calculated. ,in For loss function, Indicates the first The first data point Each feature dimension For the number of data points, For loss function pairs The partial derivatives, The number of feature dimensions is used, and then the weights are adjusted based on the importance score. ,in To adjust the coefficient, This is the mean of the importance scores for all feature dimensions.

6. The multi-dimensional network traffic anomaly detection method based on ensemble learning according to claim 1, characterized in that, In the feature extraction step, a feature fusion algorithm based on entropy weight method is adopted, for each data point composed of a feature vector extracted by a model, denoted as Firstly, the information entropy of the first feature dimension is calculated , wherein is the probability distribution of the first data point on the first feature dimension, is a constant, is the feature vector value of the first data point extracted by the first model, is the number of data points, and then the information utility value is calculated , and the feature weight is obtained , wherein is the information utility value of the first feature dimension, is the number of models, and the fused feature vector is , wherein is the fused feature vector of the first data point.​​​​ 7. The multi-dimensional network traffic anomaly detection method based on ensemble learning according to claim 1, characterized in that, In the abnormality detection step, the abnormality score of the data point is calculated by using a weighted average method, and the determination of the model weight is performed by using an optimization method based on a genetic algorithm, and the population size is set to Each individual represents a set of model weights , wherein is the number of manifold-based learning models in the ensemble learning model, and the objective function is the F1 value of the model on the validation set Through genetic operations such as selection, crossover and mutation, the population is constantly evolved to find a set of weights that maximize the value as the final model weight.

8. The multi-dimensional network traffic anomaly detection method based on ensemble learning of claim 1, wherein, After analyzing the distribution patterns of outliers in the low-dimensional manifold space, an anomaly region prediction algorithm based on density peak clustering is used for data points in the low-dimensional manifold space. Redefine its local density calculation method and calculate data points of Nearest neighbor distance , ,in Representing data points and The Euclidean distance between them is then calculated, and the data points are then analyzed. Local density , ,in To truncate the distance, then calculate its distance. Set density threshold and distance threshold ,like and Then determine the data points The cluster centers are used to cluster the data points. Data points are clustered based on the cluster centers. Clusters with few data points and far from other clusters are identified as abnormal regions.

Citation Information

Patent Citations

  • Helium leak detection method and system for switch cabinet

    CN118882947A

  • Method and system for sensing credible state of DCS (Distributed Control System) controller

    CN119045448A