Feature-based methods for detecting abnormal network traffic

By combining statistical methods, TCN networks, causal models, and principal component reinforcement learning models, network traffic features are screened and classified, solving the problems of false positives, false negatives, and high computational complexity in existing technologies, and achieving efficient and accurate detection of abnormal network traffic.

CN119814421BActive Publication Date: 2025-10-28CHONGQING INST OF GREEN & INTELLIGENT TECH CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411918437.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-10-28
Estimated Expiration
2044-12-25

AI Technical Summary

Technical Problem

Existing methods for detecting abnormal network traffic are prone to false positives or false negatives when faced with complex attacks, and they have high computational complexity and data requirements, making it difficult to achieve efficient and accurate detection.

Method used

By combining statistical methods, TCN networks, causal models, principal component reinforcement learning models, and classifiers, irrelevant features are filtered out through feature selection and feature concatenation. Reinforcement learning networks are then used for feature extraction and classification to achieve efficient and accurate detection of abnormal network traffic.

Benefits of technology

It improves the accuracy and robustness of abnormal network traffic detection, reduces computational complexity, and ensures accurate identification of abnormal network traffic.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119814421B_ABST
    Figure CN119814421B_ABST
Patent Text Reader

Abstract

This invention relates to a feature-based method for detecting abnormal network traffic, belonging to the field of information security. The method comprises the following steps: S1: Data acquisition and preprocessing; S2: Calculation of static features using statistical methods; S3: Extraction of dynamic features using a TCN network and concatenation into the features to be detected; S4: Establishment of a causal model, using the features to be detected as input; S5: Establishment of a principal component reinforcement learning model, using the features to be detected as input; S6: Concatenation of the outputs of the causal model and the principal component reinforcement learning model and input into a classifier to obtain the detection results of abnormal network traffic; S7: Establishment of a loss function, combined with the result labels, to train the network parameters; S8: Detection of abnormal network traffic using the trained network on real-time network traffic. This invention ensures accurate detection of abnormal network traffic while filtering out interfering features and reducing computational complexity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method for detecting abnormal network traffic based on feature selection, belonging to the field of information security, and particularly to a method for detecting abnormal network traffic based on feature selection. Background Technology

[0002] With the widespread adoption of the internet and various online services, the surge in network traffic has presented unprecedented challenges to network security. Network anomaly detection, as a crucial component of network security, aims to promptly identify and respond to potential security threats, protecting network resources and user data.

[0003] The background of abnormal network traffic detection can be traced back to the early stages of network security development. Initially, network security mainly relied on static protective measures such as firewalls and intrusion detection systems (IDS). While these measures could prevent known attacks to a certain extent, they proved inadequate against new and complex attack patterns. As network attack techniques continued to evolve, attackers began to use various means to circumvent traditional security measures, leading to frequent network security incidents and causing huge economic losses and reputational crises for businesses and individuals.

[0004] Against this backdrop, feature-based methods for detecting abnormal network traffic have emerged. These methods analyze the characteristics of network traffic to identify differences between normal and abnormal traffic, thereby detecting potential attacks. These methods can be categorized into several types, including statistical methods, machine learning methods, rule-based methods, and deep learning methods. Statistical methods analyze the basic characteristics of network traffic to quickly identify abnormal traffic situations. However, this method may produce false positives or false negatives when facing complex attacks. Machine learning methods identify traffic patterns by training models, achieving high detection accuracy, but requiring large amounts of labeled data and computational resources. Rule-based methods rely on predefined rules and are suitable for detecting known attacks, but are less adaptable to novel attacks. Deep learning methods, through automatic feature extraction and classification, can handle high-dimensional data and are highly adaptable, but their computational complexity and data requirements are relatively high.

[0005] In essence, network anomaly traffic detection is a binary classification problem, which can be implemented using a classifier. The accuracy of the classification variables, i.e., the features, is the decisive factor in determining the classification accuracy. Single-dimensional features lead to incomplete information, while multiple-dimensional features cause a sharp increase in computational complexity. Summary of the Invention

[0006] The purpose of this invention is to overcome the shortcomings of the prior art and provide a feature-based network anomaly traffic detection method for accurately detecting network anomaly traffic. This method aims to classify static and dynamic features based on classifier principles, and simultaneously combine causal models and principal component analysis (PCA) for mutual verification to significantly reduce irrelevant features in the data. However, fixed, selected features cannot accurately represent randomly changing traffic characteristics. Therefore, reinforcement learning networks are introduced in combination with PCA to ultimately achieve accurate, efficient, and robust identification of network anomaly traffic.

[0007] To achieve the above objectives, the present invention provides the following technical solution:

[0008] The network anomaly traffic detection method based on feature selection is implemented by a network anomaly traffic detection network based on feature selection. The network is characterized by comprising a statistical method module, a TCN network, a causal model, a principal component reinforcement learning model, and a classifier. The statistical method module contains a method for calculating statistical features and is connected in parallel with the TCN network. Both inputs are time-series data, and the outputs are concatenated using concat to obtain the features to be detected. The causal model and the principal component reinforcement learning model are connected in parallel, each using the features to be detected as input for feature selection. The outputs are concatenated using concat and then input to the connected classifier. The classifier detects network anomaly traffic.

[0009] The network anomaly traffic detection method based on feature selection is characterized by comprising the following steps:

[0010] S1: Collect network traffic data and perform preprocessing to obtain time-series data;

[0011] S2: Use the statistical methods module to perform statistics on time series data to obtain static features;

[0012] S3: Use the TCN network to extract features from time-series data to obtain dynamic features, and then concatenate the dynamic features and static features to form the features to be detected;

[0013] S4: Establish a causal model based on the Granger causality test method, using the feature to be detected as input;

[0014] S5: Establish a principal component reinforcement learning model based on the principal component analysis method and reinforcement learning network architecture, and take the feature to be detected as input;

[0015] S6: The outputs of the causal model and the principal component reinforcement learning model are concatenated and then input into a classifier to obtain the detection results of abnormal network traffic;

[0016] S7: Establish a loss function and train the feature-based abnormal traffic detection network by combining it with labeled historical time series data;

[0017] S8: Use the feature-based network anomaly detection network trained in step S7 to detect network anomalies in real-time network traffic.

[0018] Furthermore, the preprocessing described in step S1 needs to include sorting the data according to its timestamp and dividing it according to the sampling period.

[0019] Furthermore, the static features mentioned in step S2 include the mean, variance, and higher-order moments of the time series data within the period.

[0020] Furthermore, the causal model described in step S4 is a selector, and its specific working principle is as follows:

[0021] S401: Utilize the Granger causality test to establish a linear regression model for each feature and label of the feature to be detected; specifically, for

[0022] For pre-set parameters; y t x is the label value at time t; t Let be a feature at time t that is to be detected, and d be a window size set manually.

[0023] S402: Use the output of the regression model as the input of the discrete Hopfield neural network; wherein the input and output of the discrete Hopfield neural network have the same dimension as the feature to be detected, and its output is binary;

[0024] Specifically, if the output information of a neuron is greater than the threshold, then the output value of the neuron is 1; if it is less than the threshold, then the output value of the neuron is 0.

[0025] S403: Use the binary output of the discrete Hopfield neural network to select the causal features to be detected; where 0 indicates no selection and 1 indicates selection.

[0026] Furthermore, the principal component reinforcement learning model described in step S5 is composed of a reinforcement learning network combined with principal component analysis, and the reinforcement learning framework consists of (s t ,a t ,r t The system consists of a triplet; where state s is the state. t The feature to be detected is input at time t; action a t The feature to be detected at time t is selected, i.e., a vector with elements of 0 or 1; the reward r is... tThe percentage of correct detection results for abnormal network traffic output by the classifier at time t; the environment corresponding to the reinforcement learning framework is the classifier, and the corresponding experience pool is generated by the principal component analysis method.

[0027] Specifically, its working principle is as follows:

[0028] S501: Input the historical features to be detected as state inputs into the principal component reinforcement learning model;

[0029] S502: Principal component analysis is used to perform principal component analysis on the features to be detected, and the features corresponding to the first 90% of principal components are retained.

[0030] S503: Set the position corresponding to the feature retained in step S502 to 1, and the rest to 0, to generate the vector of the action, which is used as the action;

[0031] S504: Store the paired state and action in the experience pool;

[0032] S505: The principal component reinforcement learning model matches the nearest state from the replay buffer based on the current state, and then obtains the paired action;

[0033] S506: Filter out the principal component features according to the pairing action obtained in step S505, and output them to the classifier;

[0034] S507: The percentage of correct detection results for abnormal network traffic output by the classifier is used as a reward to update the parameters of the principal component reinforcement learning model.

[0035] Furthermore, the classifier described in step S6 is XGBoost.

[0036] Preferably, to reduce computational complexity and for scenarios with lower recall requirements, the concatenation described in step S6 is to use the AND operation to merge causal features and principal component features. That is, if a feature is selected as both a causal feature and a principal component feature, then this feature will be input into the classifier as a feature; otherwise, the classifier will not process this feature.

[0037] Preferably, to reduce computational complexity and for scenarios with high recall requirements, the concatenation described in step S6 is to use the OR operation to merge causal features and principal component features. That is, if a feature is selected as a causal feature or a principal component feature, this feature will be input into the classifier as a feature; otherwise, the classifier will not process this feature.

[0038] Furthermore, the loss function described in step S7 is Loss = Loss Focal -Loss recallAmong them, Loss Focal =-α t (1-p t ) γ log(p t ) represents the Focal Loss function, p t To predict the probability that the result matches the label, Loss recall α represents the false negative rate of abnormal network traffic samples, i.e., the recall rate of abnormal samples; t γ are hyperparameters.

[0039] Preferably, to improve computational efficiency, the trained feature-based abnormal traffic detection network described in step S8 has its causal model branches pruned, and only the principal component reinforcement learning model is used for feature selection.

[0040] An electronic device includes at least one processor; and a memory communicatively connected to said at least one processor; wherein,

[0041] The memory stores a computer program that is executed by the at least one processor, which enables the at least one processor to perform the feature-based network anomaly traffic detection method described above.

[0042] Finally, the present invention also discloses a computer-readable storage medium storing computer instructions for causing a processor to execute the aforementioned feature-based network abnormal traffic detection method.

[0043] The beneficial effects of this invention are as follows: It provides a network abnormal traffic detection method based on feature screening. First, it uses the static features of statistical methods and the dynamic features of TCN networks to improve the feature dimension. Then, it uses causal models and principal component reinforcement learning models to achieve accurate feature screening. Finally, it uses a classifier to accurately classify and identify the screened features. This invention can ensure the accuracy of network abnormal traffic detection, and can also screen out interfering features and reduce computational complexity. Attached Figure Description

[0044] To make the objectives and technical solutions of this invention clearer, the following figures are provided for illustration:

[0045] Figure 1 This is an architecture diagram of the network anomaly traffic detection network based on feature filtering in Embodiment 1 of the present invention;

[0046] Figure 2 This is a diagram of the principal component reinforcement learning model architecture in Embodiment 1 of the present invention;

[0047] Figure 3 This is a diagram of the architecture of the feature-based network anomaly detection network in Embodiment 2 of the present invention.

[0048] Figure 4 This is a schematic diagram of the electronic device in Embodiment 3 of the present invention. Detailed Implementation

[0049] To make the objectives and technical solutions of this invention clearer, the invention will be described in detail below with reference to the accompanying drawings and embodiments.

[0050] Example 1: To protect the security of internet devices and prevent attacks from abnormal network traffic, network traffic data, including IP address, port information, traffic type, and request frequency, is captured in real time from the web server's network interface to characterize the differences between normal and abnormal behavior. Considering the workload of manual annotation, this example uses pre-annotated open-source datasets (such as NSL-KDD, CSE-CIC-IDS2018, etc.) as the training set and real data collected directly from the web server application as the test set. This example provides a "feature-based method for detecting abnormal network traffic".

[0051] Combination Figure 1 The method is implemented by a feature-based network anomaly traffic detection network. Its features include a statistical method module, a TCN network, a causal model, a principal component reinforcement learning model, and a classifier. The statistical method module contains a method for calculating statistical features and is connected in parallel with the TCN network. Both inputs are time-series data, and the outputs are concatenated using concat to obtain the features to be detected. The causal model and the principal component reinforcement learning model are connected in parallel, each using the features to be detected as input for feature filtering. The outputs are concatenated using concat and then input to the connected classifier. The classifier detects abnormal network traffic.

[0052] Specifically, it includes the following steps:

[0053] Step 1: Perform preprocessing on the CSE-CIC-IDS2018 dataset to obtain time series data.

[0054] The CSE-CIC-IDS2018 dataset contains 80 monitoring feature vectors, covering various attributes of network traffic, aiming to provide comprehensive data support for research on intrusion detection systems. In this embodiment, the abnormal behavior could be the presence of a DDoS attack or other types of attacks. Only some features from the CSE-CIC-IDS2018 dataset are used, such as timestamps, source IP addresses, and access traffic, and specific attack category labels are also included.

[0055] The preprocessing includes, but is not limited to: data cleaning, data normalization, sorting data by timestamp for each IP address, and dividing the data according to the sampling period.

[0056] Step 2: Use the statistical methods module to process the time series data to obtain static features.

[0057] The static features include the mean, variance, and higher-order moments of the time series data within the period, wherein the order of the higher-order moments is similar to the dimension of the dynamic features.

[0058] Step 3: Use the TCN network to extract features from the time series data to obtain dynamic features, and then concatenate the dynamic features and static features to form the features to be detected.

[0059] Step 4: Establish a causal model based on the Granger causality test method, using the feature to be detected as input.

[0060] The causal model described is a selector, and its specific working principle is as follows:

[0061] S401: Use the Granger causality test to establish a linear regression model for each feature and label of the feature to be detected; specifically, the regression model for any feature is as follows: Among them, a j b j The weighting parameters are obtained through fitting analysis and are pre-defined parameters; y t x is the label value at time t; t Let be a feature at time t that is to be detected, and d be a window size set manually.

[0062] S402: Use the output of the regression model as the input of the discrete Hopfield neural network; wherein the discrete Hopfield neural network is a single-layer network, whose input and output have the same dimension as the feature to be detected, and whose output is binary;

[0063] Specifically, if the output information of a neuron is greater than the threshold, then the output value of the neuron is 1; if it is less than the threshold, then the output value of the neuron is 0.

[0064] S403: Use the binary output of the discrete Hopfield neural network to select the causal features to be detected; where 0 indicates no selection and 1 indicates selection.

[0065] Step 5: Establish a principal component reinforcement learning model based on the principal component analysis method and reinforcement learning network architecture, and take the feature to be detected as input.

[0066] Combination Figure 2The principal component reinforcement learning model is composed of a reinforcement learning network combined with principal component analysis, and the reinforcement learning framework consists of (s t ,a t ,r t The system consists of a triplet; where state s is the state. t The feature to be detected is input at time t; action a t The feature to be detected at time t is selected, i.e., a vector with elements of 0 or 1; the reward r is... t The percentage of correct detection results for abnormal network traffic output by the classifier at time t; the environment corresponding to the reinforcement learning framework is the classifier, and the corresponding experience pool is generated by the principal component analysis method.

[0067] Specifically, the steps of its working principle are as follows:

[0068] S501: Input the historical features to be detected as state inputs into the principal component reinforcement learning model;

[0069] S502: Principal component analysis is used to perform principal component analysis on the features to be detected, and the features corresponding to the first 90% of principal components are retained.

[0070] S503: Set the position corresponding to the feature retained in step S502 to 1, and the rest to 0, to generate the vector of the action, which is used as the action;

[0071] S504: Store the paired state and action in the experience pool;

[0072] S505: The principal component reinforcement learning model matches the nearest state from the experience pool based on the current state, and then obtains the paired action;

[0073] S506: Filter out the principal component features according to the pairing action obtained in step S505, and output them to the classifier;

[0074] S507: The percentage of correct detection results for abnormal network traffic output by the classifier is used as a reward to update the parameters of the principal component reinforcement learning model.

[0075] Step 6: Concatenate the outputs of the causal model and the principal component reinforcement learning model and input them into a classifier to obtain the detection results of abnormal network traffic.

[0076] The classifier mentioned is XGBoost.

[0077] Step 7: Establish a loss function and train the feature-based network anomaly traffic detection network by combining it with labeled historical time-series data.

[0078] The loss function is Loss = Loss Focal -Lossrecall Among them, Loss Focal =-α t (1-p t ) γ log(p t ) represents the Focal Loss function, p t To predict the probability that the result matches the label, Loss recall α represents the false negative rate of abnormal network traffic samples, i.e., the recall rate of abnormal samples; t γ are hyperparameters, which are determined by searching using RandomizedSearchCV or HalvingSearchCV in sklearn.

[0079] It should be noted that training the feature-based network anomaly traffic detection network refers to the joint training of the TCN network, discrete Hopfield neural network, principal component reinforcement learning model, and classifier, rather than training each of these networks individually.

[0080] Step 8: Use the feature-based network anomaly detection network trained in Step 7 to detect network anomalies in real-time network traffic.

[0081] Example 2: In view of the scenario in Example 1, this example provides a "network abnormal traffic detection method based on feature filtering" to improve the efficiency of real-time abnormal traffic detection.

[0082] The steps are the same as in Example 1, and will not be repeated here. The difference is that...

[0083] For scenarios with low recall requirements, the concatenation described in step six involves using the AND operation to merge causal features and principal component features. That is, if a feature is selected as both a causal feature and a principal component feature, then this feature will be input into the classifier as a feature; otherwise, the classifier will not process this feature.

[0084] For scenarios requiring high recall, the concatenation described in step six involves using an OR operation to merge causal features and principal component features. That is, if a feature is selected as a causal feature or a principal component feature, this feature will be input into the classifier as a feature; otherwise, the classifier will not process this feature.

[0085] Combination Figure 3 In step eight, the trained feature-based abnormal traffic detection network has its causal model branches pruned, and only the principal component reinforcement learning model is used for feature selection.

[0086] Example 3: For the scenario described in Example 1, Figure 4 A schematic diagram of an electronic device (90) that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers.

[0087] Electronic devices can also refer to various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the invention described and / or claimed herein.

[0088] like Figure 4 As shown, the electronic device (90) includes at least one processor (91) and a memory, such as a read-only memory (ROM) (92) or a random access memory (RAM) (93), which is communicatively connected to the at least one processor (91). The memory stores computer programs executable by the at least one processor. The processor (91) can perform various appropriate actions and processes based on the computer programs stored in the ROM (92) or loaded from storage unit (98) into the RAM (93). The RAM (43) can also store various programs and data required for the operation of the electronic device (90). The processor (91), ROM (42), and RAM (43) are interconnected via a bus (94). An input / output (I / O) interface (95) is also connected to the bus (94).

[0089] Multiple components in the electronic device (90) are connected to an I / O interface (95), including: input units (96), such as a keyboard, mouse, etc.; output units (97), such as various types of displays, speakers, etc.; storage units (98), such as disks, optical disks, etc.; and communication units (99), such as network cards, modems, wireless transceivers, etc. The communication unit (99) allows the electronic device (90) to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0090] The processor (91) can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processors (91) include, but are not limited to, central processing units (CPUs), graphics processing units (GPUs), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. The processor (91) performs the various methods and processes described above, such as feature-based network anomaly traffic detection methods.

[0091] In some embodiments, the feature-based network anomaly traffic detection method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as a storage unit (98). In some embodiments, part or all of the computer program may be loaded and / or installed on an electronic device (90) via a ROM (92) and / or a communication unit (99). When the computer program is loaded into RAM (93) and executed by a processor (91), one or more steps of the feature-based network anomaly traffic detection method described above may be performed. Alternatively, in other embodiments, the processor (91) may be configured to perform the feature-based network anomaly traffic detection method by any other suitable means (e.g., by means of firmware).

[0092] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0093] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0094] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0095] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0096] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0097] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.

[0098] Finally, it should be noted that the above preferred embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail through the above preferred embodiments, those skilled in the art should understand that various changes can be made to it in form and detail without departing from the scope defined by the claims of the present invention.

Claims

1. A network anomaly traffic detection method based on feature selection, characterized in that, Includes the following steps: S1: Collect network traffic data and preprocess it to obtain time-series data; S2: Use the statistical methods module to perform statistics on time series data to obtain static features; S3: Use the TCN network to extract features from time-series data to obtain dynamic features, and then concatenate the dynamic features and static features to form the features to be detected; S4: Establish a causal model based on the Granger causality test method, using the feature to be detected as input; S5: Establish a principal component reinforcement learning model based on the principal component analysis method and reinforcement learning network architecture, and take the feature to be detected as input; S6: The outputs of the causal model and the principal component reinforcement learning model are concatenated and then input into a classifier to obtain the detection results of abnormal network traffic; S7: Establish a loss function and train the feature-based abnormal traffic detection network by combining it with labeled historical time series data; S8: Use the feature-based network anomaly detection network trained in step S7 to detect network anomalies in real-time network traffic. The preprocessing described in step S1 includes sorting the data according to its timestamp and dividing it according to the sampling period; the static features described in step S2 include the mean, variance, and higher-order moments of the time-series data within the period; the classifier described in step S6 is XGBoost; and the loss function described in step S7 is Loss = Loss Focal -Loss recall Among them, Loss Focal =-α t (1-p t ) γ log(p t ) represents the FocalLoss loss function, p t To predict the probability that the result matches the label, Loss recall α represents the false negative rate of abnormal network traffic samples, i.e., the recall rate of abnormal samples; t γ are hyperparameters.

2. The network anomaly traffic detection method based on feature filtering according to claim 1, characterized in that, The feature-based filtering network includes a statistical method module, a TCN network, a causal model, a principal component reinforcement learning model, and a classifier. The statistical method module contains a method for calculating statistical features, which is connected in parallel with the TCN network. The inputs are time-series data, and the outputs are concatenated using concat to obtain the features to be detected. The causal model and the principal component reinforcement learning model are connected in parallel, and each takes the features to be detected as input for feature filtering. The outputs are concatenated using concat and then input to the classifier connected to it. The classifier detects abnormal network traffic.

3. The network anomaly traffic detection method based on feature filtering according to claim 1, characterized in that, The causal model described in step S4 is a selector, and its specific working principle is as follows: S401: Use the Granger causality test to establish a linear regression model for each feature and label of the feature to be detected; specifically, the regression model for any feature is as follows: Among them, a j 、b j The weighting parameters are obtained through fitting analysis and are pre-defined parameters; y t x is the label value at time t; t Let be a feature at time t that is to be detected, and d be a window size set manually. S402: Use the output of the regression model as the input of the discrete Hopfield neural network; wherein the input and output of the discrete Hopfield neural network have the same dimension as the feature to be detected, and its output is binary; S403: Use the binary output of the discrete Hopfield neural network to select the causal features to be detected; where 0 indicates no selection and 1 indicates selection.

4. The network anomaly traffic detection method based on feature filtering according to claim 1, characterized in that, The principal component reinforcement learning model described in step S5 is composed of a reinforcement learning network combined with principal component analysis. The reinforcement learning model consists of (s...) t ,a t ,r t The system consists of a triplet; where state s is the state. t The feature to be detected is input at time t; action a t The feature to be detected at time t is selected, i.e., a vector with elements of 0 or 1; the reward r is... t The percentage of correct network anomaly traffic detection results output by the classifier at time t; the environment corresponding to the reinforcement learning model is the classifier, and the corresponding experience pool is generated by principal component analysis.

5. The network anomaly traffic detection method based on feature filtering according to claim 4, characterized in that, The working principle of the principal component reinforcement learning model is as follows: S501: Input the historical features to be detected as state inputs into the principal component reinforcement learning model; S502: Principal component analysis is used to perform principal component analysis on the features to be detected, and the features corresponding to the first 90% of principal components are retained. S503: Set the position corresponding to the feature retained in step S502 to 1, and the rest to 0, to generate the vector of the action, which is used as the action; S504: Store the paired state and action in the experience pool; S505: The principal component reinforcement learning model matches the nearest state from the replay buffer based on the current state, and then obtains the paired action; S506: Filter out the principal component features according to the pairing action obtained in step S505, and output them to the classifier; S507: The percentage of correct detection results for abnormal network traffic output by the classifier is used as a reward to update the parameters of the principal component reinforcement learning model.

6. The network anomaly traffic detection method based on feature filtering according to claim 1, characterized in that, To reduce computational complexity and for scenarios with lower recall requirements, the concatenation described in step S6 is performed by using the AND operation to merge causal features and principal component features. That is, if a feature is selected as both a causal feature and a principal component feature, then this feature will be input into the classifier as a feature; otherwise, the classifier will not process this feature.

7. The network anomaly traffic detection method based on feature filtering according to claim 1, characterized in that, To reduce computational complexity, for scenarios with high recall requirements, the concatenation described in step S6 is to use the OR operation to merge causal features and principal component features. That is, if a feature is selected as a causal feature or a principal component feature, this feature will be input into the classifier as a feature; otherwise, the classifier will not process this feature.

8. The network anomaly traffic detection method based on feature filtering according to claim 1, characterized in that, To improve computational efficiency, the trained feature-based abnormal traffic detection network described in step S8 has its causal model branches pruned, and only the principal component reinforcement learning model is used for feature selection.

9. An electronic device, characterized in that, The electronic device includes at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the feature-based network abnormal traffic detection method according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the feature-based network abnormal traffic detection method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Causal network discovery system based on reinforcement learning

    CN115171773A

  • Improved DQN-based interpretable monitoring data identification method

    CN116304855A