Abnormal traffic detection network training method, abnormal traffic detection method and device
By generating normal and abnormal traffic characteristics, using the traffic encoder and detection language model to construct positive and negative sample pairs, and calculating the ternary loss function, the problem of strong dependence on the distribution of initial data characteristics in the prior art is solved, and an abnormal traffic detection with high accuracy is achieved.
Patent Information
- Application Number
- CN202510874933.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-06-27
AI Technical Summary
The anomaly detection algorithm based on deep learning in the prior art has strong dependence on the initial data feature distribution, making it difficult to adapt to the dynamically changing network traffic characteristics and emerging unknown patterns in the computing power network, resulting in low accuracy of abnormal traffic detection.
By generating normal and abnormal traffic characteristics, using the first and second traffic encoders to perform hidden spatial variable mapping, construct positive and negative sample pairs, calculate the function value of the ternary loss function, based on this, the abnormal traffic detection network is weighted, and a comparison learning mechanism is introduced to achieve automatic generation and dynamic adjustment of prompt words.
It improves the accuracy of abnormal traffic detection, can adapt to the abnormal detection needs in diverse network operation and maintenance scenarios, and effectively identify various abnormal traffic.
Smart Images

Figure CN120389910A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of network operation and maintenance technologies, and particularly to an abnormal traffic detection network training method, an abnormal traffic detection method, and an apparatus. Background Art
[0002] With the continuous access of new computing power network element entities, the computing power network has evolved into a multi-level and high-dimensional heterogeneous network architecture jointly composed of macro base stations, micro base stations, and various small base stations. The complexity of this architecture directly leads to an exponential growth in network data traffic, increasing the difficulty of detecting abnormal traffic in the operation and maintenance system.
[0003] In related technologies, deep learning-based anomaly detection algorithms are adopted. However, such algorithms usually rely on static training sets for model training, so they have a strong dependence on the initial data feature distribution. When the network traffic characteristics change dynamically, the detection performance drops sharply. Moreover, when facing newly emerging unknown traffic patterns in the computing power network, the model generalization ability is significantly insufficient, and it is difficult to adapt to the spatio-temporal differences in data distribution in a large-scale heterogeneous network environment. These reasons together lead to a low accuracy in detecting abnormal traffic in the computing power network with large-scale data patterns. Summary of the Invention
[0004] The main purpose of the embodiments of this application is to propose an abnormal traffic detection network training method, an abnormal traffic detection method, and an apparatus, so as to improve the accuracy of detecting abnormal traffic in the computing power network.
[0005] To achieve the above object, a first aspect of the embodiments of this application proposes an abnormal traffic detection network training method. The abnormal traffic detection network includes a first traffic encoder and a detection language model. The method includes: Generating normal traffic features based on the network traffic data extracted from the computing power network, and generating abnormal traffic features based on the normal traffic features; Inputting the normal traffic features into the first traffic encoder for encoding to obtain normal latent space variables, and obtaining abnormal latent space variables according to the normal latent space variables and the abnormal traffic features; Inputting the normal latent space variables and the abnormal latent space variables into the detection language model for reconstruction respectively to obtain normal traffic representations and abnormal traffic representations, and constructing positive sample pairs and negative sample pairs according to the normal traffic features, the normal traffic representations, and the abnormal traffic representations; Calculating the function value of the triplet loss function according to the first distance of the positive sample pairs and the second distance of the negative sample pairs, and based on the function value, adjusting the weights of at least the abnormal traffic detection network until the trained abnormal traffic detection network is obtained.
[0006] In one embodiment, obtaining the abnormal latent space variable according to the normal latent space variable and the abnormal traffic characteristics includes: Inputting the abnormal flow characteristics into a second flow encoder for encoding to obtain an abnormal intermediate variable, wherein the model structure of the second flow encoder is consistent with that of the first flow encoder; The abnormal latent space variable is obtained according to the sum of the normal latent space variable and the abnormal intermediate variable.
[0007] In one embodiment, generating abnormal traffic features based on the normal traffic features includes: Based on a preset sliding window, calculating a sliding average corresponding to each characteristic value in the normal flow characteristic, and forming a seasonal series component according to the sliding average; Calculating a difference sequence between the normal flow characteristic and the seasonal sequence component, performing local weighted regression fitting on the difference sequence to obtain a trend sequence component; The abnormal flow feature is obtained by subtracting the seasonal sequence component and the trend sequence component from the normal flow feature.
[0008] In one embodiment, constructing positive sample pairs and negative sample pairs based on the normal traffic characteristics, the normal traffic representation, and the abnormal traffic representation includes: Taking the normal traffic feature as an anchor sample, taking the normal traffic representation as a normal sample, and taking the abnormal traffic representation as an abnormal sample; The positive sample pair is formed according to the anchor point feature value of the anchor point sample and the normal feature value of the normal sample, and the negative sample pair is formed according to the anchor point feature value and the abnormal feature value of the abnormal sample.
[0009] In one embodiment, calculating the function value of the ternary loss function according to the first distance of the positive sample pair and the second distance of the negative sample pair includes: Calculating a first similarity between the anchor point feature value and the normal feature value to obtain the first distance; Calculating a second similarity between the anchor point feature value and the abnormal feature value in all the negative sample pairs corresponding to the anchor point sample, selecting the negative sample pair corresponding to the minimum value of the second similarity as the minimum negative sample pair, and using the second similarity of the minimum negative sample pair as the second distance; The function value of the ternary loss function is calculated based on the first distance and the second distance.
[0010] In one embodiment, the calculating the function value of the ternary loss function according to the first distance and the second distance includes: Calculate the distance difference between the first distance and the second distance; Calculate the sum of the distance difference and a preset margin parameter to obtain the loss value. If the loss value is negative, set the loss value to zero.
[0011] In one embodiment, the detection language model at least includes a feed-forward network layer, a normalization layer, and a multi-head attention module. Based on the function value, at least adjust the weights of the abnormal traffic detection network until the trained abnormal traffic detection network is obtained, including: During training, freeze the model parameters of the multi-head attention module; Based on the function value, adjust the model parameters corresponding to the first traffic encoder, the second traffic encoder, the feed-forward network layer, and the normalization layer until the iteration termination condition is reached, and obtain the trained abnormal traffic detection network.
[0012] To achieve the above object, a second aspect of the embodiments of the present application proposes an abnormal traffic detection method, which is executed by an abnormal traffic monitoring network trained by using the abnormal traffic detection network training method according to any one of the first aspect. The method includes: Obtain initial traffic data and convert the initial traffic data into initial traffic features; Input the initial traffic features into the first traffic encoder for feature encoding to obtain latent space representation features; Input the latent space representation features into the detection language model for reconstruction to obtain target traffic features, and calculate the reconstruction error between the target traffic features and the latent space representation features. If the reconstruction error is greater than a preset threshold, the detection result indicates that the initial traffic data is abnormal traffic.
[0013] To achieve the above object, a third aspect of the embodiments of the present application proposes an abnormal traffic detection network training device. The abnormal traffic detection network includes a first traffic encoder and a detection language model. The device includes: A traffic feature generation module: used to generate normal traffic features according to the network traffic data extracted from the computing power network, and generate abnormal traffic features based on the normal traffic features; A latent space mapping module: used to input the normal traffic features into the first traffic encoder for encoding to obtain normal latent space variables, and obtain abnormal latent space variables according to the normal latent space variables and the abnormal traffic features; A comparison sample pair generation module is configured to input the normal latent space variables and the abnormal latent space variables into the detection language model for reconstruction, thereby obtaining normal traffic representation and abnormal traffic representation, and constructing positive sample pairs and negative sample pairs based on the normal traffic features, the normal traffic representation, and the abnormal traffic representation; Loss adjustment module: used to calculate the function value of the ternary loss function according to the first distance of the positive sample pair and the second distance of the negative sample pair, and based on the function value, at least adjust the weight of the abnormal traffic detection network until the trained abnormal traffic detection network is obtained.
[0014] To achieve the above-mentioned purpose, the fourth aspect of an embodiment of the present application proposes an electronic device, which includes a memory and a processor, the memory stores a computer program, and the processor implements the method described in the first or second aspect above when executing the computer program.
[0015] To achieve the above-mentioned purpose, the fifth aspect of the embodiment of the present application proposes a storage medium, which is a storage medium. The storage medium stores a computer program, and when the computer program is executed by a processor, it implements the method described in the first or second aspect above.
[0016] The abnormal traffic detection network training method, abnormal traffic detection method and device proposed in the embodiments of the present application generate normal traffic features based on network traffic data extracted from the computing power network, and generate abnormal traffic features based on the normal traffic features, input the normal traffic features into the first traffic encoder for encoding to obtain normal latent space variables, and obtain abnormal latent space variables based on the normal latent space variables and the abnormal traffic features, respectively input the normal latent space variables and the abnormal latent space variables into the detection language model for reconstruction to obtain normal traffic representation and abnormal traffic representation, construct positive sample pairs and negative sample pairs based on the normal traffic features, the normal traffic representation and the abnormal traffic representation, calculate the function value of the ternary loss function based on the first distance of the positive sample pair and the second distance of the negative sample pair, and based on the function value, at least adjust the weight of the abnormal traffic detection network until a trained abnormal traffic detection network is obtained. The embodiments of the present application train the abnormal traffic detection network based on comparative prompt word learning for complex and diverse computing power network operation and maintenance scenarios. During the training process, the normal traffic features are first adjusted by the first traffic encoder to adjust the neural network weights, converting them into learnable dynamic prompt words to achieve adaptive matching for specific operation and maintenance tasks. On this basis, a contrastive learning mechanism is introduced to automatically generate and dynamically adjust prompt words. A triplet loss function is used to construct a fine-tuning optimization objective for the detection language model. This enhances the feature separability between normal and abnormal traffic, reduces the reconstruction error between different traffic types, and employs a transfer learning strategy, enabling the detection language model to adapt to new scenarios with fine-tuning of a small number of samples based on pre-training. The resulting anomaly traffic detection network can effectively identify various types of abnormal traffic and meet the anomaly detection needs in diverse network operation and maintenance scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 This is a flowchart of the abnormal traffic detection network training method provided in an embodiment of the present application.
[0018] Figure 2 This is a flowchart of generating abnormal traffic characteristics based on the normal traffic characteristics provided by an embodiment of the present application.
[0019] Figure 3 This is a flowchart of obtaining abnormal latent space variables based on normal latent space variables and abnormal traffic characteristics provided by an embodiment of the present application.
[0020] Figure 4 This is a schematic diagram of latent space mapping provided in an embodiment of the present application.
[0021] Figure 5 This is a flowchart of constructing positive sample pairs and negative sample pairs based on normal traffic characteristics, normal traffic representation and abnormal traffic representation provided by an embodiment of the present application.
[0022] Figure 6This is a flowchart of an embodiment of the present application for calculating the function value of a ternary loss function based on the first distance of a positive sample pair and the second distance of a negative sample pair.
[0023] Figure 7 This is a schematic diagram of data processing for positive and negative sample alignment provided in an embodiment of the present application.
[0024] Figure 8 This is a flowchart of an embodiment of the present application for calculating the function value of a ternary loss function based on the first distance and the second distance.
[0025] Figure 9 This is a schematic diagram of the calculation process of the ternary loss function provided in the embodiment of the present application.
[0026] Figure 10 This is an overall flow chart of the abnormal traffic detection network training method provided in an embodiment of the present application.
[0027] Figure 11 This is a flow chart of the abnormal traffic detection method provided in an embodiment of the present application.
[0028] Figure 12 This is a structural block diagram of an abnormal traffic detection network training device provided by another embodiment of the present application.
[0029] Figure 13 This is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0030] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0031] It should be noted that although the functional modules are divided in the device schematic and the logical order is shown in the flowchart, in some cases, the steps shown or described can be performed in a different order than the module division in the device or the order in the flowchart.
[0032] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0033] First, let’s analyze some of the terms used in this application: Artificial Intelligence (AI): It is a new technical science that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence; artificial intelligence is a branch of computer science. Artificial intelligence attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. The research in this field includes robots, speech recognition, image recognition, natural language processing, and expert systems, etc. Artificial intelligence can simulate the information process of human consciousness and thinking. Artificial intelligence also refers to the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.
[0034] With the continuous access of new computing power network element entities, the computing power network has evolved into a multi-level and high-dimensional heterogeneous network architecture jointly composed of macro base stations, micro base stations, and various small base stations. The complexity of this architecture directly leads to an exponential growth in network data traffic, requiring the operation and maintenance system to have higher performance, more efficient algorithms, and more powerful hardware support. In addition, the access of new computing power network element entities makes the allocation and scheduling of computing power resources more complex, and it is impossible to directly achieve fast and efficient computing power resource allocation, resulting in resource waste and task delay. It is difficult to achieve effective resource scheduling in a cross-regional and cross-vendor environment, and there are compatibility and stability problems. All of the above problems increase the difficulty of detecting abnormal traffic in the operation and maintenance system.
[0035] In related technologies, anomaly detection algorithms based on deep learning are adopted. However, these algorithms usually rely on static training sets for model training, so they have a strong dependence on the initial data feature distribution and can only be effectively used in several fixed detection scenarios. When the network traffic characteristics change dynamically, the detection performance drops sharply. Moreover, when facing new unknown traffic patterns in the computing power network, the model generalization ability is significantly insufficient, and it is difficult to adapt to the spatio-temporal differences in data distribution in a large-scale heterogeneous network environment. In addition, more and more complex network operation and maintenance tasks make it impossible to generate prompt words one by one using the method of artificial prior knowledge. These reasons together lead to a low accuracy of anomaly detection in the computing power network with large-scale data patterns.
[0036] Based on this, the embodiments of the present application provide an abnormal traffic detection network training method, an abnormal traffic detection method, and a device. For the complex and diverse operation and maintenance scenarios of the computing power network, the abnormal traffic detection network is trained based on contrastive prompt learning. During the training process, first, the neural network weights of the normal traffic features are adjusted by the first traffic encoder, and they are transformed into learnable dynamic prompts to achieve adaptive matching for specific operation and maintenance tasks. On this basis, a contrastive learning mechanism is introduced to realize the automatic generation and dynamic adjustment of prompts, and a fine-tuning optimization objective for the detection language model is constructed based on the triplet loss function. Thus, the feature separability between normal traffic and abnormal traffic is enhanced, the reconstruction error between different types of traffic is enlarged, and a transfer learning strategy is adopted to enable the detection language model to adapt to new scenarios through fine-tuning with a small number of samples on the basis of pre-training. The finally constructed abnormal traffic detection network can effectively identify various types of abnormal traffic and meet the abnormal detection requirements in diverse network operation and maintenance scenarios.
[0037] The embodiments of the present application provide an abnormal traffic detection network training method, an abnormal traffic detection method, and a device, which are specifically described through the following embodiments. First, the abnormal traffic detection network training method in the embodiments of the present application is described.
[0038] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is a theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.
[0039] Artificial intelligence technology is an interdisciplinary subject that involves a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0040] The abnormal traffic detection network training method provided by the embodiments of the present application relates to the technical field of network operation and maintenance. The abnormal traffic detection network training method provided by the embodiments of the present application can be applied to a terminal, or to a server side, or can be a computer program running on the terminal or the server side. For example, the computer program can be a native program or software module in an operating system; it can be a local (Native) application (Application, APP), that is, a program that needs to be installed in the operating system to run, such as a client supporting the training of an abnormal traffic detection network, that is, a program that only needs to be downloaded into a browser environment to run; it can also be a small program that can be embedded into any APP. In short, the above computer program can be any form of application program, module or plug-in. Among them, the terminal communicates with the server through a network. The abnormal traffic detection network training method can be executed by the terminal or the server, or jointly executed by the terminal and the server.
[0041] In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, or a smart watch, etc. The server can be an independent server, or can be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms; it can also be a service node in a blockchain system, and the service nodes in the blockchain system form a peer-to-peer (Peer To Peer, P2P) network, and the P2P protocol is an application layer protocol running on top of the Transmission Control Protocol (TCP) protocol. The connection between the terminal and the server can be made through communication connection methods such as Bluetooth, Universal Serial Bus (USB), or network, and this embodiment does not limit this here.
[0042] This application can be used in numerous general or specific computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are executed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0043] It should be noted that in each specific embodiment of this application, when it comes to relevant processing based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiments of this application need to obtain the user's sensitive personal information, the user's separate permission or separate consent will be obtained through methods such as pop-up windows or redirecting to a confirmation page. After clearly obtaining the user's separate permission or separate consent, the necessary user-related data for the normal operation of the embodiments of this application will be obtained.
[0044] The following describes the abnormal traffic detection network training method in the embodiments of this application.
[0045] Figure 1 is an optional flowchart of the abnormal traffic detection network training method provided by the embodiments of this application. Figure 1 The method in may include but is not limited to steps 110 to 140. At the same time, it can be understood that this embodiment does not specifically limit the order of steps 110 to 140 in, and the order of steps can be adjusted according to actual needs, or some steps can be reduced or added. Figure 1 The order of steps 110 to 140 in can be adjusted according to actual needs, or some steps can be reduced or added.
[0046] Step 110: Generate normal traffic features based on the network traffic data extracted from the computing power network, and generate abnormal traffic features based on the normal traffic features.
[0047] In one embodiment, tools are first used to extract technical and statistical features related to network traffic in computing networks (such as cloud computing, edge computing, and data center networks). Basic features are directly extracted from packet or flow header information and are typically used to describe basic traffic attributes. For example, packet capture tools can be used to capture traffic source and destination IP addresses, ports, and protocol types. Flow-based statistical features are also extracted for bidirectional flows, such as duration, number of packets, and byte size. These basic and statistical features are then used as network traffic data. It should be understood that the network traffic data presented here is for illustrative purposes only and does not constitute a limitation.
[0048] Next, the network traffic data extracted in the previous step is converted into a format suitable for processing by the large language model. Specifically, the extracted network traffic data is vectorized and mapped to the input space of the large language model through an embedding layer.
[0049] First, by defining a feature template, we convert network traffic features (basic features + statistical features) into a structured natural language description. For example, a feature template might be: "Source IP is {src_ip}, Destination IP is {dst_ip}, Source port is {src_port}, Destination port is {dst_port}, Protocol is {protocol}, Flow duration is {duration} seconds, Total number of packets is {packets}, Total number of bytes is {bytes}, Average packet size is {avg_pkt_size} bytes." Then, by filling in the feature template, we can obtain the corresponding natural language description. For example, "Source IP is 192.168.1.1, Destination IP is 10.0.0.5, Source port is 54321, Destination port is 80, Protocol is TCP, Flow duration is 120 seconds, Total number of packets is 1000, Total number of bytes is 5242880, Average packet size is 5242 bytes." After obtaining the natural language description, it is mapped into a semantic vector through a pre-trained language model (such as BERT, Sentence-BERT) to obtain normal traffic features.
[0050] In one embodiment, after obtaining the normal traffic characteristics, it is necessary to generate abnormal traffic characteristics based on the normal traffic characteristics. Figure 2 , Figure 2 This is a flow chart of generating abnormal traffic features based on the normal traffic features provided by an embodiment of the present application, which specifically includes the following steps: Step 210: Calculate the sliding average corresponding to each characteristic value in the normal flow characteristic based on the preset sliding window, and construct a seasonal series component according to the sliding average.
[0051] In one embodiment, the normal traffic characteristics are represented as: , using the addition principle to perform STL decomposition on the normal traffic characteristics. Among them, STL decomposition decomposes the normal traffic characteristics of the time series into three main components: seasonal sequence component, trend sequence component, and residual component, and takes the residual component as the abnormal traffic characteristics.
[0052] Among them, the seasonal sequence component can reflect the long-term change trend of the data and the overall trend of the data over a long period of time. For the calculation process of the seasonal sequence component, based on a preset sliding window, calculate the sliding average value corresponding to each eigenvalue in the normal traffic characteristics, and form the seasonal sequence component according to the sliding average value. For the i-th eigenvalue in the normal traffic characteristics , its sliding average value is expressed as:
[0053] Among them, k represents the preset sliding window, so the seasonal sequence component is expressed as .
[0054] Step 220: Calculate the difference sequence between the normal traffic characteristics and the seasonal sequence component, and perform locally weighted regression fitting on the difference sequence to obtain the trend sequence component.
[0055] In one embodiment, calculate the difference sequence between the normal traffic characteristics and the trend sequence component, and then perform locally weighted regression fitting on the difference sequence to obtain the trend sequence component . Among them, the process of locally weighted regression fitting LOESS is to select a group of neighboring points around a certain point in the difference sequence, and after weighting the neighboring points, perform local prediction fitting. At this time, the samples (points) closer to the prediction point have larger weights, and the samples (points) farther from the prediction point have smaller weights. A cubic weight function can be used to determine the weight values in the weighting process. Since the prediction of each point is based on its surrounding neighboring points, the obtained trend sequence component can well capture the local trend.
[0056] Step 230: Subtract the seasonal sequence component and the trend sequence component from the normal traffic characteristics to obtain the abnormal traffic characteristics.
[0057] In one embodiment, after removing the seasonal sequence component and the trend sequence component from the normal traffic characteristics, the random fluctuation or noise part is obtained as the residual sequence, and it is used as the abnormal traffic characteristics , which is expressed as:
[0058] In one embodiment, in the abnormal traffic detection scenario of the computing power network, since white noise is a random noise with a mean of 0, a constant variance, and no autocorrelation, it has no task relevance. Therefore, if white noise is used directly, it will lead to the loss of key information. However, the real network traffic may contain local burst traffic (such as short-term high load), protocol-specific behavior (such as TCP retransmission, DNS query delay) or covert attack signals (such as low-frequency DDoS, scanning behavior), etc. Therefore, the embodiment of the present application selects the residual sequence As a reflection of abnormal traffic characteristics corresponding to normal traffic, the residual sequence dynamically reflects the deviation of current traffic from the historical model. It preserves local features of normal traffic that are not explained by the global model, making it more consistent with actual traffic. A sudden increase in the data value in the residual sequence can indicate a sudden anomaly, while a persistent deviation from the mean can indicate a low-frequency attack. This shows that using the residual sequence can better reflect the traffic characteristics of the current task.
[0059] Step 120: Input the normal traffic characteristics into the first traffic encoder for encoding to obtain normal latent space variables, and obtain abnormal latent space variables based on the normal latent space variables and the abnormal traffic characteristics.
[0060] In one embodiment, an abnormal traffic detection network for abnormal traffic detection includes a first traffic encoder and a detection language model. The first traffic encoder may be composed of a Transformer encoder and a fully connected layer. The fully connected layer is located in the feedforward neural network portion of each encoder layer of the Transformer encoder and is primarily used to process features at each position in time series data to achieve dimensional alignment.
[0061] The detection language model can be a GPT-2Medium model. The GPT-2Medium model is a pre-trained large language model based on the Transformer architecture. The structure of the GPT-2Medium model includes at least a multi-head self-attention module, a feed-forward network layer (FFN) and a normalization layer (Layer Normalization, LayerNorm). Taking a Transformer layer as an example, the input of its multi-head self-attention module is the output or word embedding of the previous layer (the first layer). After processing in the multi-head self-attention module, it is sent to the feed-forward network layer through the residual connection and the normalization layer. After processing in the feed-forward network layer, it is sent to the next multi-head self-attention module through the residual connection and the normalization layer. In the GPT-2Medium model, the multi-head self-attention module is the core part of the model and is used to capture the dependencies within the sequence, but the computational overhead of this module is large. Therefore, in the embodiment of the present application, the weights of the multi-head self-attention module in the GPT-2Medium model are frozen, and only the feed-forward network layer and the normalization layer are activated, thereby reducing the training overhead and maintaining acceptable detection accuracy.
[0062] In one embodiment, the normal flow characteristics Input the first flow encoder Encode to obtain normal latent space variables , expressed as:
[0063]
[0064] in, represents the i-th eigenvalue in the normal latent space variable.
[0065] The embodiment of the present application uses the first traffic encoder as an autoencoder to perform latent space mapping on normal traffic features, which can ensure that the mapping result is automatically updated, so that when it is used as a prompt word, it can be adapted to a specific task in real time.
[0066] Next, we get the abnormal latent space variables based on the normal latent space variables. Figure 3 , Figure 3 This is a flowchart of obtaining abnormal latent space variables based on normal latent space variables and abnormal traffic characteristics provided by an embodiment of the present application, which specifically includes the following steps: Step 310: Input the abnormal flow characteristics into the second flow encoder for encoding to obtain abnormal intermediate variables.
[0067] In one embodiment, a second flow encoder having the same model structure as the first flow encoder is set. , the model parameters of both can be obtained according to actual training. Here, the second traffic encoder is used to map the abnormal traffic features to the latent space to obtain an abnormal intermediate variable . And since the structures of the first traffic encoder and the second traffic encoder are the same, the dimensions of the abnormal intermediate variable and the normal latent space variable are the same.
[0068] Step 320: Obtain an abnormal latent space variable according to the sum of the normal latent space variable and the abnormal intermediate variable.
[0069] In one embodiment, the normal latent space variable and the abnormal intermediate variable are added together so that the abnormal latent space variable follows a distribution regarding the abnormal traffic features, and then the abnormal latent space variable can be obtained , which is expressed as:
[0070]
[0071] where represents the i-th value of the abnormal intermediate variable, represents the i-th value of the abnormal latent space variable.
[0072] In one embodiment, referring to Figure 4 , Figure 4 is a schematic diagram of latent space mapping provided by an embodiment of the present application. Figure 4 First, the normal traffic feature is subjected to STL decomposition to obtain a seasonal sequence component , a trend sequence component and an abnormal traffic feature . Then, the normal traffic feature is input into the first traffic encoder for encoding to obtain a normal latent space variable , the abnormal traffic feature is input into the second traffic encoder for encoding to obtain an abnormal intermediate variable , and the normal latent space variable and the abnormal intermediate variable are added together to obtain an abnormal latent space variable .
[0073] The normal latent space variable and the abnormal latent space variable obtained through the above process jointly form the prompt word input for the subsequent detection language model. Compared with the text-based display prompt words, in the embodiment of the present application, the normal latent space variable and the abnormal latent space variable respectively undergo feature extraction by the first traffic encoder and the second traffic encoder, which contain more deep traffic semantic features and can improve the accuracy of anomaly detection.
[0074] Step 130: Input the normal latent space variables and the abnormal latent space variables into the detection language model for reconstruction, obtain normal traffic representation and abnormal traffic representation, and construct positive sample pairs and negative sample pairs based on the normal traffic characteristics, normal traffic representation and abnormal traffic representation.
[0075] In one embodiment, after generating contrasting prompt words based on normal latent space variables and abnormal latent space variables, the contrasting prompt words are used as positive and negative samples to guide the detection language model to generate the desired output. In this process, the detection language model will learn to distinguish the prompt words of normal traffic from the prompt words of abnormal traffic, thereby achieving fine-tuning of the large language model. Because the normal latent space variables and the abnormal latent space variables in the embodiments of the present application can be optimized during the training process and are not fixed prompt words, the contrasting learning effect can be maximized, the difference between the two types of traffic is maximized, and the accuracy of the anomaly detection results is improved.
[0076] Next, the normal latent space variables and abnormal latent space variables are input into the detection language model for reconstruction. Assuming that the detection language model is represented by G(·), the normal traffic representation is reconstructed as and abnormal traffic indication Among them, the normal flow represents the i-th value for: , the i-th value of abnormal traffic for: .
[0077] In one embodiment, since the output dimension of the GPT-2Medium model used as the detection language model is 1024, and the feature number d used to characterize the traffic may not be equal to 1024, the embodiment of the present application also adds a fully connected layer to achieve dimensional conversion, converting the output dimension of the GPT-2Medium model from 1024 to the target feature number d, where d is the dimension of normal traffic features, normal traffic representation, and abnormal traffic representation. At this time, the i-th value of the normal traffic representation is and abnormal traffic represents the i-th value Updated to: and , correspondingly, the normal flow is expressed as , abnormal traffic is expressed as .
[0078] In one embodiment, after having the normal flow representation and the abnormal flow representation, positive and negative samples can be constructed. Figure 5 , Figure 5 This is a flowchart of constructing positive sample pairs and negative sample pairs based on normal traffic characteristics, normal traffic representation, and abnormal traffic representation provided by an embodiment of the present application, specifically including the following steps: Step 510: Use normal traffic features as anchor samples, represent normal traffic as normal samples, and represent abnormal traffic as abnormal samples.
[0079] Step 520: construct a positive sample pair based on the anchor point feature value of the anchor point sample and the normal feature value of the normal sample, and construct a negative sample pair based on the anchor point feature value and the abnormal feature value of the abnormal sample.
[0080] In one embodiment, the normal flow characteristics As an anchor sample, normal traffic represents As a normal sample, abnormal traffic indicates As an abnormal sample. For an anchor point sample, it has v anchor point feature values. According to each anchor point feature value and the normal feature value of the normal sample at the corresponding position and the abnormal feature value of the abnormal sample, a set of positive and negative sample pairs can be obtained.
[0081] Take the i-th anchor point feature value in the anchor point sample For example, the normal feature value of the i-th normal sample is , the i-th abnormal feature value in the abnormal sample is , corresponding to the positive sample pair and negative sample pairs .
[0082] According to the above process, multiple groups of positive sample pairs and negative sample pairs corresponding to the anchor point samples are obtained.
[0083] Step 140: Calculate the function value of the ternary loss function according to the first distance of the positive sample pair and the second distance of the negative sample pair. Based on the function value, at least adjust the weight of the abnormal traffic detection network until a trained abnormal traffic detection network is obtained.
[0084] In one embodiment, referring to Figure 6 , Figure 6 This is a flowchart of calculating the function value of the ternary loss function based on the first distance of the positive sample pair and the second distance of the negative sample pair provided in an embodiment of the present application, which specifically includes the following steps: Step 610: Calculate a first similarity between the anchor point feature value and the normal feature value to obtain a first distance.
[0085] In one embodiment, for a certain anchor point sample, the i-th anchor point feature value The corresponding positive sample pair is , the negative sample pair is First, calculate the positive sample pair Anchor point feature value and normal eigenvalues The first similarity between them, get the first distance Among them, the distance function Used to measure the similarity or difference between positive sample pairs or negative sample pairs.
[0086] It is understandable that clustering samples of the same category, that is, positive sample pairs, together in the feature space can increase the similarity of similar samples, thereby enhancing the detection language model's ability to identify samples of the same category. Therefore, the embodiment of the present application can maximize the first distance to bring positive sample pairs closer to each other in the embedding space.
[0087] Step 620: Calculate the second similarity between the anchor feature value and the abnormal feature value in all negative sample pairs corresponding to the anchor sample, select the negative sample pair corresponding to the minimum value of the second similarity as the minimum negative sample pair, and use the second similarity of the minimum negative sample pair as the second distance.
[0088] In one embodiment, for each anchor point sample, when the identifier i of the anchor point feature value has different values, it corresponds to different negative sample pairs. At this time, all negative sample pairs constitute a negative sample pair set, and the negative sample pair set is expressed as: { }.
[0089] At this time, the second similarity between the anchor point feature value and the abnormal feature value in each negative sample pair in the negative sample pair set is calculated. Taking the i-th negative sample pair as an example, the second similarity is expressed as: Then, the minimum value of the second similarity is selected, the negative sample pair corresponding to the minimum value is used as the minimum negative sample pair, and the second similarity of the minimum negative sample pair is used as the second distance.
[0090] Therefore, the second distance is expressed as:
[0091] It can be understood that since the negative sample pair refers to a sample of a different category from the anchor sample, the embodiment of the present application needs to minimize the second distance so that samples of different categories are further away in the embedding space, thereby increasing the distinction between samples of different categories and improving the detection language model's ability to recognize negative sample pairs.
[0092] In one embodiment, referring to Figure 7 , Figure 7 This is a diagram of data processing for positive and negative samples provided in the embodiment of the present application. As an example, it is input into the detection language model G(·) for data processing to obtain the i-th eigenvalue of normal traffic representation , and then process it with the fully connected layer to get the updated , and the anchor point eigenvalue Constitute a positive sample pair Then, input the distance function Calculate the first distance .
[0093] For the i-th eigenvalue of the abnormal latent space variable , input it into the detection language model G(·) for data processing, and obtain the i-th eigenvalue of abnormal traffic representation , and then process it with the fully connected layer to get the updated , and the anchor point eigenvalue Constitute a negative sample pair Then, input the distance function Calculate the second similarity , and further calculate the second distance based on the second similarity.
[0094] Step 630: Calculate the function value of the ternary loss function according to the first distance and the second distance.
[0095] In one embodiment, referring to Figure 8 , Figure 8 This is a flowchart of calculating the function value of the ternary loss function according to the first distance and the second distance provided in an embodiment of the present application, which specifically includes the following steps: Step 810: Calculate the distance difference between the first distance and the second distance.
[0096] Step 820: Calculate the sum of the distance difference and the preset margin parameter to obtain a loss value. If the loss value is negative, set the loss value to zero.
[0097] In one embodiment, the distance difference is expressed as:
[0098] Assume that the preset margin parameter is m ( ), which is used to ensure that the first distance between positive sample pairs is not only smaller than the minimum distance between negative sample pairs, but also smaller by at least m units, which helps to provide additional separation between positive and negative sample pairs, making positive sample pairs more closely clustered in the embedding space, while negative sample pairs are more dispersed.
[0099] Therefore, the loss value of the ternary loss function is expressed as:
[0100] In one embodiment, to ensure that the loss value is not negative, if the loss value is negative, the loss value is set to zero. Therefore, the loss value L is expressed as:
[0101] The above process uses max(..., 0) to ensure non - negative values.
[0102] In one embodiment, referring to Figure 9 , Figure 9 is a schematic diagram of the calculation process of the triplet loss function provided by the embodiments of the present application. First, the originally extracted normal traffic features are encoded by a first traffic encoder with updatable weights into normal latent space variables regarding normal traffic. At the same time, the normal traffic features are decomposed into residual components strongly related to the detection task to obtain abnormal traffic features. Then, the abnormal traffic features are input into a second traffic encoder for encoding to obtain abnormal intermediate variables, and based on the sum of the normal latent space variables and the abnormal intermediate variables, abnormal latent space variables are obtained. Next, the normal latent space variables and the abnormal latent space variables are respectively input into a detection language model for reconstruction to obtain normal traffic representations and abnormal traffic representations, and then positive sample pairs and negative sample pairs serving as contrast sample pairs are constructed based on the normal traffic features, normal traffic representations, and abnormal traffic representations. Finally, the loss value of the triplet loss function is calculated based on the positive and negative sample pairs as the optimization objective to fine - tune the large language model serving as the detection language model.
[0103] In one embodiment, the triplet loss function consists of three parts: anchor samples, positive samples, and negative samples. The objective of this function is to make the distance between the anchor sample and the positive sample as small as possible, while making the distance between the anchor and the negative sample as large as possible, thereby effectively improving the discrimination ability of the model.
[0104] If, when calculating the function value of the triplet loss function in the embodiments of the present application, only the condition that the distance between the anchor sample and a certain negative sample exceeds the threshold is judged, for the abnormal traffic representations generated due to the interference of the residual components it is too easy to meet this condition, resulting in poor fine - tuning effect of the detection language model and being unable to be applied to the downstream abnormal traffic detection task. Therefore, in the embodiments of the present application, the minimum distance of all negative sample pairs in each anchor sample is emphasized. In the set of negative sample pairs { }, the negative sample pair corresponding to the abnormal feature value with the closest distance to the anchor feature value is selected as the minimum negative sample pair, making the detection language model pay more attention to the difficult - to - distinguish negative samples and improving the discrimination ability of the detection language model.
[0105] It can be seen that in the process of calculating the loss value in the embodiments of the present application, the discrimination difficulty is increased through the minimum negative sample pair, and at the same time, a preset margin parameter m is introduced to optimize the distribution of samples in the embedding space, finally making the positive sample pairs closer and the negative sample pairs farther away, thereby improving the classification performance of the detection language model.
[0106] In one embodiment, once the function value is obtained, the weights of at least the abnormal traffic detection network can be adjusted based on the function value until a trained abnormal traffic detection network is obtained. The training process is described as follows: during training, the model parameters of the multi-head attention module are frozen. Based on the function value, the model parameters corresponding to the first traffic encoder, the second traffic encoder, the feedforward network layer, and the normalization layer are adjusted until an iteration termination condition is met, thereby obtaining a trained abnormal traffic detection network. The iteration termination condition here can be when the number of iterations reaches a preset number or when the optimization performance reaches a preset standard.
[0107] In the above process, the feedforward network layer and normalization layer of the detection language model, as well as the first flow encoder and the second flow encoder are fine-tuned using the ternary loss function, while the fully connected layer used to update the normal flow representation and abnormal flow representation is used. The original random weights are retained and not updated. By adjusting the network weights corresponding to the first and second traffic encoders corresponding to the normal and abnormal traffic representations as prompt words, the abnormal traffic detection network can more flexibly adapt to specific tasks, avoiding the high computational cost of fine-tuning the entire model. Only a small number of prompt word parameters need to be fine-tuned to quickly adapt to multiple tasks, thereby improving the generalization ability of the abnormal traffic detection network.
[0108] It is understandable that after the training is completed, only the first traffic encoder needs to be retained in the abnormal traffic detection network.
[0109] In one embodiment, referring to Figure 10 , Figure 10 This is an overall flow chart of the abnormal traffic detection network training method provided in an embodiment of the present application.
[0110] After obtaining normal traffic features, normal and abnormal traffic features are generated based on them. The normal and abnormal latent space variables are respectively input into the detection language model for reconstruction, resulting in normal and abnormal traffic representations. These representations are then dimensionalized to obtain reconstructed normal and abnormal traffic representations. Next, positive and negative sample pairs are constructed based on the normal traffic features, normal traffic representation, and abnormal traffic representation as comparison samples. Finally, the loss value of the ternary loss function is calculated based on the positive and negative sample pairs as the optimization target to train the abnormal traffic detection network.
[0111] The technical solution provided by the embodiment of the present application generates normal traffic features based on the network traffic data extracted from the computing power network, and generates abnormal traffic features based on the normal traffic features, inputs the normal traffic features into the first traffic encoder for encoding to obtain normal latent space variables, and obtains abnormal latent space variables based on the normal latent space variables and the abnormal traffic features, respectively inputs the normal latent space variables and the abnormal latent space variables into the detection language model for reconstruction to obtain normal traffic representation and abnormal traffic representation, constructs positive sample pairs and negative sample pairs based on the normal traffic features, the normal traffic representation and the abnormal traffic representation, calculates the function value of the ternary loss function based on the first distance of the positive sample pair and the second distance of the negative sample pair, and based on the function value, at least adjusts the weight of the abnormal traffic detection network until a trained abnormal traffic detection network is obtained. The embodiment of the present application trains the abnormal traffic detection network based on comparative prompt word learning for complex and diverse computing power network operation and maintenance scenarios. During the training process, the normal traffic features are first adjusted by the first traffic encoder to adjust the neural network weights and convert them into learnable dynamic prompt words to achieve adaptive matching for specific operation and maintenance tasks. On this basis, a contrastive learning mechanism is introduced to automatically generate and dynamically adjust prompt words. A triplet loss function is used to construct a fine-tuning optimization objective for the detection language model. This enhances the feature separability between normal and abnormal traffic, reduces the reconstruction error between different traffic types, and employs a transfer learning strategy, enabling the detection language model to adapt to new scenarios with fine-tuning of a small number of samples based on pre-training. The resulting anomaly traffic detection network can effectively identify various types of abnormal traffic and meet the anomaly detection needs in diverse network operation and maintenance scenarios.
[0112] In one embodiment, after the abnormal traffic detection network is trained, abnormal traffic monitoring network is used to perform abnormal traffic detection. Figure 11 , Figure 11 This is a flow chart of the abnormal traffic detection method provided by an embodiment of the present application, which specifically includes the following steps: Step 1110: Acquire initial flow data and convert the initial flow data into initial flow features.
[0113] In one embodiment, initial traffic data is obtained in a manner consistent with network traffic data, and the initial traffic data is converted into initial traffic features through a pre-trained language model. The calculation method of the initial traffic features is consistent with that of normal traffic features.
[0114] Step 1120: Input the initial traffic features into the first traffic encoder for feature encoding to obtain latent space representation features.
[0115] In one embodiment, the initial flow characteristics Input the first flow encoder Perform feature encoding to obtain latent space representation features 。
[0116] Step 1130: Input the latent space representation features into the detection language model for reconstruction to obtain the target traffic features, and calculate the reconstruction error between the target traffic features and the latent space representation features. If the reconstruction error is greater than the preset threshold, the detection result indicates that the initial traffic data is abnormal traffic.
[0117] In one embodiment, the detection language model reconstructs based on the latent space representation features to obtain the target traffic features , and the purpose of reconstruction is to make the target traffic features as close as possible to the initial traffic features . At this time, the difference between the target traffic features and the latent space representation features is calculated as the reconstruction error.
[0118] Specifically, during the training process, the training purpose of the detection language model is to make the output normal traffic representation closer and closer to the actual normal traffic features. The closer it is, the more it indicates that its feature extraction ability matches the normal traffic scenario. At this time, if the input is abnormal traffic features, the obtained output should be quite different from the input. Therefore, during the inference process, the reconstruction error can be used to measure the fitting degree of the output of the detection language model to the input. The smaller the reconstruction error, the closer the input and output of the detection language model are, the closer the input data is to the normal traffic scenario, and the detection result is normal traffic data. The larger the reconstruction error, the more deviated the input and output of the detection language model are. Since the detection language model is trained to output data closer to the normal traffic scenario, it can be inferred that the input data is abnormal traffic data.
[0119] In one embodiment, the preset threshold can be determined by the three - standard - deviation principle or the ROC curve. Among them, the ROC curve (Receiver Operating Characteristic Curve) is a tool for evaluating the performance of classification models, and the ROC curve can help select the optimal classifier and the best threshold. Then, the reconstruction error is compared with the preset threshold. If the reconstruction error is greater than the preset threshold, it indicates that the input initial traffic data may be abnormal, and the detection result indicates that the initial traffic data is abnormal traffic. This is because abnormal data usually does not conform to the normal data pattern learned by the detection language model, resulting in a larger reconstruction error. Therefore, the detection of abnormal traffic in the computing power network can be realized, and finally an abnormal detection report is generated.
[0120] It can be seen that the abnormal traffic detection method proposed in the embodiment of the present application combines prompt word fine-tuning with large model technology, which can achieve efficient and accurate anomaly detection. First, the triple loss function based on the minimum negative sample pair is adopted as the optimization target of network fine-tuning, which not only enhances the feature separability of normal traffic and abnormal traffic, but also improves the detection sensitivity by expanding the reconstruction error between categories. Secondly, it integrates the feature extraction advantages of deep neural networks and the generalization ability of large language models, and accurately captures the distribution pattern of normal data through guided learning. This hybrid architecture enables the abnormal traffic detection network to have excellent anomaly recognition capabilities in downstream network operation and maintenance tasks. Not only does it improve detection efficiency and detection accuracy, but it also has good interpretability. It is especially suitable for abnormal traffic detection under complex computing power network architecture.
[0121] The embodiment of the present application also provides an abnormal traffic detection network training device, which can implement the above abnormal traffic detection network training method. Figure 12 , the device comprises: Traffic feature generation module 1210: used to generate normal traffic features based on network traffic data extracted from the computing power network, and generate abnormal traffic features based on the normal traffic features.
[0122] Latent space mapping module 1220: used to input normal traffic features into the first traffic encoder for encoding to obtain normal latent space variables, and obtain abnormal latent space variables based on the normal latent space variables and abnormal traffic features.
[0123] Comparison sample pair generation module 1230: used to input normal latent space variables and abnormal latent space variables into the detection language model for reconstruction, obtain normal traffic representation and abnormal traffic representation, and construct positive sample pairs and negative sample pairs based on normal traffic characteristics, normal traffic representation and abnormal traffic representation.
[0124] Loss adjustment module 1240: used to calculate the function value of the ternary loss function based on the first distance of the positive sample pair and the second distance of the negative sample pair, and based on the function value, at least adjust the weight of the abnormal traffic detection network until a trained abnormal traffic detection network is obtained.
[0125] The specific implementation of the abnormal traffic detection network training device of this embodiment is basically the same as the specific implementation of the abnormal traffic detection network training method described above, and will not be repeated here.
[0126] In one embodiment, the abnormal traffic flow detection method also includes a corresponding abnormal traffic flow detection device, which includes: Initial traffic acquisition module: used to obtain initial traffic data and convert the initial traffic data into initial traffic features.
[0127] Feature encoding module: used to input the initial traffic features into the first traffic encoder for feature encoding to obtain the latent space representation features.
[0128] Reconstruction analysis module: used to input the latent space representation features into the detection language model for reconstruction to obtain the target traffic features, and calculate the reconstruction error between the target traffic features and the latent space representation features. If the reconstruction error is greater than the preset threshold, the detection result indicates that the initial traffic data is abnormal traffic.
[0129] The specific implementation manner of the abnormal traffic detection device in this embodiment is basically the same as that of the above abnormal traffic detection method, and will not be elaborated here.
[0130] This application embodiment also provides an electronic device, including: At least one memory; at least one processor; at least one program; the program is stored in the memory, and the processor executes the at least one program to implement the abnormal traffic detection network training method or the abnormal traffic detection method described above in this application. This electronic device can be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (PDA), an in-vehicle computer, etc.
[0131] Please refer to Figure 13 , Figure 13 which shows the hardware structure of an electronic device in another embodiment. The electronic device includes: Processor 1301, which can be implemented in ways such as a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in this application embodiment; Memory 1302, which can be implemented in forms such as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). Memory 1302 can store an operating system and other application programs. When implementing the technical solutions provided in this specification embodiment through software or firmware, the relevant program codes are stored in memory 1302, and are called by processor 1301 to execute the abnormal traffic detection network training method or the abnormal traffic detection method in this application embodiment; Input / output interface 1303, which is used to implement information input and output; Communication interface 1304, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.); Bus 1305 , which transmits information between various components of the device (e.g., processor 1301 , memory 1302 , input / output interface 1303 , and communication interface 1304 ); The processor 1301 , the memory 1302 , the input / output interface 1303 and the communication interface 1304 are connected to each other in communication within the device via a bus 1305 .
[0132] An embodiment of the present application further provides a storage medium, which is a storage medium that stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned abnormal traffic detection network training method, or the abnormal traffic detection method.
[0133] The memory, as a non-transient storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0134] The abnormal traffic detection network training method, abnormal traffic detection method and device proposed in the embodiments of the present application generate normal traffic features based on the network traffic data extracted from the computing power network, and generate abnormal traffic features based on the normal traffic features. The normal traffic features are input into the first traffic encoder for encoding to obtain normal latent space variables, and abnormal latent space variables are obtained based on the normal latent space variables and the abnormal traffic features. The normal latent space variables and the abnormal latent space variables are respectively input into the detection language model for reconstruction to obtain normal traffic representations and abnormal traffic representations. Positive sample pairs and negative sample pairs are constructed based on the normal traffic features, the normal traffic representations and the abnormal traffic representations. The function value of the triplet loss function is calculated according to the first distance of the positive sample pairs and the second distance of the negative sample pairs. Based on the function value, at least the weights of the abnormal traffic detection network are adjusted until a trained abnormal traffic detection network is obtained. The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0135] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0136] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figures, or combine some steps, or different steps. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0137] Those of ordinary skill in the art will understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or a suitable combination thereof.
[0138] As used in the specification of this application and the above-mentioned drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of this application described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices.
[0139] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may mean: only A exists, only B exists, and both A and B exist simultaneously. Here, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (item) of the following" or a similar expression means any combination of these items, including any combination of single items (items) or plural items (items). For example, at least one (item) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0140] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above-mentioned unit division is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical, or other forms.
[0141] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0142] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0143] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes multiple instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store programs.
[0144] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but are not intended to limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the present invention should be within the scope of the present invention.
Claims
1. A method for training an abnormal traffic detection network, characterized in that The abnormal traffic detection network includes a first traffic encoder and a detection language model, and the method includes: Generating normal traffic features based on the network traffic data extracted from the computing power network, and generating abnormal traffic features based on the normal traffic features; Inputting the normal traffic features into the first traffic encoder for encoding to obtain normal latent space variables, and obtaining abnormal latent space variables according to the normal latent space variables and the abnormal traffic features; Respectively inputting the normal latent space variables and the abnormal latent space variables into the detection language model for reconstruction to obtain normal traffic representations and abnormal traffic representations, and constructing positive sample pairs and negative sample pairs according to the normal traffic features, the normal traffic representations and the abnormal traffic representations; Calculating the function value of the triplet loss function according to the first distance of the positive sample pair and the second distance of the negative sample pair, and based on the function value, at least adjusting the weights of the abnormal traffic detection network until the trained abnormal traffic detection network is obtained.
2. The abnormal traffic detection network training method according to claim 1, wherein The obtaining the abnormal latent space variables according to the normal latent space variables and the abnormal traffic features includes: Inputting the abnormal traffic features into a second traffic encoder for encoding to obtain abnormal intermediate variables, and the model structures of the second traffic encoder and the first traffic encoder are the same; Obtaining the abnormal latent space variables according to the sum of the normal latent space variables and the abnormal intermediate variables.
3. The abnormal traffic detection network training method according to claim 1, wherein, The generating the abnormal traffic features based on the normal traffic features includes: Calculating the moving average corresponding to each eigenvalue in the normal traffic features based on a preset sliding window, and forming a seasonal sequence component according to the moving average; Calculating the difference sequence between the normal traffic features and the seasonal sequence component, and performing locally weighted regression fitting on the difference sequence to obtain a trend sequence component; Subtracting the seasonal sequence component and the trend sequence component from the normal traffic features to obtain the abnormal traffic features.
4. The abnormal traffic detection network training method according to claim 1, wherein The constructing the positive sample pairs and the negative sample pairs according to the normal traffic features, the normal traffic representations and the abnormal traffic representations includes: Taking the normal traffic features as anchor samples, taking the normal traffic representations as normal samples, and taking the abnormal traffic representations as abnormal samples; Constructing the positive sample pairs according to the anchor feature values of the anchor samples and the normal feature values of the normal samples, and constructing the negative sample pairs according to the anchor feature values and the abnormal feature values of the abnormal samples.
5. The abnormal traffic detection network training method according to claim 4, wherein The calculating the function value of the triplet loss function according to the first distance of the positive sample pair and the second distance of the negative sample pair includes: Calculating the first similarity between the anchor feature values and the normal feature values to obtain the first distance; Calculating the second similarity between the anchor feature values and the abnormal feature values in all the negative sample pairs corresponding to the anchor samples, and selecting the negative sample pair corresponding to the minimum value of the second similarity as the minimum negative sample pair, and taking the second similarity of the minimum negative sample pair as the second distance; Calculating the function value of the triplet loss function according to the first distance and the second distance.
6. The abnormal traffic detection network training method according to claim 5, characterized in that Calculating the function value of the triplet loss function according to the first distance and the second distance includes: Calculating the distance difference between the first distance and the second distance; Calculating the sum of the distance difference and a preset margin parameter to obtain the loss value, and setting the loss value to zero if the loss value is negative.
7. The abnormal traffic detection network training method according to claim 2, characterized in that The detection language model at least includes a feed-forward network layer, a normalization layer, and a multi-head attention module. Based on the function value, at least adjusting the weights of the abnormal traffic detection network until the trained abnormal traffic detection network is obtained, including: Freezing the model parameters of the multi-head attention module during training; Based on the function value, adjusting the model parameters corresponding to the first traffic encoder, the second traffic encoder, the feed-forward network layer, and the normalization layer until an iteration termination condition is reached, and obtaining the trained abnormal traffic detection network.
8. An abnormal traffic detection method, characterized in that, Executed by an abnormal traffic monitoring network trained by the abnormal traffic detection network training method according to any one of claims 1 to 7, the method includes: Obtaining initial traffic data and converting the initial traffic data into initial traffic features; Inputting the initial traffic features into the first traffic encoder for feature encoding to obtain latent space representation features; Inputting the latent space representation features into the detection language model for reconstruction to obtain target traffic features, and calculating the reconstruction error between the target traffic features and the latent space representation features. If the reconstruction error is greater than a preset threshold, the detection result indicates that the initial traffic data is abnormal traffic.
9. An abnormal traffic detection network training device, characterized in that The abnormal traffic detection network includes a first traffic encoder and a detection language model, and the device includes: A traffic feature generation module: configured to generate normal traffic features according to network traffic data extracted from a computing power network, and generate abnormal traffic features based on the normal traffic features; A latent space mapping module: configured to input the normal traffic features into the first traffic encoder for encoding to obtain normal latent space variables, and obtain abnormal latent space variables according to the normal latent space variables and the abnormal traffic features; A contrast sample pair generation module: configured to input the normal latent space variables and the abnormal latent space variables into the detection language model for reconstruction respectively to obtain normal traffic representations and abnormal traffic representations, and construct positive sample pairs and negative sample pairs according to the normal traffic features, the normal traffic representations, and the abnormal traffic representations; A loss adjustment module: configured to calculate the function value of the triplet loss function according to the first distance of the positive sample pair and the second distance of the negative sample pair, and based on the function value, at least adjust the weights of the abnormal traffic detection network until the trained abnormal traffic detection network is obtained.
10. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the abnormal traffic detection network training method according to any one of claims 1 to 7, or the abnormal traffic detection method according to claim 8.
11. A storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method for training an abnormal traffic detection network described in any one of claims 1 to 7, or the abnormal traffic detection method described in claim 8.
Citation Information
Patent Citations
Semi-supervised network abnormal behavior detection method based on behavior feature coding
CN113032778A
Industrial Internet of Things anomaly detection method based on space-time variation auto-encoder
CN119004328A
C2F application data anomaly detection model training method, anomaly detection method and related equipment
CN119884759A
Anomaly detection method and apparatus, electronic device, computer readable storage medium, computer program, and computer program product
WO2023015843A1