Abnormal behavior detection method based on TransGAN network

Through a method based on the TransGAN network, using a time sliding window and XGBoost to filter features, combined with deep separable convolution and Transformer, the generator and discriminator are optimized to achieve efficient anomaly detection, solving the problems of dynamic analysis of time series features and global-local feature fusion in existing technologies, and improving the accuracy and robustness of detection.

CN120692076APending Publication Date: 2025-09-23INFORMATION & COMM COMPANY OF QINGHAI ELECTRIC POWER
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510853456.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing technologies have deficiencies in dynamic analysis of time series features, global and local feature fusion, model training stability, and adaptive detection mechanisms, resulting in low accuracy and robustness of anomaly detection.

Method used

A method based on the TransGAN network is adopted. By introducing the time sliding window mechanism and the XGBoost algorithm to screen features, a TransGAN model in which the generator and discriminator work together is constructed. Combining depthwise separable convolution and Transformer, and using joint loss function optimization, a dual-channel detection pipeline is implemented to fusion calculate anomaly scores.

Benefits of technology

It significantly improves the accuracy and robustness of anomaly detection, reduces the false alarm rate, and provides an efficient end-to-end detection solution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120692076A_ABST
    Figure CN120692076A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of network abnormal behavior detection, and particularly relates to an abnormal behavior detection method based on a TransGAN network. The method comprises the following steps: firstly, dynamically evaluating feature importance by utilizing a time sliding window mechanism in combination with a time sequence sensitive XGBoost algorithm, and screening a basic feature set and an attention key feature set through double thresholds; then, a TransGAN model in which a generator and a discriminator cooperatively work is constructed, the generator adopts a position sensing embedding technology to strengthen key feature focusing and introduces a gating mechanism to dynamically adjust basic feature weight, and the discriminator fuses depth separable convolution to extract local features and carries out global distribution comparison with a lightweight Transform decoder; and finally, through a dual-channel detection assembly line and a reconstruction channel, calculating residual features of input data and generating output, judging channel generation anomaly confidence, and fusing the two to form a comprehensive anomaly score to realize accurate judgment. The accuracy, robustness and real-time performance of anomaly detection are improved, and the method is suitable for various anomaly detection scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of network abnormal behavior detection, and specifically relates to an abnormal behavior detection method based on a TransGAN network. Background Art

[0002] In fields such as computer vision, network security, and industrial monitoring, abnormal behavior detection, as a core technology for risk early warning, remains a research priority, with its accuracy and robustness remaining a key focus. While traditional machine learning-based detection methods demonstrate certain advantages in static feature analysis, they suffer from significant drawbacks when processing time series data. For one thing, their feature importance assessment mechanisms lack dynamic time series awareness and are unable to capture sudden changes in feature importance over time, leading to delayed detection of short-term anomalies. Furthermore, traditional algorithms typically employ global feature averaging, making it difficult to distinguish the temporal correlations of features within different time windows. For example, in network attack detection, they are unable to effectively identify multi-stage attack feature combinations based on time series.

[0003] While deep learning-based methods have achieved breakthroughs in feature representation capabilities, they still face multiple technical bottlenecks. Traditional GAN ​​models are prone to "mode collapse" during training, making it difficult for the generator to cover the full distribution of real data, resulting in increased false positive rates during anomaly detection. Furthermore, their discriminators rely solely on convolutional layers to extract local features and lack the ability to model global feature dependencies, leading to significant fluctuations in detection accuracy in complex scenarios. Furthermore, the loss function design of existing deep learning models often focuses on a single dimension, adversarial loss, lacking constraints on data reconstruction accuracy. This makes it difficult to quantify the detailed differences between generated and real data, further impacting the reliability of anomaly determination.

[0004] In summary, existing technologies have significant deficiencies in dynamic analysis of time series features, fusion of global and local features, model training stability, and adaptive detection mechanisms. There is an urgent need to improve the accuracy and robustness of abnormal behavior detection through innovations in algorithm architecture and detection processes. Summary of the Invention

[0005] To address the problems of insufficient sensitivity of traditional machine learning algorithms in dynamic recognition of temporal features, low efficiency of global and local feature fusion in deep learning models, proneness to mode collapse and gradient instability during training, and the inability of fixed threshold detection mechanisms to adapt to changes in network status, the present invention provides a method for detecting abnormal behavior based on the TransGAN network. This method obtains basic and key feature sets through dynamic feature screening, constructs a generator that integrates position-aware embedding and gating mechanisms, and a discriminator that combines deep separable convolution with the Transformer. Finally, a dual-channel detection pipeline is used to fuse residual features with discrimination confidence to achieve accurate anomaly determination. This method improves the accuracy, robustness, and real-time performance of anomaly detection and is applicable to a variety of anomaly detection scenarios.

[0006] The technical solutions of the present invention are as follows:

[0007] The technical solution of the present invention is to provide an abnormal behavior detection method based on the TransGAN network, comprising:

[0008] The time-sensitive XGBoost algorithm with a time sliding window mechanism is used to analyze the original network data, dynamically identify key features, and screen out the basic feature set and the attention key feature set;

[0009] Build a TransGAN model where the generator and discriminator work together;

[0010] The generator uses a feature-enhanced Transformer to initialize the key feature set of the attention module into the position-aware embedding vector, allowing the model to prioritize the key path. The generator input layer introduces a feature selection gating mechanism for the basic feature set and outputs it through the generator output layer.

[0011] The discriminator integrates convolution and self-attention mechanisms. The front end uses depth-wise separable convolution to extract local features, and the back end uses a lightweight Transformer decoder to perform global distribution comparison.

[0012] A joint loss function is used to guide TransGAN model optimization, which includes adversarial loss, residual loss, and gradient penalty.

[0013] Among them, the adversarial loss drives the generator to generate realistic samples and optimizes the discriminator's discrimination ability; the residual loss measures the similarity between the generated data and the real data; the gradient penalty term stabilizes the training process;

[0014] The data to be detected is input into the trained TransGAN model, and the comprehensive anomaly score is calculated by fusing the dual-channel detection pipeline of the reconstruction channel and the discrimination channel to determine abnormal behavior.

[0015] As a further optimization option of this method, in the time sliding window mechanism, the size of the time sliding window and the sliding step are set according to the dynamic change characteristics of the network data flow to ensure the continuity of the time window and the ability to capture dynamic changes in data, where the size of the time sliding window is τ and the sliding step is s, satisfying τ>s to form overlapping windows.

[0016] As a further optimization option of this method, the XGBoost algorithm calculates the feature importance score by combining the split gain and the feature occurrence frequency, and obtains the feature importance score through weighted calculation. The feature importance score calculation formula is:

[0017] I(f)=α·Gain(f)+(1-α)·Frequency(f);

[0018] Among them, α∈[0,1] is the weight coefficient, Gain(f) is the split gain, and Frequency(f) is the feature occurrence frequency.

[0019] As a further optimization of this method, in the feature-enhanced Transformer architecture of the generator, the high-importance features of the attention key feature set are initialized through a position-aware embedding mechanism, so that the high-importance features are given higher weights in the embedding space, thereby guiding the model to give priority to these features during the generation process; the attention key feature set is F attention ={f1,f2,...,f k}, whose feature importance score is I(f i ), then the position-aware embedding vector The calculation formula is:

[0020] P i =PE(f i )·I(f i );

[0021] Among them, PE(f i ) is the standard positional encoding function, and d is the embedding dimension.

[0022] As a further optimization option of this method, the feature selection gating mechanism of the generator includes: base Each feature f j ∈F base Calculate the gating weights and weight the input features by gating weights, dynamically suppress redundant features and enhance the focus on key features;

[0023] Among them, for the basic feature set F base Each feature f j ∈F base , calculate its gating weight wj The formula is:

[0024] w j =σ(W g ·f j +b g );

[0025] Where W g and b g is a learnable parameter, σ is the Sigmoid activation function; f j is the jth feature.

[0026] As a further optimization option of this method, the discriminator adopts a convolution and self-attention fusion architecture, in which the depth-wise separable convolution reduces the computational complexity by decomposing the standard convolution into depth-wise convolution and point-wise convolution, while retaining the local feature extraction capability. The depth-wise convolution is used to extract local features channel by channel, and the point-wise convolution is used to adjust the feature dimension.

[0027] As a further optimization option of this method, the adversarial loss adopts the improved WGAN-GP framework and combines the idea of ​​least squares loss to solve the gradient vanishing problem of traditional GAN ​​through continuous gradient signals. It also improves the discriminator's sensitivity to abnormal patterns by minimizing the mean square error between generated samples and real data. The basic form of the adversarial loss is:

[0028] L adv =D(x real )-D(x fake );

[0029] Among them, L adv To combat the loss value, D(x real ) is the output value of the discriminator for the real data; D(x fake ) is the output value of the discriminator for the fake data generated by the generator;

[0030] The residual loss adopts a multi-scale residual metric, a weighted combination of L1 loss, L2 loss and SSIM, to balance the similarity between local and global features, ensuring that the generated samples are highly consistent with the real data at both the pixel level and the structure level;

[0031] The gradient penalty term adopts the gradient penalty strategy in WGANGP and combines it with the local smoothness constraint. The calculation formula of the gradient penalty term is:

[0032]

[0033] in, is a sample obtained by randomly interpolating between the real data and the generated data, Interpolate samples for the discriminator The gradient of , ‖·‖2 is the L2 norm of the gradient, and λ is the penalty coefficient.

[0034] As a further optimization option of this method, the method for generating the comprehensive anomaly score includes: combining the residual feature vector with the discriminant confidence through weighted summation, attention mechanism or neural network fusion to generate a comprehensive anomaly score, wherein the calculation formula of the residual feature vector is:

[0035]

[0036] Among them, r is the residual vector, x test is the original input data, To reconstruct data

[0037] The calculation formula for the discrimination confidence is:

[0038] C(x)=σ(D(x));

[0039] Where C(x) is the discrimination confidence, σ(·) is the Sigmoid function, and D(x) is the original output value of the discriminator for the input data x;

[0040] The comprehensive anomaly score combines the residual feature vector with the discriminant confidence using weighted summation, attention mechanism or neural network fusion.

[0041] As a further optimization option of this method, the abnormal behavior determination logic is: setting an abnormality score threshold, when the comprehensive abnormality score of the input data exceeds the threshold, it is determined to be abnormal behavior; otherwise, it is determined to be normal traffic.

[0042] The beneficial effects brought about by the technical solutions provided in the embodiments of the present application include at least the following:

[0043] Beneficial effects:

[0044] The present invention significantly improves the accuracy of anomaly detection through the dynamic screening of timing-sensitive features and the collaborative optimization of the TransGAN architecture: the generator's position-aware embedding and gating mechanism focus on key feature paths, and the discriminator fuses local feature extraction with global distribution comparison to enhance complex pattern recognition capabilities; the model adopts a joint loss function (adversarial loss, multi-scale residual and gradient penalty) to improve training stability and reduce computational complexity; the dual-channel detection pipeline synchronously captures local residual features and global discrimination confidence, outputs a comprehensive score after fusion and combines it with dynamic threshold judgment, achieving efficient end-to-end detection while ensuring a low false alarm rate, providing a highly robust solution for real-time monitoring of large-scale networks. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 This is a schematic diagram of the overall process of the abnormal behavior detection method based on the TransGAN network;

[0046] Figure 2 Detailed flowchart of step S100 of the abnormal behavior detection method based on the TransGAN network;

[0047] Figure 3 Detailed flowchart of step S200 of the abnormal behavior detection method based on the TransGAN network;

[0048] Figure 4 Detailed flowchart of step S300 of the abnormal behavior detection method based on the TransGAN network;

[0049] Figure 5 Detailed flowchart of step S400 of the abnormal behavior detection method based on the TransGAN network. DETAILED DESCRIPTION

[0050] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this application.

[0051] See also Figure 1 , which shows an abnormal behavior detection method based on a TransGAN network provided by an embodiment of the present invention, the method comprising:

[0052] S100: The time-sensitive XGBoost algorithm with a time sliding window mechanism is used to analyze the original network data, dynamically identify key features, and screen out the basic feature set and the attention key feature set.

[0053] S200: Constructing a TransGAN model where the generator and discriminator work together:

[0054] The generator uses a feature-enhanced Transformer to initialize the key feature set of the attention module into the position-aware embedding vector, allowing the model to prioritize the key path. The generator input layer introduces a feature selection gating mechanism for the basic feature set and outputs it through the generator output layer.

[0055] The discriminator integrates convolution and self-attention mechanisms. The front end uses depth-wise separable convolution to extract local features, and the back end uses a lightweight Transformer decoder to perform global distribution comparison.

[0056] S300: A joint loss function is used to guide TransGAN model optimization, which includes three parts: adversarial loss, residual loss, and gradient penalty. The adversarial loss drives the generator to generate realistic samples and optimizes the discriminator's discrimination ability; the residual loss measures the similarity between the generated data and the real data; and the gradient penalty stabilizes the training process.

[0057] S400: Input the data to be detected into the trained TransGAN model, calculate the comprehensive anomaly score through the dual-channel detection pipeline fusion of the reconstruction channel and the discrimination channel, and determine the abnormal behavior.

[0058] The specific plan is as follows:

[0059] In an abnormal behavior detection method based on the TransGAN network, S100 performs feature analysis on the original network data through the improved XGBoost algorithm, and dynamically screens key features in combination with the time sliding window mechanism to form a basic feature set and an attention key feature set, providing high-quality input for subsequent model training.

[0060] Please refer to Figure 2 , which shows a flowchart of an exemplary abnormal behavior detection method S100 based on the TransGAN network of the present application, including:

[0061] S110: Design a time sliding window mechanism.

[0062] The time sliding window divides the continuous network data stream into multiple overlapping time windows. Each window contains a certain amount of historical data and slides forward as new data flows in. This captures the dynamic changes of the data and improves the time series sensitivity of feature extraction.

[0063] Specifically, the parameters of the time sliding window include: the size τ of the time sliding window and the sliding step s of the time sliding window. τ is made larger than s to form overlapping windows.

[0064] Under the time sliding window mechanism, the original network data is divided into multiple consecutive time windows, and feature extraction is performed in each window.

[0065] S120: Use the XGBoost algorithm to evaluate the importance of original network data features.

[0066] Based on the time sliding window mechanism, the XGBoost algorithm is used to evaluate the importance of original network data features and dynamically identify the features that have the greatest impact on abnormal behavior detection.

[0067] In one possible implementation, the main method for the XGBoost algorithm to calculate feature importance includes split gain and feature occurrence frequency.

[0068] The splitting gain measures the contribution of a feature to the model's predictions during the decision tree splitting process, i.e., the information gain that the feature brings when splitting the node. Specifically, when building a decision tree, XGBoost calculates the sum of the gains of each feature across all splitting nodes and uses this to assess the importance of the feature.

[0069] Among them, the feature occurrence frequency measures the number of times a feature is selected as a split node in all decision trees, reflecting the frequency of use of the feature in the entire model.

[0070] XGBoost calculates the feature importance score by the split gain and frequency of each feature.

[0071] In one possible implementation, combining split gain and frequency, the feature importance score calculation formula is:

[0072] I(f)=α·Gain(f)+(1-α)·Frequency(f);

[0073] Among them, α∈[0,1] is the weight coefficient, usually α=0.5 to balance the influence of gain and frequency, Gain(f) is the splitting gain, and Frequency(f) is the feature occurrence frequency.

[0074] S130: Screen the features based on the set threshold values ​​θ1 and θ2 to construct a basic feature set and an attention key feature set.

[0075] Set thresholds θ1 and θ2, where θ1 < θ2. By comparing the feature importance scores with the thresholds θ1 and θ2, the basic feature set and the attention-critical feature set are screened out.

[0076] Specifically, the base feature set includes all features with a feature importance score above a lower threshold θ1. Features with a lower threshold θ1 contribute to abnormal behavior detection overall and can serve as the basic input of the model. The attention-critical feature set consists of features with a feature importance score above a higher threshold θ2. Features above the higher threshold θ2 have stronger discriminative power for identifying abnormal behavior within a specific time window and will be given higher attention in subsequent TransGAN models.

[0077] In an abnormal behavior detection method based on the TransGAN network, S200 is based on the TransGAN model of feature-enhanced Transformer generator and convolutional self-attention fusion discriminator. Through dynamic feature focusing and multi-scale local-global feature analysis, it achieves high-precision and robust detection of network abnormal behavior.

[0078] Please refer to Figure 3 , which shows a flowchart of an exemplary abnormal behavior detection method S200 based on the TransGAN network of the present application, including:

[0079] S210: Generator Design — Feature-Enhanced Transformer Architecture.

[0080] In the construction of the TransGAN model, the generator generates realistic network data samples to assist the discriminator in learning the distribution characteristics of normal traffic.

[0081] In one possible implementation, a feature-enhanced Transformer is used as the basic architecture of the generator, and a position-aware embedding mechanism for the attention key feature set and a feature selection gating mechanism for the basic feature set are introduced to enhance the model's ability to focus on the key path and optimize feature utilization efficiency.

[0082] Specifically, the location-aware embedding mechanism design includes:

[0083] Let the key feature set of attention be F attention ={f1,f2,...,f k}, whose feature importance score is I(f i ), then the position-aware embedding vector The calculation formula is:

[0084] P i =PE(f i )·I(f i );

[0085] Where PE(f i ) is the standard positional encoding function, and d is the embedding dimension. Through this formula, the high-importance features of the attention key feature set are given higher weights in the embedding space, thereby guiding the model to prioritize these features during the generation process.

[0086] Specifically, the feature selection gating mechanism includes:

[0087] Gating weight calculation: For the basic feature set F base Each feature f j ∈F base , calculate its gating weight w j :

[0088] w j =σ(W g ·f j +b g );

[0089] Where W g and b gis a learnable parameter and σ is the Sigmoid activation function.

[0090] Weighted Input Features: Weight the input features:

[0091] x′ j =w j ·f j ;

[0092] Through this mechanism, the model can dynamically suppress redundant features and enhance attention to key features.

[0093] In one possible implementation, the main structure of the generator consists of multiple layers of Transformer blocks, each of which includes a multi-head self-attention module and a feedforward neural network. To enhance feature representation, a feature enhancement module is introduced within the multi-head self-attention module. This module adjusts the method for determining attention weights based on the weight information of the key feature set.

[0094] Through the above design, the generator can dynamically focus on key feature paths during the generation process, and at the same time optimize the utilization efficiency of basic features through the gating mechanism, thereby improving the quality of generated samples and the ability to model normal traffic distribution.

[0095] S220: Discriminator design - convolution and self-attention fusion architecture.

[0096] The discriminator is used to distinguish samples generated by the generator from real network data and capture abnormal patterns through global distribution comparison.

[0097] In one possible implementation, the discriminator uses a convolutional and self-attention fusion architecture, extracting local features through depthwise separable convolutions and then combining them with a lightweight Transformer decoder for global distribution comparison, thereby achieving efficient and accurate anomaly detection. After global distribution comparison, a feature fusion layer is used to integrate local and global features, and the discrimination result is output through a fully connected layer.

[0098] Specifically, depthwise separable convolution reduces computational complexity by decomposing standard convolution into depthwise convolution and pointwise convolution, while retaining the ability to extract local features. The lightweight Transformer decoder consists of a multi-head self-attention module and a cross-attention module to compare the global feature distribution differences between generated samples and real data.

[0099] Through the above design, the discriminator can capture local and global features at the same time at a low computational cost, thereby improving the ability to discriminate abnormal behaviors.

[0100] In an abnormal behavior detection method based on the TransGAN network, S300 constructs an optimization objective by combining adversarial loss, residual loss, and gradient penalty terms, enabling the Trans-GAN model to maintain training stability while generating realistic samples, thereby improving the robustness and accuracy of anomaly detection.

[0101] Please refer to Figure 4 , which shows a flowchart of an exemplary abnormal behavior detection method S300 based on the TransGAN network of the present application, which includes:

[0102] S310: Design adversarial loss to optimize the generator and discriminator.

[0103] Adversarial loss is the core mechanism driving the collaborative optimization of the generator and discriminator. It enables the generator to produce realistic samples while enhancing the discriminator's ability to distinguish between real and generated data. Adversarial loss employs a modified WassersteinGAN with Gradient Penalty (WGAN-GP) framework, combined with the concept of least squares loss, to improve training stability and detection accuracy.

[0104] In one possible implementation, the basic form of the adversarial loss is as follows:

[0105] L adv =D(x real )-D(x fake );

[0106] Where D(x real ) is the output value of the discriminator for the real data, and the goal is to maximize this value, that is, to make the discriminator think that the real data is real; D(x fake ) is the discriminator's output value for fake data generated by the generator. The goal is to minimize this value, which means that the discriminator believes the generated data is fake. Through adversarial training, the generator attempts to minimize this loss to generate more realistic data, while the discriminator attempts to maximize this loss to better distinguish between real and fake data.

[0107] As an optimization of the above implementation, to further enhance the robustness of the model, the adversarial loss combines the Wasserstein distance and the least squares loss. The Wasserstein distance solves the vanishing gradient problem of traditional GANs by providing a continuous gradient signal, while the least squares loss minimizes the mean squared error between generated samples and real data, making the discriminator more sensitive to abnormal patterns.

[0108] S320: Design residual loss to measure the similarity between generated data and real data.

[0109] The residual loss is used to measure the similarity between the samples generated by the generator and the real data, ensuring that the generator can accurately capture the distribution characteristics of normal traffic.

[0110] In one possible implementation, the residual loss adopts a multi-scale residual metric, combined with L1 loss, L2 loss, and structural similarity index (SSIM), to improve the accuracy of anomaly detection.

[0111] Among them, L1 loss is used to measure the absolute difference between generated data and real data at the pixel level or feature level, and is suitable for the detection of sparse anomalies.

[0112] L2 loss is used to measure the overall distribution fitting accuracy between generated data and real data, and is suitable for the detection of dense anomalies.

[0113] The structural similarity index (SSIM) is a similarity measure based on a local window that captures the spatial correlation of the data and ensures that the generated samples are highly consistent with the real data in structure.

[0114] In one possible implementation, a multi-scale residual loss balances the similarity between local and global features by weightedly combining the three metrics described above. By optimizing this loss function, the generator generates samples that are highly similar to real data at both the pixel and structural levels, thereby improving the accuracy of anomaly detection.

[0115] S330: Design gradient penalty terms to ensure the stability of the training process.

[0116] The gradient penalty term is used to prevent the gradient of the discriminator from exploding or disappearing, ensuring the stability of the training process.

[0117] In one possible implementation, the gradient penalty term adopts the gradient penalty strategy in WGANGP and combines it with a local smoothness constraint to further improve the robustness of the model.

[0118] Based on the above implementation, the calculation formula of the gradient penalty term is:

[0119]

[0120] in, is a sample obtained by randomly interpolating between real data and generated data, used to constrain the gradient of the discriminator, Interpolate samples for the discriminator The gradient of , which measures the impact of input changes on the output, ‖·‖2 is the L2 norm of the gradient, that is, the length of the vector, indicating the magnitude of the gradient; λ is the penalty coefficient.

[0121] S340: The optimization of the generator and the discriminator adopts an alternating update strategy.

[0122] The optimization of the generator and the discriminator adopts an alternating update strategy:

[0123] Fix the generator and update the discriminator: optimize the discriminator parameters by gradient ascent, maximize the adversarial loss and minimize the gradient penalty term.

[0124] Fix the discriminator and update the generator: optimize the generator parameters by gradient descent to minimize the adversarial loss and residual loss.

[0125] In an abnormal behavior detection method based on TransGAN network, S400

[0126] Please refer to Figure 5 , which shows a flowchart of an exemplary abnormal behavior detection method S400 based on the TransGAN network of the present application, including:

[0127] S410: Construct a reconstruction channel and calculate the residual feature vector based on the generator output.

[0128] After the TransGAN model training is completed, the reconstruction channel uses the generator to reconstruct the input data to generate the corresponding reconstructed data.

[0129] Specifically, the data to be tested first passes through a feature selection gating mechanism to suppress redundant features and enhance the influence of key features. Subsequently, the input data enters a multi-head self-attention module, which, combined with position-aware embedding vectors, enables the generator to prioritize high-importance features during the reconstruction process. Finally, the generator outputs the reconstructed data, which should be in the same form as the input data to facilitate the subsequent calculation of the residual feature vector.

[0130] In one possible implementation, the calculation of the residual feature vector is based on the difference between the reconstructed data output by the generator and the original input data. The residual calculation formula is:

[0131]

[0132] Among them, r is the residual vector, which reflects the reconstruction error of the generator on the input data, x test is the original input data, To reconstruct the data.

[0133] S420: Construct a discrimination channel and generate discrimination confidence based on the discriminator output.

[0134] In the TransGAN model, the discriminant channel uses a discriminator to distinguish between samples generated by the generator and real network data, and detects anomalous patterns through global distribution comparison. After training, the discriminator can be used to assess the abnormality of the input data. The discriminant output reflects the degree to which the data matches the normal traffic distribution. A higher confidence level indicates that the input data is closer to normal traffic patterns; conversely, a lower confidence level may indicate abnormal behavior in the data.

[0135] The reasoning process of the discriminant channel is:

[0136] During the inference phase, the data to be tested is first fed into the discriminator's depthwise separable convolutional module to extract local features. These local features are then fed into a lightweight Transformer decoder to capture the data's global feature distribution and compare it with the distribution of normal traffic. Finally, the discriminator's output layer integrates the local and global features through a fully connected layer and generates a scalar value, the confidence level. This confidence level typically ranges from [0, 1], where values ​​close to 1 indicate a high match with normal traffic, while values ​​close to 0 indicate that the data may deviate from the normal pattern and have a high probability of being abnormal.

[0137] In a possible implementation, the calculation formula for the judgment confidence is:

[0138] C(x)=σ(D(x));

[0139] Where C(x) is the discrimination confidence, which indicates the degree of matching between the input data x and the normal traffic distribution, and its value range is [0, 1]. σ(·) is the Sigmoid function, which is used to normalize the original output of the discriminator to the interval [0, 1]. D(x) is the original output value of the discriminator for the input data x, reflecting the global matching degree between the data and the normal traffic.

[0140] If the judgment confidence is high, it means that the global feature distribution of the input data is similar to normal traffic, and it is therefore judged to be normal behavior; if the judgment confidence is low, it means that the global feature distribution of the input data is significantly different from normal traffic, and it may be abnormal behavior.

[0141] S430: Design a dual-channel feature fusion mechanism to integrate the residual feature vector and the discriminant confidence to generate a comprehensive anomaly score.

[0142] After calculating the reconstruction and discrimination channels, the residual feature vector is fused with the discrimination confidence to generate a comprehensive anomaly score. This fusion mechanism leverages information from both local feature deviations and global distribution differences to improve the accuracy and robustness of anomaly detection. Because the residual feature vector reflects the reconstruction error between the input data and normal traffic in local features, while the discrimination confidence measures the degree of match between the input data and the global feature distribution, the combination of the two provides a more comprehensive basis for anomaly determination.

[0143] In a possible implementation, the residual feature vector and the discrimination confidence can be fused using a variety of methods such as weighted summation, attention mechanism, or neural network fusion.

[0144] For example, a neural network fusion method is used, through a multi-layer perceptron or Transformer architecture, to take the residual feature vector and the discriminant confidence as input, and output a comprehensive anomaly score after nonlinear transformation.

[0145] S440: Determine abnormal behavior based on the comprehensive abnormality score and threshold.

[0146] The comprehensive anomaly score reflects the degree to which the input data matches normal traffic patterns. A higher score indicates a greater likelihood that the data deviates from the normal pattern. To make the final determination, a threshold for the anomaly score is set. When the comprehensive anomaly score of the input data exceeds this threshold, it is considered abnormal behavior; otherwise, it is considered normal traffic.

[0147] In a possible implementation, the threshold is set based on statistical characteristics of historical data.

[0148] The mean and standard deviation of the comprehensive anomaly score for normal traffic are optimized in combination with the false alarm rate control strategy. If you want to control the false alarm rate within 5%, you can set the threshold to the mean + 2σ to ensure that only 5% of normal data is misclassified as anomaly.

[0149] The basic principles of the present application have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, and effects mentioned in this application are merely illustrative and not restrictive, and it should not be assumed that these advantages, strengths, and effects are required of each embodiment of this application. In addition, the specific details disclosed above are merely illustrative and facilitating understanding, and are not restrictive. The above details do not limit this application to necessarily being implemented using the above specific details.

[0150] The block diagrams of the devices, devices, equipment, and systems involved in this application are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As will be appreciated by those skilled in the art, these devices, devices, equipment, and systems can be connected, arranged, or configured in any manner. Words such as "include," "comprise," "have," and the like are open-ended words, meaning "including but not limited to," and can be used interchangeably therewith. The words "or" and "and" used herein refer to the words "and / or" and can be used interchangeably therewith, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to," and can be used interchangeably therewith.

[0151] It should also be noted that in the apparatus, device, and method of the present application, each component or each step can be decomposed and / or recombined, and such decomposition and / or recombination should be regarded as equivalent solutions of the present application.

[0152] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present application. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present application. Therefore, the present application is not intended to be limited to the aspects shown herein, but rather to be applied in the widest sense consistent with the principles and novel features of the present invention.

[0153] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. A method for detecting abnormal behavior based on a TransGAN network, characterized in that: include: The time-sensitive XGBoost algorithm with a time sliding window mechanism is used to analyze the original network data, dynamically identify key features, and screen out the basic feature set and the attention key feature set; Build a TransGAN model where the generator and discriminator work together; The generator uses a feature-enhanced Transformer to initialize the key feature set of the attention module into the position-aware embedding vector, allowing the model to prioritize the key path. The generator input layer introduces a feature selection gating mechanism for the basic feature set and outputs it through the generator output layer. The discriminator integrates convolution and self-attention mechanisms. The front end uses depthwise separable convolution to extract local features, and the back end uses a lightweight Transformer decoder to perform global distribution comparison. A joint loss function is used to guide TransGAN model optimization, which includes adversarial loss, residual loss, and gradient penalty. Among them, the adversarial loss drives the generator to generate realistic samples and optimizes the discriminator's discrimination ability; the residual loss measures the similarity between the generated data and the real data; the gradient penalty term stabilizes the training process; The data to be detected is input into the trained TransGAN model, and the comprehensive anomaly score is calculated by fusing the dual-channel detection pipeline of the reconstruction channel and the discrimination channel to determine abnormal behavior.

2. The abnormal behavior detection method based on the TransGAN network according to claim 1 is characterized in that: In the time sliding window mechanism, the size of the time sliding window and the sliding step are set according to the dynamic change characteristics of the network data flow to ensure the continuity of the time window and the ability to capture dynamic changes in data. The size of the time sliding window is τ, and the sliding step is s, satisfying τ>s to form overlapping windows.

3. The abnormal behavior detection method based on the TransGAN network according to claim 1 is characterized in that: When the XGBoost algorithm calculates the feature importance score, it combines the split gain and the feature occurrence frequency to obtain the feature importance score through weighted calculation. The feature importance score calculation formula is: I(f)=α·Gain(f)+(1-α)·Frequency(f); Among them, α∈[0,1] is the weight coefficient, Gain(f) is the split gain, and Frequency(f) is the feature occurrence frequency.

4. The abnormal behavior detection method based on the TransGAN network according to claim 1 is characterized in that In the generator’s feature-enhanced Transformer architecture, high-importance features of the attention-critical feature set are initialized through a position-aware embedding mechanism, which gives high-importance features higher weights in the embedding space, thereby guiding the model to prioritize these features during generation. in The key feature set of attention is F attention ={f1,f2,...,f k }, whose feature importance score is I(f i ), then the position-aware embedding vector The calculation formula is: Q i =PE(f i )·I(f i ); Among them, PE(f i ) is the standard positional encoding function, and d is the embedding dimension.

5. The abnormal behavior detection method based on the TransGAN network according to claim 1 is characterized in that: The feature selection gating mechanism of the generator includes: base Each feature f j ∈F base Calculate the gating weights and weight the input features by gating weights, dynamically suppress redundant features and enhance the focus on key features; Among them, for the basic feature set F base Each feature f j ∈F base , calculate its gating weight w j The formula is: w j =σ(W g ·f j +b g ); Where W g and b g is a learnable parameter, σ is the Sigmoid activation function; f j is the jth feature.

6. The abnormal behavior detection method based on the TransGAN network according to claim 1, characterized in that: The discriminator adopts a convolution and self-attention fusion architecture, in which the depthwise separable convolution reduces the computational complexity while retaining the local feature extraction capability by decomposing the standard convolution into depthwise convolution and pointwise convolution. The depthwise convolution is used to extract local features channel by channel, and the pointwise convolution is used to adjust the feature dimension.

7. The abnormal behavior detection method based on the TransGAN network according to claim 1, characterized in that: The adversarial loss adopts the improved WGAN-GP framework and combines the idea of ​​least squares loss to solve the gradient vanishing problem of traditional GAN ​​through continuous gradient signals. It also improves the discriminator's sensitivity to abnormal patterns by minimizing the mean square error between generated samples and real data. The basic form of the adversarial loss is: L adv =D(x real )-D(x fake ); Among them, L adv To combat the loss value, D(x real ) is the output value of the discriminator for the real data; D(x fake ) is the output value of the discriminator for the fake data generated by the generator; The residual loss adopts a multi-scale residual metric, a weighted combination of L1 loss, L2 loss and SSIM, to balance the similarity between local and global features, ensuring that the generated samples are highly consistent with the real data at both the pixel level and the structure level; The gradient penalty term adopts the gradient penalty strategy in WGANGP and combines it with the local smoothness constraint. The calculation formula of the gradient penalty term is: in, is a sample obtained by randomly interpolating between the real data and the generated data, Interpolate samples for the discriminator The gradient of , ‖·‖2 is the L2 norm of the gradient, and λ is the penalty coefficient.

8. The abnormal behavior detection method based on the TransGAN network according to claim 1, characterized in that: The method for generating the comprehensive anomaly score includes: combining the residual feature vector and the discrimination confidence by weighted summation, attention mechanism or neural network fusion to generate the comprehensive anomaly score, wherein the calculation formula of the residual feature vector is: Among them, r is the residual vector, x test is the original input data, To reconstruct data The calculation formula for the discrimination confidence is: C(x)=σ(D(x)); Where C(x) is the discrimination confidence, σ(·) is the Sigmoid function, and D(x) is the original output value of the discriminator for the input data x; The comprehensive anomaly score combines the residual feature vector with the discriminant confidence using weighted summation, attention mechanism or neural network fusion.

9. The abnormal behavior detection method based on the TransGAN network according to claim 1, characterized in that: The abnormal behavior determination logic is: set an abnormality score threshold, and when the comprehensive abnormality score of the input data exceeds the threshold, it is determined to be abnormal behavior; otherwise, it is determined to be normal traffic.

Citation Information

Cited By

  • Bulk product source flow direction modeling method based on point source list and Beidou trajectory

    CN121544144A