Group-aware diffusion data augmentation method for solving data imbalance problem
Patent Information
- Application Number
- CN202610841657.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-11
- Publication Date
- 2026-09-25
AI Technical Summary
[0006]本发明所要解决的技术问题是提供一种解决数据不均衡问题的群体感知扩散数据增强方法,其解决的技术问题是,在内部威胁行为数据增强场景下,现有生成方法难以同步保持群体级统计分布一致性、解耦建模复杂时空依赖,且批次内随机噪声采样导致训练目标对比失稳
[0062]本发明的有益效果是:由于本发明在步骤三中引入基于MMD的群体感知训练目标,显式约束了合成数据在各特征维度的边缘分布与特征间的条件依赖分布,从物理分布层面消除了传统生成模型常见的分布偏移现象;同时,步骤2采用双通道架构将时序演变与跨维关联解耦建模,确保了生成序列在时间逻辑与特征协同上的语义连贯性;步骤3的同一步扩散采样策略统一了批次内样本的噪声基准,保障了群体分布对比的公平性与梯度更新的稳定性。
Smart Images

Figure CN122818346A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the interdisciplinary field of cybersecurity and artificial intelligence, specifically involving technologies for enhancing internal threat behavior data, sample equalization, and training downstream detection models based on multivariate time series data. Background Technology
[0002] In internal threat behavior detection scenarios, the scarcity of anomalous samples and the extreme imbalance in class distribution are the core bottlenecks restricting the generalization ability of detection models. Current data augmentation methods for time series data are mainly divided into two categories: traditional perturbation resampling and deep generative models (such as GANs and standard diffusion models). However, existing technologies have the following functional limitations:
[0003] (1) Lack of population statistical attributes leads to distribution shift: Existing diffusion models focus on point-by-point reconstruction error during training, neglecting the marginal value distribution and functional dependencies between features at the overall dataset level. This neglect results in synthetic data being realistic on individual samples, but shifting from the original data in population statistical distribution, thus weakening the decision reliability of downstream detection models.
[0004] (2) The coupling between temporal dynamics and cross-dimensional correlation modeling is chaotic: internal threat behavior includes both the evolutionary pattern of the time dimension and the linkage relationship between the feature dimensions. Existing methods often use a single network to process temporal and feature information, or focus only on a single dimension, which makes it impossible to effectively decouple and capture complex spatiotemporal dependency structures, resulting in poor semantic coherence of the generated data.
[0005] (3) Inconsistent diffusion sampling noise leads to training instability: Traditional diffusion models randomly assign independent diffusion time steps to each sample in a batch during training, resulting in samples in the same batch being at different noise levels. Under this condition, when calculating the population distribution constraint loss, the noise benchmark is not uniform, which makes the distribution comparison ineffective. It is difficult to balance diversity and distribution alignment during the training process, affecting the generation quality. Summary of the Invention
[0006] The technical problem this invention aims to solve is to provide a group-aware diffusion data augmentation method that addresses the issue of data imbalance. Specifically, in scenarios involving augmentation of internal threat behavior data, existing generation methods struggle to maintain consistency in group-level statistical distribution, decouple complex spatiotemporal dependencies in modeling, and suffer from instability in training target comparisons due to random noise sampling within batches. This invention provides a group-aware diffusion generation method that explicitly constrains group distribution, independently models temporal and feature dependencies, and ensures training stability through co-frequency sampling, thereby overcoming the aforementioned shortcomings.
[0007] The technical solution of this invention to solve the above-mentioned technical problems is as follows: A method for group perception diffusion data enhancement to solve the problem of data imbalance, comprising the following steps:
[0008] Step 1: Obtain the time-series priority view through multi-source log preprocessing and dual-view reconstruction. With feature-priority view ;
[0009] Step 2, based on the results obtained in Step 1 and A preliminary synthesized sequence is obtained through dual-channel independent encoding and feature fusion. ;
[0010] Step 3, based on the original data from Step 1 and Step 2... The updated model parameters are obtained through simultaneous diffusion sampling and group-aware constraint training. ;
[0011] Step 4, based on the model parameters obtained in Step 3 By reverse-generating high-fidelity threat sequences and constructing datasets, a class-balanced augmented training set is obtained.
[0012] In some possible implementations, step 1 specifically involves:
[0013] Step 1.1 involves performing field mapping, timestamp unification, linear interpolation to impute missing values, and quantile outlier pruning on the original multivariate time series of internal threat behaviors to eliminate format ambiguity and acquisition noise, resulting in a regularized multivariate time series tensor. ,in The total number of samples, The time step length, The number of dimensions for behavioral features;
[0014] Step 1.2: Use a sliding window to truncate long sequences and zero-padding / repetitive sampling to complete short sequences, ensuring a fixed sequence length. Perform Z-score normalization on each feature dimension:
[0015] ;
[0016] In the formula, and The first The sample mean and standard deviation of the dimensional feature;
[0017] Step 1.3, standardize the tensor Perform dimensional permutation: Preserve the original dimensional order to obtain a time-series priority view. A feature-first view is obtained by replacing the feature dimension with the time dimension. .
[0018] In step 1, cleaning and standardization eliminate dimensional differences and noise to avoid interfering with subsequent gradient optimization; the dual-view physical separation of temporal evolution and feature linkage provides structured input for subsequent decoupled encoding, solving the problem of incomplete dependency capture caused by a single network mixing temporal and feature information.
[0019] In some possible implementations, step 2 specifically involves:
[0020] Step 2.1: The dual-view input is mapped to a high-dimensional feature space through independent fully connected linear layers:
[0021] ;
[0022] ;
[0023] In the formula, , For learnable weight matrix, , It is the bias vector;
[0024] Step 2.2, introduce sinusoidal position coding injection sequence prior:
[0025] ;
[0026] ;
[0027] In the formula, For periodic position encoding functions;
[0028] Step 2.3: Input the encoded features into the base Transformer encoder for global dependency capture:
[0029] ;
[0030] ;
[0031] Step 2.4: The feature stream enters the multi-level diffusion transform block for iterative denoising and reconstruction. For the first... layer:
[0032] ;
[0033] ;
[0034] In the formula, ;
[0035] Step 2.5: The dual-channel outputs are element-wise weighted summed and fused, then mapped back to the original dimension via the output projection layer.
[0036] ;
[0037] In the formula, , These are the output layer parameters.
[0038] In step 2, dual-channel independent encoding decouples temporal dynamics and cross-feature associations, accurately capturing complex spatiotemporal dependency structures; positional encoding is combined with Transformer and multi-level diffusion transform block (DiT Block) to improve long-distance dependency modeling capabilities and denoising reconstruction accuracy; feature fusion aggregates temporal and feature information to ensure the temporal logic rationality and feature association consistency of the generated sequence.
[0039] In some possible implementations, step 3 specifically involves:
[0040] Step 3.1, Simultaneous Diffusion Sampling: Uniform sampling diffusion time step within the batch. Generate standard normal noise Perform forward noise addition:
[0041] ;
[0042] In the formula, This is the cumulative signal-to-noise ratio coefficient. Noise injection coefficient;
[0043] Step 3.2, Construct the group perception training target loss:
[0044] One-dimensional value distribution MMD: for the th Each feature dimension is used to calculate the original data. With synthetic data Distance from the mean in RKHS:
[0045] ;
[0046] In the formula, For kernel mapping functions, The square of the norm;
[0047] Feature-dependent distribution (MMD): For any pair of features Introduce dependency mapping functions Capture linkage relationships:
[0048] ;
[0049] Total PAT loss:
[0050] ;
[0051] In the formula, , Hyperparameters that are greater than zero;
[0052] Step 3.3, Joint Optimization: The total loss combines the standard diffusion loss and the PAT loss.
[0053] ;
[0054] In the formula, Control the strength of the group constraints; use the AdamW optimizer to update the model parameters and obtain the convergent parameters. .
[0055] In step 3, a unified noise benchmark is sampled in the same step to solve the problems of distribution comparison failure and training instability caused by inconsistent noise within batches; the loss explicitly constrains the marginal distribution and the conditional dependency distribution between features to completely eliminate the group-level distribution offset of the generated data; joint optimization balances the generation diversity and distribution alignment to improve the stability of gradient update and generation quality.
[0056] In some possible implementations, step 4 specifically involves:
[0057] Step 4.1, input pure Gaussian noise ,from Iterate to Perform reverse diffusion denoising:
[0058] ;
[0059] In the formula, For the inverse process variance, To assist with Gaussian noise, noise is gradually removed to generate a synthetic time-series dataset. ;
[0060] Step 4.2, will The dataset is mixed with the original imbalanced dataset at a preset ratio to construct a class-balanced augmented training set, which can be directly input into the downstream detection model for training.
[0061] Step 4 uses backdiffusion to generate high-fidelity, statistically aligned threat behavior sequences to supplement scarce anomalous samples. The balanced training set effectively alleviates the data imbalance problem in internal threat detection. While achieving a suboptimal recall rate (only 0.002 away from the best), the downstream detection model outperforms comparison methods such as Diffusion-TS and TimeGAN in terms of precision and F1 score, achieving dual suppression of false negatives and false positives.
[0062] The beneficial effects of this invention are as follows: In step three, the invention introduces a group perception training objective based on MMD, which explicitly constrains the marginal distribution of synthetic data in each feature dimension and the conditional dependency distribution between features, thus eliminating the distribution shift phenomenon commonly seen in traditional generative models from the perspective of physical distribution. At the same time, step two adopts a dual-channel architecture to decouple the temporal evolution and cross-dimensional correlation modeling, ensuring the semantic coherence of the generated sequence in terms of temporal logic and feature collaboration. The same-step diffusion sampling strategy in step three unifies the noise benchmark of samples within a batch, ensuring the fairness of group distribution comparison and the stability of gradient update.
[0063] In comparative experiments (see the experimental section of the detailed implementation), the synthetic data generated by this invention highly overlaps with the original real data in the PCA and t-SNE dimensionality reduction spaces. Figure 3 , Figure 4 This objectively demonstrates the solution to the distribution imbalance problem. In downstream GRU classifier testing, our method outperforms Diffusion-TS in precision, recall, and F1 score. In recall, our method achieves the second-best result with a score 0.002 lower than Diffusion-TS. In the other two categories, it outperforms Diffusion-TS, TimeGAN, and other comparable methods, with precision reaching 0.971, recall 0.963, and F1 score 0.967. These results directly verify the effectiveness of our invention in mitigating class imbalance, improving minority class coverage, and controlling false alarm rates, achieving a closed-loop technical effect of "high-fidelity distribution alignment + strong spatiotemporal dependency capture." Attached Figure Description
[0064] Figure 1 This is the overall flowchart of the present invention;
[0065] Figure 2 Overall architecture diagram for enhancing insider threat behavior data;
[0066] Figure 3 Visualize the PCA results;
[0067] Figure 4 Visualize the t-SNE results;
[0068] Figure 5 Comparison results of different data augmentation methods. Detailed Implementation
[0069] The principles and features of the present invention are described below with reference to the accompanying drawings. The examples given are only for explaining the present invention and are not intended to limit the scope of the present invention.
[0070] This invention provides a group perception diffusion data augmentation method to solve the data imbalance problem. The overall process is as follows: Figure 1 As shown. The macroscopic steps are as follows:
[0071] Step 1: Multi-source log preprocessing and dual-view reconstruction. The original multivariate time series of insider threat behaviors are cleaned, missing values imputed, aligned with a sliding window, and standardized to zero mean and unit variance. Then, through dimensionality replacement, a time-series priority view focusing on temporal evolution is physically separated. Feature Priority View Linked to Focused Features The purpose of this step is to provide structured input for subsequent decoupling coding.
[0072] Step 2: Dual-channel independent encoding and feature fusion. The dual-viewpoint views are input into the temporal and feature channels respectively. Initial dependency patterns are extracted through linear projection, positional encoding, and a standard Transformer. These patterns are then input into a multi-level diffusion transform block (DiT Block) for deep denoising feature learning. Finally, the dual-channel outputs are fused by weighted summation and mapped back to the original dimensions to obtain the preliminary synthesized sequence. For its overall architecture, please refer to [link / reference]. Figure 2 .
[0073] Step 3: Simultaneous Diffusion Sampling and Group Perception Constraint Training. In each training batch, a fixed diffusion time step is used for uniform sampling. Perform forward noise addition to obtain noisy samples. The model is based on Predict noise and estimate denoised data Calculate the maximum mean difference (MMD) to construct the group-aware training objective (PAT) loss. The distribution of the constrained one-dimensional values and the distribution of the functional dependencies between features are consistent with the original data. Total loss Combined with standard diffusion loss Optimize model parameters.
[0074] Step 4: High-fidelity threat sequence reverse generation and dataset construction. After the model training converges, a complete reverse diffusion denoising process is performed starting from pure Gaussian noise. The generated high-quality threat behavior time series is iteratively output and mixed with the original positive and negative samples to construct a balanced training set for use by downstream detection models.
[0075] The implementation steps of the present invention will be described in detail below with reference to the accompanying drawings and specific operational logic.
[0076] (1) Step 1: Multi-source behavior log preprocessing and dual-view construction
[0077] Input: A sequence of internal threat behavior logs collected from the original heterogeneous system.
[0078] Processing and formulas:
[0079] First, field mapping, timestamp unification, and linear interpolation / quantile outlier pruning are performed on the original logs to eliminate format ambiguity and collection noise, resulting in a well-ordered multivariate time series tensor. ,in The total number of samples, The time step length, The number of dimensions for behavioral features.
[0080] Secondly, a sliding window technique is used to truncate long sequences and to perform zero-padding or repeated sampling on short sequences, thus unifying the sequence length to a fixed value. Then, Z-score standardization was performed on each feature dimension to eliminate dimensional differences:
[0081] (1)
[0082] In the formula, and The first The physical function of the sample mean and standard deviation of the dimensional features is to make the data conform to a zero-mean, unit-variance distribution, and to prevent numerical scale differences from interfering with subsequent gradient optimization.
[0083] Finally, to adapt to the dual-channel architecture, the standardized tensor was modified. Perform a dimension replacement operation: Preserve the original dimension order to obtain a time-series priority view. The physical meaning is to focus on the continuous evolution trajectory of the same feature at different time steps. By replacing the feature dimension with the time dimension, we obtain a feature-priority view. The physical meaning is to focus on the instantaneous linkage relationship between different feature dimensions at the same time step.
[0084] Output: Dual-view input tensor and The result is passed to step 2 for decoupling encoding.
[0085] Attached image reference: See Figure 1 The data preprocessing module and Figure 2 The input view section.
[0086] (2) Step 2: Dual-channel independent coding and deep denoising feature fusion
[0087] Input: Timing Priority View With feature-priority view .
[0088] Processing and formulas:
[0089] First, the dual-view input is mapped to a high-dimensional feature space through independent fully connected linear layers:
[0090] (2)
[0091] (3)
[0092] In the formula, , For learnable weight matrix, , This is the bias vector. The physical function of this operation is to upscale the low-dimensional original behavioral features to a high-dimensional manifold space where the model can easily capture nonlinear relationships.
[0093] Subsequently, a sinusoidal positional coding injection sequence prior is introduced:
[0094] (4)
[0095] (5)
[0096] In the formula, For periodic positional encoding functions, the physical meaning is to provide absolute positional identifiers of time steps or feature indices for self-attention mechanisms that lack sequence perception capabilities, ensuring that the model can distinguish the order of occurrence of behaviors or the arrangement logic of feature dimensions.
[0097] Next, the encoded features are input into the base Transformer encoder for global dependency capture:
[0098] (6)
[0099] (7)
[0100] The basic encoder is computed alternately through multi-layer self-attention and feedforward networks. Its physical function is to establish long-distance temporal correlations and cross-feature collaborative patterns.
[0101] Then, the feature stream enters the multi-level diffusion transform block (DiT Block) for iterative denoising and reconstruction. For the ... layer( ):
[0102] (8)
[0103] (9)
[0104] Each DiT Block contains layer normalization, conditional multi-head self-attention, feedforward neural network and residual connection. Its physical function is to gradually strip away superimposed Gaussian noise and reconstruct the multimodal distribution structure of the underlying data under the guidance of a given diffusion time step condition.
[0105] Finally, the dual-channel outputs are summed and fused element-wise, and then mapped back to the original feature space through the output projection layer:
[0106] (10)
[0107] In the formula, , These are the output layer parameters. The physical function of this step is to re-aggregate the temporal dynamics and feature dependencies of the decoupled learning, generating a synthetic sequence that simultaneously satisfies temporal logical rationality and feature association consistency.
[0108] Output: Predicted synthetic sequence value at the current diffusion step That is, the estimated .
[0109] Attached image reference: See Figure 2 The dual-channel encoder architecture section.
[0110] (3) Step 3: Simultaneous diffusion sampling and group perception constraint training
[0111] Input: Original clean sample within the batch Current parameters of the model .
[0112] Processing and formulas:
[0113] 1) Simultaneous Diffusion Sampling: To avoid distribution distortion caused by inconsistent noise levels within a batch, this invention enforces a unified batch time step. Random Sampling And generate standard normal noise. The forward noise addition formula is:
[0114] (11)
[0115] In the formula, The cumulative signal-to-noise ratio coefficient, in physical terms, controls the proportion of the original signal retained in each step; This is the noise injection coefficient, controlling the intensity of random perturbations. Simultaneous sampling ensures all samples within a batch are under the same signal-to-noise ratio baseline, providing a fair physical comparison environment for subsequent population distribution calculations.
[0116] 2) Group Awareness Training Objective (PAT) Construction: Model Reception Predicted noise And reverse-engineer the cleaning data To constrain population distribution consistency, a maximum mean difference (MMD) loss based on the regenerating kernel Hilbert space (RKHS) is introduced.
[0117] • One-dimensional value distribution MMD: For the first Each feature dimension is used to calculate the original data. With synthetic data Distance from the mean in RKHS:
[0118] (12)
[0119] In the formula, For kernel mapping function, the physical function is to project the original eigenvalues to a high-dimensional Hilbert space and implicitly calculate the differences of complex nonlinear distributions through inner product operation; The norm square quantifies the degree of statistical deviation between two distributions along this dimension.
[0120] • Feature-specific functional dependency distribution (MMD): For any pair of features Introduce dependency mapping functions Capture linkage relationships (such as cross-correlation, conditional probability):
[0121] (13)
[0122] The physical function of this step is to ensure that the synthetic data not only has a reasonable distribution of individual values, but also that the collaborative change patterns between features (such as the strong correlation between login on a specific device and access to a specific file) are consistent with real threat behavior.
[0123] • PAT total loss synthesis: weighted fusion of two types of distribution constraints:
[0124] (14)
[0125] In the formula, , The hyperparameter is greater than zero and its function is to dynamically balance the optimization weights of one-dimensional fidelity and multi-dimensional dependency consistency.
[0126] 3) Joint optimization: Combining PAT loss with standard denoising mean square error loss:
[0127] (15)
[0128] In the formula, Control the strength of the group constraints. Use the AdamW optimizer to differentiate and update the parameters. The physical function of this step is to explicitly pull the model toward the population statistical manifold of the real data while gradually denoising it, preventing the generated trajectory from deviating from the true distribution.
[0129] Output: Updated model parameters This is used for the next iteration or the final generation.
[0130] (4) Step 4: High-fidelity threat sequence reverse generation and dataset construction
[0131] Input: Model parameters after training convergence Pure Gaussian noise .
[0132] Processing: Perform the complete reverse diffusion process. From Iterate to Each step uses the trained model to predict the noise residual. And update the sample state according to the formula for the inverse diffusion process:
[0133] (16)
[0134] In the formula, For the inverse process variance, To assist Gaussian noise ( (Time). The physical function of this iterative process is to gradually remove random perturbations and, based on the temporal and feature joint distribution learned by the model, "carve" out a structurally complete and statistically aligned sequence of threat behaviors from disordered noise.
[0135] Output: Synthetic time series dataset The dataset is then mixed with the original imbalanced dataset in a predetermined ratio to form a class-balanced augmented training set, which is then directly input into downstream detection models (such as GRU, LSTM, etc.) for training.
[0136] Experimental verification section
[0137] To verify the effectiveness of this invention in solving practical technical problems, experiments were conducted based on the CMU-CERT v4.2 insider threat dataset. This dataset contains 545 days of behavioral logs from 1000 employees, extracted as (1000, 545, 48) multivariate time series, and sliced into 28-day sliding window segments as input. The downstream detector uses a benchmark GRU network, and the evaluation metrics are precision, recall, and F1 score.
[0138] (1) Distribution alignment verification (corresponding to solving the "distribution offset" problem): such as Figure 3 (PCA dimensionality reduction) and Figure 4 As shown in the t-SNE dimensionality reduction diagram, the synthetic data (blue dots) generated by this invention exhibits a high degree of overlap and interleaving with the original real data (red dots) in the two-dimensional projection space. The consistent global profile of PCA demonstrates that this invention successfully preserves the core variance structure of the original data; the local clustering of t-SNE proves that this invention restores the multimodal distribution of the real data in a fine-grained neighborhood. This visualization directly verifies the physical effectiveness of PAT loss in eliminating distribution shifts.
[0139] (2) Downstream detection performance verification (addressing the issues of "dependency modeling and class imbalance"): The enhanced dataset of this invention, along with data generated by Diffusion-TS, TimeGAN, DoppelGANger, and SigDiffusions, were used to train the GRU detector. The results are as follows: Figure 5 As shown, this invention outperforms all other core metrics except Recall, which is only slightly better: Precision reaches 0.971, Recall reaches 0.963, and F1 reaches 0.967. Compared to TimeGAN (lowest in all three metrics) and SigDiffusions (high precision but low recall), this invention achieves dual suppression of false positives and false negatives. In internal threat behavior detection, low recall means missed detection of critical threats, posing a significant risk; this invention achieves a recall of 96.3%, combined with a precision of 0.971, demonstrating that the dual-channel architecture effectively captures the spatiotemporal dependencies of complex threats. Furthermore, simultaneous sampling and PAT constraints ensure the diversity and authenticity of generated samples, completely resolving the technical contradictions of traditional methods where "high coverage inevitably leads to high false positives" or "high precision inevitably leads to low coverage."
[0140] (3) Experimental closed-loop conclusion: The above experimental data and visualization results form a complete logical chain: PAT distribution constraint → elimination of distribution offset ( Figure 3 , 4 Validation) → Dual-channel decoupling + stable training at the same frequency → Generation of high-fidelity diverse samples → Downstream detector Precision / Recall dual boost ( Figure 5 (Verification) → Completely overcome the three major defects in the background technology.
[0141] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A group perception diffusion data augmentation method for solving the problem of data imbalance, characterized in that, The steps include the following: Step 1: Obtain the time-series priority view through multi-source log preprocessing and dual-view reconstruction. With feature-priority view ; Step 2, based on the results obtained in Step 1 and A preliminary synthesized sequence is obtained through dual-channel independent encoding and feature fusion. ; Step 3, based on the original data from Step 1 and Step 2... The updated model parameters are obtained through simultaneous diffusion sampling and group-aware constraint training. ; Step 4, based on the model parameters obtained in Step 3 By reverse-generating high-fidelity threat sequences and constructing datasets, a class-balanced augmented training set is obtained.
2. The group perception diffusion data augmentation method for solving the data imbalance problem according to claim 1, characterized in that, Step 1 is as follows: Step 1.1 involves performing field mapping, timestamp unification, linear interpolation to impute missing values, and quantile outlier pruning on the original multivariate time series of internal threat behaviors to eliminate format ambiguity and acquisition noise, resulting in a regularized multivariate time series tensor. ,in The total number of samples, The time step length, The number of dimensions for behavioral features; Step 1.2: Use a sliding window to truncate long sequences and zero-padding / repetitive sampling to complete short sequences, ensuring a fixed sequence length. Perform Z-score normalization on each feature dimension: ; In the formula, and The first The sample mean and standard deviation of the dimensional feature; Step 1.3, standardize the tensor Perform dimensional permutation: Preserve the original dimensional order to obtain a time-series priority view. ; The feature-first view is obtained by replacing the feature dimension with the time dimension. .
3. The group perception diffusion data augmentation method for solving the data imbalance problem according to claim 1, characterized in that, Step 2 is as follows: Step 2.1: The dual-view input is mapped to a high-dimensional feature space through independent fully connected linear layers: ; ; In the formula, , For learnable weight matrix, , It is the bias vector; Step 2.2, introduce sinusoidal position coding injection sequence prior: ; ; In the formula, For periodic position encoding functions; Step 2.3: Input the encoded features into the base Transformer encoder for global dependency capture: ; ; Step 2.4: The feature stream enters the multi-level diffusion transform block for iterative denoising and reconstruction. For the first... layer: ; ; In the formula, ; Step 2.5: The dual-channel outputs are element-wise weighted summed and fused, then mapped back to the original dimension via the output projection layer. ; In the formula, , These are the output layer parameters.
4. The group perception diffusion data augmentation method for solving the data imbalance problem according to claim 1, characterized in that, Step 3 specifically involves: Step 3.1, Simultaneous Diffusion Sampling: Uniform sampling diffusion time step within the batch. Generate standard normal noise Perform forward noise addition: ; In the formula, This is the cumulative signal-to-noise ratio coefficient. Noise injection coefficient; Step 3.2, Construct the group perception training target loss: One-dimensional value distribution MMD: for the th Each feature dimension is used to calculate the original data. With synthetic data Distance from the mean in RKHS: ; In the formula, For kernel mapping functions, The square of the norm; Feature-dependent distribution (MMD): For any pair of features Introduce dependency mapping functions Capture linkage relationships: ; Total PAT loss: ; In the formula, , Hyperparameters that are greater than zero; Step 3.3, Joint Optimization: The total loss combines the standard diffusion loss and the PAT loss. ; In the formula, Control the strength of group constraints; The AdamW optimizer is used to update the model parameters, resulting in converged parameters. .
5. The group perception diffusion data augmentation method for solving the data imbalance problem according to claim 1, characterized in that, Step 4 specifically involves: Step 4.1, input pure Gaussian noise ,from Iterate to Perform reverse diffusion denoising: ; In the formula, For the inverse process variance, To supplement Gaussian noise; Gradually remove noise to generate synthetic time series datasets ; Step 4.2, will The dataset is mixed with the original imbalanced dataset at a preset ratio to construct a class-balanced augmented training set, which can be directly input into the downstream detection model for training.