Multi-branch depth anomaly traffic discrimination method based on disturbance response regulation and control mechanism

By constructing a multi-branch deep anomaly traffic discrimination method based on a disturbance response control mechanism, the problems of insufficient robustness of feature disturbances and insufficient guidance of multi-branch structures in existing technologies are solved, and high accuracy and stability of anomaly traffic identification are achieved in complex network environments.

CN121145084APending Publication Date: 2025-12-16JIANGSU COLLEGE OF INFORMATION TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511244494.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-02
Publication Date
2025-12-16

AI Technical Summary

Technical Problem

Existing abnormal traffic detection methods struggle to simultaneously achieve robustness against feature perturbations, adaptive channel response, and guidance of multi-branch structures in complex network environments, resulting in insufficient accuracy and stability in identifying various types of abnormal traffic.

Method used

A multi-branch deep anomaly traffic discrimination method based on disturbance response control mechanism is adopted. Through data preprocessing and disturbance-compatible coding, multi-path structure feature modeling and dynamic response-guided fusion strategy, a multi-path response space is constructed, and feature fusion and discrimination are performed using three types of functional sub-network structures.

Benefits of technology

It significantly improves the accuracy and robustness of anomaly identification in complex network traffic scenarios, enhances the adaptability to unknown disturbances and anomaly patterns, and improves the consistent expressive ability across feature channels and the generalization performance of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121145084A_ABST
    Figure CN121145084A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-branch depth anomaly traffic discrimination method based on a disturbance response regulation and control mechanism. The method comprises the following steps: A, data preprocessing and disturbance compatible coding modeling; b, heterogeneous representation sub-network initialization and multi-path parallel construction are carried out; c, a phi MLSE branch multi-layer nonlinear expansion and explicit response channel regulation and control mechanism; d, a phi TEC branch local sensing and global structure coupling mapping mechanism; e. a phi BFC branch dual-channel disturbance compression mapping network; f. a psi AGIF multi-branch guide type response weighted fusion mechanism; g.psi DEC nonlinear discriminant mapping and normalization response output mechanism; h, carrying out unbalanced multi-class distribution perception loss modeling; i, a disturbance perception training temperature control scheduling and multi-mode self-stabilization updating mechanism; and J, model parameter freezing and deployment export construction. The method has the advantages that the problem of feature inconsistency between heterogeneous fields is effectively relieved, the adaptability of the model to unknown disturbance and abnormal modes is enhanced, the consistency expression ability of cross-feature channels is improved, and each branch structure of a subsequent network can be directly connected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to an abnormal flow discrimination method, and more particularly to a multi-branch deep abnormal flow discrimination method based on a disturbance response control mechanism. Background Technology

[0002] This paper addresses the technical bottlenecks in current abnormal traffic detection tasks, including insufficient representation of multi-source heterogeneous features, structural inconsistencies, sensitivity to perturbations and noise, and weak ability to identify long-tail categories. Existing methods, such as shallow models and traditional attention mechanisms, often fail to simultaneously achieve robustness to feature perturbations, adaptive channel response, and guidance through multi-branch structures, making it difficult to achieve high accuracy and stability in identifying multiple types of abnormal traffic in complex network environments. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to provide a multi-branch deep anomaly traffic discrimination method based on a disturbance response control mechanism, which significantly improves the accuracy and robustness of anomaly identification in complex network traffic scenarios by integrating a disturbance perception preprocessing mechanism, multi-path structural feature modeling and dynamic response guidance fusion strategy.

[0004] To address the aforementioned technical problems, this invention provides a multi-branch deep anomaly flow discrimination method based on a disturbance response control mechanism, comprising the following steps:

[0005] A. Data preprocessing and perturbation-compatible coding modeling

[0006] a. Raw traffic collection and feature initialization: Collect multiple types of traffic samples. Each sample contains multiple category features and numerical features. All samples are used to form an initial raw feature set X0, where each sample contains multi-dimensional information with fixed dimensions.

[0007] b. Perturbation-compatible one-hot encoding for categorical fields: Perturbation-compatible one-hot encoding is performed on all categorical fields. The specific encoding process includes constructing a standard category set C and introducing an additional category c. unk Representing unknown or noisy inputs, each category field is encoded as a sparse vector of dimension |C|+1. After multi-field encoding, a tensor X is generated. cat ;

[0008] c. Bit-level expansion mapping for highly dynamic numerical fields: For numerical fields with drastically fluctuating value ranges, a perturbation decomposition-based bit expansion strategy is introduced. The maximum number of digits *d* for each field is calculated, and each numerical value is mapped to a *d*-bit integer array, forming a high-dimensional sparse tensor. Multiple fields are expanded and then merged to form tensor X. digit ;

[0009] d. Multi-source feature fusion and structural splicing: The original numerical tensor X is fused and structured. numEncoding tensor X cat With position expansion tensor X digit Perform a concatenation operation to generate an input tensor with uniform structural features:

[0010] X = Concat(X) num ,X cat ,X digit ),X∈R N×D

[0011] Where N represents the number of samples, and D is the dimension after fusion;

[0012] e. Input tensor normalization and tuning: Constructing the normalization control operator T norm This is used to perform standard normalization on the input tensor, ultimately yielding a normalized tensor. The strategy for normalization is defined as follows:

[0013]

[0014] Where, μ∈R 1×D Let σ represent the feature mean calculated along the sample dimension, where σ∈R 1×D Let X be the standard deviation, ∈ be a numerically stable constant, and the output tensor X retains dimensional information, resulting in a normalized tensor.

[0015] B. Initialization and multi-path parallel construction of heterogeneous representation subnetworks: Normalized feature tensors... As a unified input interface, three types of functional sub-network structures are simultaneously injected to construct a multi-path response space. The sub-network structures include Φ MLSE Branch, Φ TEC Branches and Φ BFC Branching, output feature Z k Output feature Z k The dimensions are unified to 128, which facilitates structurally consistent fusion and subsequent joint training.

[0016] Φ MLSE The branch is constructed using a multi-layer nonlinear activation mapping, a channel saliency enhancement mechanism, and a local context interaction module structure, with an input dimension of R. N×D The final output is mapped to R. N×128 , denoted as:

[0017]

[0018] Φ TEC The branch is constructed based on pseudo-structure modeling and cross-embedding perception mechanism, firstly by... Projected to the L·E dimension, then rearranged into the structure tensor R N×L×E After being nested with attention structures, it is compressed to R N×128 , denoted as:

[0019]

[0020] Φ BFC The branch adopts a two-layer compression + sparse modulation mechanism, Φ BFC Branches are based on the architecture-aware path, directly... Process into low-dimensional discriminant vectors:

[0021]

[0022] C.Φ MLSE Branched multilayer nonlinear expansion and explicit response channel modulation mechanism:

[0023] a. First-layer structure expansion and response control: First, the normalized input tensor... Mapping to an intermediate higher-order space via linear transformation Let D1 = 512:

[0024]

[0025] Subsequently, a self-stabilizing nonlinear activation ψ is introduced. swish With the perturbation-weighted mechanism, the response regulation tensor H1 is constructed as follows:

[0026]

[0027] in For Dropout operations, ⊙ indicates element-wise weighting by channel dimension;

[0028] b. Explicit Disturbance Response Weighting Module definition The module calculates the channel disturbance response value per channel, assuming the input... Let represent the intermediate feature vector of the i-th sample, and its channel control vector is defined as follows:

[0029]

[0030] For each channel c:

[0031]

[0032] This represents a channel-specific perturbation mapping network; η is the perturbation amplitude hyperparameter; α controls the compression ratio; φ c (·) is a nonlinear transderivative function based on variance adjustment;

[0033] c. Second-layer feature expansion and re-modulation: H1 is further processed and projected onto... Let D2 = 256, and implement perturbation response control symmetrical to the first layer:

[0034]

[0035] d. Enhanced contextual interaction and structural convergence further reduce the dimensionality of the feature map H2 to Set D3 = 128 and send it to the local context interaction module:

[0036]

[0037] A lightweight local attention mechanism is employed to uncover local structural consistency and improve discriminative power;

[0038] D.Φ TEC Branch-local perception and global structure coupling mapping mechanism: To model the inter-layer dependencies of input features in a pseudo-structured space, a table structure enhancement coupling module is proposed. Through structure-aware embedding rearrangement, multi-head interactive aggregation, perturbation compression, and nested feedback mechanisms, the original tensor is processed... Perform multi-granularity structure modeling, and the entire process output is as follows:

[0039] a. Subspace structure activation embedding Through linear embedding Map the input to an L·E-dimensional pseudo-structure space, where L=6 is the structure step size and E=6 is the embedding dimension per step:

[0040]

[0041] Where ψ ReLU (·) represents the ReLU activation function. This means reorganizing the flattened vector into L structural sub-units, each sub-unit having dimension E;

[0042] b. Structurally Coupled Interaction Module Employing a multi-head coupled attention mechanism With Feedforward Network

[0043]

[0044] Where H=4 represents the number of structure-sensing heads, each... All are defined as:

[0045]

[0046] Note the weight matrix. It can be trained independently, enhancing the ability to model inter-channel responses;

[0047] Then through a structural feedback feedforward network Input structure delay and hierarchical feature interaction:

[0048]

[0049] in For a linear mapping, ψ GELU It is a higher-order activation function;

[0050] c. Disturbance response compression and regularized suppression Introducing the disturbance control compression mechanism DERI to H ff Random gating and structural response modulation are applied to each structural unit:

[0051]

[0052] ∈ i ~Bernoulli(1-α) i ),

[0053]

[0054] Where: α i The dynamic sparse control law can be adjusted through training or statistical structural complexity. For shared perturbation projection networks; output maintains tensor shape invariance, H deri ∈R N×L×E ;

[0055] d. Unify the dimensional projection output, flattening the structure tensor to R. N×(L·E) And compressed into a unified fusion dimension:

[0056]

[0057] Compression operation ensures compatibility with Φ MLSE ,Φ BFC The consistent output dimensions of branches facilitate subsequent branch merging.

[0058] E.Φ BFC Branched dual-channel perturbation compression mapping network:

[0059] a. Basic structure projection and disturbance gating on normalized feature input Apply initial structure mapping And via the bootstrap activation function Adjust response range:

[0060]

[0061] in: For trainable fully connected linear mappings It is a sigmoid-like monotonic response function with adjustable curvature, γ1>0; For perturbation compression unit;

[0062] b. Compressed mapping and sparsity control further map H1 to a low-dimensional space and introduce structural-level L1 sparsity constraints to control invalid responses:

[0063]

[0064] in: Activation curve The curvature is slightly less than γ1 for smoother post-response control; a term λ is added to H2 during training. sp ‖H2‖1, where λ sp For sparsity regularization weights;

[0065] c. Unified Projection and Output Representation: The final output vector will be compressed to a dimension R consistent with other branches. 128 As a basic structural guidance pathway:

[0066]

[0067] This feature will participate in multi-branch weighted fusion, providing a benchmark for higher-order representations;

[0068] F.ΨAGIF multi-branch guided response weighted fusion mechanism:

[0069] Design an adaptive guided weight fusion mechanism to integrate the three heterogeneous substructure paths. The system dynamically learns the importance weights of each element to achieve structure-aware feature fusion.

[0070] Set the output of each sub-path For each branch k, the guiding weight ω k ∈R N×128 The generation method is as follows:

[0071]

[0072] in: ρ(·) represents the local nested structure compression mapping of each branch; ρ(·) represents the normalized perceptron function, such as LayerNorm or a custom statistical normalization function; σ(·) is the Sigmoid activation function, which implements element-level gating. The trainable guiding matrix is ​​mapped back to the weight dimension;

[0073] The final fused feature output is:

[0074]

[0075] G.Ψ DEC Nonlinear discriminant mapping and normalized response output mechanism:

[0076] a. Discriminative feature mapping path construction: First, nonlinear response and perturbation suppression processing is applied to the fused features to enhance robustness to edge samples and low-confidence discriminative regions.

[0077]

[0078] Where: ψ γ To guide the nonlinear response function, such as To map a linear projection to 128 dimensions; This is a disturbance response suppression mechanism;

[0079] b. Multi-class probability normalization output: The final output is mapped to the probability space R through the normalized softmax function NormSoftmax with a temperature parameter. N×C Where C is the number of categories:

[0080]

[0081] in: For classification projection head; τ∈(0,1] is the temperature adjustment parameter used to adjust the entropy sensitivity of the probability distribution, usually taken as τ=0.5; output This represents the prediction confidence level of each sample for each category;

[0082] H. Modeling of impaired multi-class perceptual loss:

[0083] We propose a weighted sparse distribution-aware loss function that integrates class weights, output sparsity constraints, and numerical stability adjustment, defined as follows:

[0084]

[0085] Where: N is the number of samples, and C is the number of categories; y represents the predicted probability distribution of sample i; i ∈{1,…,C} are the true labels; For category y i The weighting factor is usually defined as: f y Represents the sample frequency of category y; ∈ = 10 -7 The first term is a stability constant to avoid log(0); the second term is an l1 sparsity regularization to suppress overly averaged prediction tendencies; μ∈R + As a hyperparameter, it adjusts the sparsity constraint strength;

[0086] I. Disturbance-aware training temperature control scheduling and multimodal self-stabilizing update mechanism:

[0087] Multi-strategy controller with disturbance response, self-stabilizing adjustment and modal coordination capabilities Its core comprises the following three innovative sub-modules:

[0088] a. Disturbance temperature control controller This module dynamically monitors the error change rate during training and calculates the perturbation temperature index T accordingly. dyn (t) is used to reflect the current training stability. When the perturbation amplitude is large, the system will automatically reduce the parameter update rate γ. t ;

[0089] b. Multi-scale convergence prediction and structural backflow module This module constructs a sliding window observation sequence.

[0090] c. Construct a modal transition cooling mechanism Specifically, this includes: detecting the current submodal. The training state is temporarily frozen; parameter updates for this mode are temporarily frozen, and the training objective is switched to another sub-mode. like If the error trend recovers and improves, then the freeze is lifted and the training path is restored;

[0091] J. Model Parameter Freezing and Deployment Export Construction: After completing the full-cycle multi-round training scheduling, the system freezes the current optimal discriminant structure and its weight set to construct the final deployable discriminant model.

[0092] The multi-dimensional information includes communication direction, protocol identifier, access status, and number of bytes.

[0093] Advantages of this invention:

[0094] (1) Perturbation-compatible one-hot encoding is performed on all categorical fields to construct a normalized control operator, which effectively alleviates the problem of feature inconsistency between heterogeneous fields, enhances the model's adaptability to unknown perturbations and abnormal patterns, improves the consistent expression capability across feature channels, and can be directly connected to the subsequent network branch structures to support end-to-end gradient propagation and structural optimization.

[0095] (2) The normalized feature tensor is injected into the three types of functional sub-network structures at the same time to construct a multi-path response space. The multi-branch structure coupling strategy realizes complementary modeling of feature subspace and improves the response capability to boundary samples and weak abnormal signals.

[0096] (3) An explicit response trade-off and feature perturbation control module is proposed to replace the conventional compressed activation channel attention structure. This module has stronger channel-independent modeling and nonlinear suppression capabilities, and can adapt to the dynamic response requirements of different input samples. It does not rely on channel average compression (unlike the SE model), and uses perturbation control instead of channel compression activation to avoid information loss. The activation response is adjustable, and it has sample dynamics, avoiding static weights of a single channel. The tensor dimensions of the entire process are transparent and controllable, and all dimensions are clear and scalable, which is extremely friendly to engineering deployment.

[0097] (4) Branched dual-channel perturbation compression mapping network: This branch aims to construct a low-order auxiliary feature path to extract the structural basis response of non-high-order expressions in the source data as a discriminant compensation term in the multi-branch fusion mechanism. It has asymmetry and structural constraints, reducing redundancy in the representation space; Multi-layer perturbation compression chain: Both consecutive layers adopt Gating mechanism to enhance response suppression to low-confidence features; Guided activation tunability: using ψ γ Instead of conventional ReLU or Swish, the function has an adjustable response curvature, enhancing the network's ability to distinguish signals in different intervals; controllable sparse compression: adding an adjustable regularization constraint ‖·‖1 controls the output sparse density, which helps subsequent classifiers focus on salient substructures;

[0098] (5) Traditional methods use fixed weights or averaging, while this method introduces structural projection, gated response, and normalization adjustment; cross-branch dimensional consistency alignment: ensuring that the outputs of each sub-path are unified to R after guided modulation. 128 The guided response is trainable and interpretable: ω k It can be used for subsequent analysis of branch importance and characteristic contribution regions;

[0099] (6) To address the common problems of convergence stagnation, overfitting, and training oscillation in deep anomaly flow discrimination models during training with long-tailed categories, multimodal structures, and high complexity, a multi-strategy controller with disturbance response, self-stabilizing adjustment, and modal coordination capabilities is designed. The classifier has temperature-controlled normalization capability and adjustable sparsity suppression mechanism, which can adapt to the long-tail class discrimination requirements and improve the generalization performance and engineering deployability of the model. Attached Figure Description

[0100] Figure 1 This is a flowchart of the multi-branch deep anomaly flow discrimination method based on the disturbance response control mechanism of the present invention. Detailed Implementation

[0101] The multi-branch deep anomaly flow discrimination method based on disturbance response control mechanism of the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0102] Example:

[0103] A multi-branch deep anomaly flow discrimination method based on disturbance response control mechanism, A. Data preprocessing and disturbance-compatible coding modeling.

[0104] a. Raw traffic collection and feature initialization: Collect multiple types of traffic samples. Each sample contains multiple category features and numerical features. All samples are used to form an initial raw feature set X0. Each sample contains multi-dimensional information with fixed dimensions. The multi-dimensional information includes communication direction, protocol identifier, access status, and number of bytes.

[0105] b. Perturbation-Compatible One-Hot Encoding for Categorical Fields: To eliminate the unsuitable impact of non-numerical features on model training, perturbation-compatible one-hot encoding is performed on all categorical fields (such as communication protocols, service port types, connection flags, etc.). The specific encoding process includes constructing a standard category set C and introducing an additional category c. unk Representing unknown or noisy inputs, each category field is encoded as a sparse vector of dimension |C|+1. After multi-field encoding, a tensor X is generated. cat This method can improve the ability to express unknown categories and perturbation mutations, and enhance training stability.

[0106] c. Bit-level expansion mapping for highly dynamic numerical fields: For numerical fields with drastically fluctuating value ranges, a perturbation decomposition-based bit expansion strategy is introduced. The maximum number of digits *d* for each field is calculated, and each numerical value is mapped to a *d*-bit integer array, forming a high-dimensional sparse tensor. Multiple fields are expanded and then merged to form tensor X. digit This strategy helps to capture the sensitivity of low-level disturbances to high-level stability, and is suitable for the nonlinear table requirements of anomaly detection.

[0107] d. Multi-source feature fusion and structural splicing: The original numerical tensor X is fused and structured. num Encoding tensor X cat With position expansion tensor X digit Perform a concatenation operation to generate an input tensor with uniform structural features:

[0108] X = Concat(X) num ,X cat ,X digit ),X∈R N×D

[0109] Where N represents the number of samples, and D is the dimension after fusion;

[0110] e. Input tensor normalization and tuning: To further unify the numerical dimensions of different features and eliminate the disturbance caused by cross-dimensional shifts, a normalization tuning operator T is constructed. normThis is used to perform standard normalization on the input tensor, ultimately yielding a normalized tensor. The strategy for normalization is defined as follows:

[0111]

[0112] Where, μ∈R 1×D Let σ represent the feature mean calculated along the sample dimension, where σ∈R 1×D The standard deviation is denoted as ∈, where ∈ is a numerically stable constant (e.g., 1×10⁵). The output tensor X retains dimensional information and is fully compatible with subsequent network structures, resulting in a normalized tensor. This mechanism enhances the consistent representation capability across feature channels and can directly interface with subsequent network branches, supporting end-to-end gradient propagation and structural optimization.

[0113] Normalized feature tensor This will serve as a unified input interface, simultaneously injecting three types of functional sub-network structures to construct a multi-path response space. The processing logic, target function, and output dimension of each sub-network are shown below:

[0114] B. Initialization and multi-path parallel construction of heterogeneous representation subnetworks: Normalized feature tensors... As a unified input interface, three types of functional sub-network structures are simultaneously injected to construct a multi-path response space. The sub-network structures include Φ MLSE Branch, Φ TEC Branches and Φ BFC Branching, output feature Z k The dimensions are unified to 128, which facilitates structurally consistent fusion and subsequent joint training.

[0115] Φ MLSE The branch is constructed using a multi-layer nonlinear activation mapping, a channel saliency enhancement mechanism, and a local context interaction module structure, with an input dimension of R. N×D The final output is mapped to R. N×128 , denoted as:

[0116]

[0117] Φ TEC The branch is constructed based on pseudo-structure modeling and cross-embedding perception mechanism, firstly by... Projected to the L·E dimension, then rearranged into the structure tensor R N×L×E After being nested with attention structures, it is compressed to R N×128 , denoted as:

[0118]

[0119] Φ BFC The branch adopts a two-layer compression + sparse modulation mechanism, ΦBFC Branches are based on the architecture-aware path, directly... Process into low-dimensional discriminant vectors:

[0120]

[0121] Φ MLSE This branch proposes an explicit response trade-off and feature perturbation modulation module to replace the conventional compressed activation channel attention structure. This module has stronger channel-independent modeling and nonlinear suppression capabilities, and can adapt to the dynamic response requirements of different input samples.

[0122] C.Φ MLSE Branched multilayer nonlinear expansion and explicit response channel modulation mechanism:

[0123] a. First-layer structure expansion and response control: First, the normalized input tensor... Mapping to an intermediate higher-order space via linear transformation Let D1 = 512:

[0124]

[0125] Subsequently, a self-stabilizing nonlinear activation ψ is introduced. swish With the perturbation-weighted mechanism, the response regulation tensor H1 is constructed as follows:

[0126]

[0127] in For Dropout operations, ⊙ indicates element-wise weighting by channel dimension;

[0128] b. Explicit Disturbance Response Weighting Module definition The module calculates the channel disturbance response value per channel, assuming the input... Let represent the intermediate feature vector of the i-th sample, and its channel control vector is defined as follows:

[0129]

[0130] For each channel c:

[0131]

[0132] This represents a channel-specific perturbation mapping network; η is the perturbation amplitude hyperparameter; α controls the compression ratio; φ c (·) is a nonlinear transderivative function based on variance adjustment;

[0133] c. Second-layer feature expansion and re-modulation: H1 is further processed and projected onto... Let D2 = 256, and implement perturbation response control symmetrical to the first layer:

[0134]

[0135] d. Enhanced contextual interaction and structural convergence further reduce the dimensionality of the feature map H2 to Set D3 = 128 and send it to the local context interaction module:

[0136]

[0137] A lightweight local attention mechanism is employed to uncover local structural consistency and enhance discriminative power; this approach does not rely on channel averaging compression (unlike the SE model): perturbation modulation replaces channel compression activation, avoiding information loss; the activation response is adjustable, exhibiting sample dynamics: the perturbation factor ξ for each sample i... c Adaptive adjustment to avoid static weighting of a single channel; response variation adjustment mechanism: driven by variance φ c Nonlinear transformations suppress activation skewness and enhance robustness to small-sample features; the tensor dimensions throughout the process are transparent and controllable: all dimensions are clear and scalable, making them extremely friendly to engineering deployment;

[0138] D.Φ TEC Branch-based local perception and global structure coupling mapping mechanism: To model the inter-layer dependencies of input features in the pseudo-structured space, this branch proposes a Tabular Enhanced Coupler (TEC) module. Through structure-aware embedding rearrangement, multi-head interactive aggregation, perturbation compression, and nested feedback mechanisms, it maps the original tensor... Perform multi-granularity structure modeling, and the entire process output is as follows:

[0139] a. Subspace structure activation embedding Through linear embedding Map the input to an L·E-dimensional pseudo-structure space, where L=6 is the structure step size and E=6 is the embedding dimension per step:

[0140]

[0141] Where ψ ReLU (·) represents the ReLU activation function. This means reorganizing the flattened vector into L structural sub-units, each sub-unit having dimension E;

[0142] b. Structurally Coupled Interaction Module This module enables multi-scale interaction between local perception and global structure, employing a multi-head coupled attention mechanism. With Feedforward Network

[0143]

[0144] Where H=4 represents the number of structure-sensing heads, each... All are defined as:

[0145]

[0146] Note the weight matrix. It can be trained independently, enhancing the ability to model inter-channel responses;

[0147] Then through a structural feedback feedforward network Input structure delay and hierarchical feature interaction:

[0148]

[0149] in For a linear mapping, ψ GELU It is a higher-order activation function;

[0150] c. Disturbance response compression and regularized suppression The disturbance control compression mechanism DERI (Disturbance Enhanced Response Inhibition) is introduced to control H. ff Random gating and structural response modulation are applied to each structural unit:

[0151]

[0152] ∈ i ~Bernoulli(1-α) i ),

[0153]

[0154] Where: α i The dynamic sparse control law can be adjusted through training or statistical structural complexity. For shared perturbation projection networks (such as a single fully connected layer); the output maintains the tensor shape invariant, H deri ∈R N×L×E ;

[0155] d. Unify the dimensional projection output, flattening the structure tensor to R. N×(L·E) And compressed into a unified fusion dimension:

[0156]

[0157] Compression operation ensures compatibility with Φ MLSE ,Φ BFCThe consistent output dimensions of each branch facilitate subsequent branch fusion; this differs from the positional encoding of the traditional TabTransformer, as this module explicitly constructs a position-aware subspace based on the structural step size L and structural dimension E; and it introduces a local-global coupled attention mechanism: multi-head interactive attention. Feedforward structure feedback Constructing a fusion representation of local and global features; a perturbation gating mechanism to replace Dropout: DERI can explicitly model the reliability of structural units and improve selectivity for low-confidence feature regions.

[0158] E.Φ BFC The proposed dual-channel perturbation and compression network (DSCN) aims to construct a low-order auxiliary feature pathway to extract the structural basis responses of non-high-order expressions in the source data, serving as discriminant compensation terms in the multi-branch fusion mechanism. The proposed DSCN design emphasizes structural sparsity, nonlinear perturbation suppression, and dynamically guided activation.

[0159] a. Basic structure projection and disturbance gating on normalized feature input Apply initial structure mapping And via the bootstrap activation function Adjust response range:

[0160]

[0161] in: For trainable fully connected linear mappings It is a sigmoid-like monotonic response function with adjustable curvature, γ1>0; For perturbation compression unit;

[0162] b. Compressed mapping and sparsity control further map H1 to a low-dimensional space and introduce structural-level L1 sparsity constraints to control invalid responses:

[0163]

[0164] in: Activation curve The curvature is slightly less than γ1 for smoother post-response control; a term λ is added to H2 during training. sp ‖H2‖1, where λ sp For sparsity regularization weights;

[0165] c. Unified Projection and Output Representation: The final output vector will be compressed to a dimension R consistent with other branches. 128 As a basic structural guidance pathway:

[0166]

[0167] This feature will participate in multi-branch weighted fusion, providing a benchmark for higher-order representations; the benefits of doing so include asymmetry and structural constraints: compared to Φ MLSE and Φ TEC Higher-order feature paths form structural asymmetric relationships, reducing representation space redundancy; multi-layer perturbation compression chains: both consecutive layers employ... Gating mechanism to enhance response suppression to low-confidence features; Guided activation tunability: using ψ γ Instead of conventional ReLU or Swish, the function allows for adjustable response curvature, enhancing the network's ability to distinguish signals in different intervals; controllable sparse compression: adding adjustable regularization constraints ‖·‖1 controls the output sparse density, which helps subsequent classifiers focus on salient substructures.

[0168] F.ΨAGIF Multi-Branch Guided Importance Fusion Mechanism: This step designs an adaptive guided weight fusion mechanism (AGIF) to integrate three heterogeneous substructure paths. The importance weights of each element are dynamically learned to achieve structure-aware feature fusion.

[0169] Design an adaptive guided weight fusion mechanism to integrate the three heterogeneous substructure paths. The system dynamically learns the importance weights of each element to achieve structure-aware feature fusion.

[0170] Set the output of each sub-path For each branch k, the guiding weight ω k ∈R N×128 The generation method is as follows:

[0171]

[0172] in: R 128 →R 64 ρ(·) represents the local nested structure compression mapping of each branch; ρ(·) represents the normalized perceptron function, such as LayerNorm or a custom statistical normalization function; σ(·) is the Sigmoid activation function, which implements element-level gating. The trainable guiding matrix is ​​mapped back to the weight dimension;

[0173] The final fused feature output is:

[0174]

[0175] This approach offers advantages over static weighted fusion: traditional methods use fixed weights or averaging, while this mechanism introduces structural projection, gated response, and normalization adjustment; cross-branch dimensional consistency alignment: ensuring that the outputs of each sub-path are unified to R after guided modulation. 128 The guided response is trainable and interpretable: ω k It can be used for subsequent analysis of branch importance and feature contribution regions.

[0176] G.Ψ DEC Nonlinear discriminant mapping and normalized response output mechanism: This step designs a normalized discriminative encoder classifier (DEC) to complete the transformation from the fused feature space. The final decoding mapping for multi-class label prediction;

[0177] a. Construction of discriminative feature mapping path: First, nonlinear response and perturbation suppression processing is applied to the fused features to enhance robustness to edge samples and low-confidence discriminative regions:

[0178]

[0179] Where: ψ γ To guide the nonlinear response function, such as To map a linear projection to 128 dimensions; This is a disturbance response suppression mechanism;

[0180] b. Multi-class probability normalization output: The final output is mapped to the probability space R through the normalized softmax function NormSoftmax with a temperature parameter. N×C Where C is the number of categories:

[0181]

[0182] in: For classification projection head; τ∈(0,1] is the temperature adjustment parameter used to adjust the entropy sensitivity of the probability distribution, usually taken as τ=0.5; output This represents the prediction confidence level of each sample for each class; the advantage of doing so is robustness in discriminating path structure: leveraging the DERI suppression mechanism and ψ γ Activation provides dynamic correction capabilities for low-quality regions; temperature-controlled normalized mapping enhances the confidence and discrimination of the model by adjusting τ; it is compatible with knowledge distillation / uncertainty modeling: this structure is naturally suitable for teacher prediction transfer or confidence assessment tasks within the distillation framework.

[0183] H. Modeling of Imbalanced Multi-Class Distribution Perceptual Loss: For classification tasks with a severe imbalance in the number of samples across multiple classes, traditional cross-entropy loss struggles to balance the discriminative power between the backbone class and the long-tail class. Therefore, a weighted sparse distribution perceptual loss function is proposed, integrating class weights, output sparsity constraints, and numerical stability adjustment, defined as follows:

[0184]

[0185] Where: N is the number of samples, and C is the number of categories; y represents the predicted probability distribution of sample i; i ∈{1,…,C} are the true labels; For category y i The weighting factor is usually defined as: f y Represents the sample frequency of category y; ∈ = 10 -7 The first term is a stability constant to avoid log(0); the second term is an l1 sparsity regularization to suppress overly averaged prediction tendencies; μ∈R + For hyperparameters, adjust the sparsity constraint strength, with μ = 0.01 to 0.1 recommended. The advantages of doing so include: inverse weighting of class frequencies: suppressing frequent classes and strengthening the learning of low-frequency classes; sparse control of the prediction space: encouraging the model to produce more explicit and separate response probabilities through the l1 term; and enhanced stability: avoiding the risk of gradient explosion and improving robustness to anomalous samples.

[0186] I. Disturbance-Aware Training Temperature Control Scheduling and Multimodal Self-Stabilizing Update Mechanism: To address common issues in deep anomaly flow discrimination models, such as convergence stagnation, exacerbated overfitting, and training oscillations during training with long-tailed categories, multimodal structures, and high complexity, this step designs a multi-strategy controller with disturbance response, self-stabilizing adjustment, and modal coordination capabilities. Its core comprises the following three innovative sub-modules:

[0187] a. Disturbance temperature control controller This module dynamically monitors the error change rate during training and calculates the perturbation temperature index T accordingly. dyn (t) is used to reflect the current training stability. When the perturbation amplitude is large, the system will automatically reduce the parameter update rate γ. t This system implements "cooling-type scheduling," effectively suppressing phenomena such as severe jitter, performance fluctuations, or early overfitting. As the disturbance converges and stabilizes, the system gradually restores its normal update rate, ensuring convergence efficiency. This temperature control mechanism possesses continuous adjustability and numerical stability, and is the core fundamental module of this training controller.

[0188] b. Multi-scale convergence prediction and structural backflow module This module constructs a sliding window observation sequence. The system predicts the trend of validation set loss in the next few rounds based on methods such as exponential moving average. If the prediction results indicate that the performance is about to deteriorate, the system will automatically trigger the structure backflow mechanism, extract the nearest neighbor perturbation version from the historical best weights, and perform local weight backtracking and fine-tuning training on the specified structure or branch. This strategy provides reversible self-repair capability without destroying the current convergence path, and significantly improves the training robustness under complex structures.

[0189] c. Construct a modal transition cooling mechanism To address the issue of gradient freezing or performance degradation in some branches during the later stages of training for multi-branch or multi-modal structures, this mechanism designs a dynamically activated modality jump control process; specifically, it includes: detecting the current submodality. If the training state is such that there are no updates for a long period of time, temporarily freeze the parameter updates of that modality and switch the training objective to another submodality. like If the error trend recovers and improves, the freeze is lifted and the training path is restored; this mechanism effectively maintains the collaborative training capability of the multimodal system throughout the entire cycle.

[0190] J. Model Parameter Freezing and Deployment Export Construction: After completing the full-cycle multi-round training scheduling, the system freezes the current optimal discriminant structure and its weight set to construct the final deployable discriminant model.

[0191] The application process of the ERDF-Net network proposed in this invention during the actual deployment phase includes the following core steps:

[0192] Step 1: Traffic Data Acquisition and Preprocessing

[0193] The system collects raw network traffic data in real time and extracts key features; it performs One-Hot encoding on character fields and bit-extended encoding on numeric fields; it normalizes and tunes all features to ensure consistent input data dimensions and range; and it outputs a normalized feature tensor as input for model inference.

[0194] Step 2: Model Loading and Inference Execution

[0195] The pre-trained ERDF-Net anomaly detection model is invoked; the pre-processed input data is fed into the model, and feature extraction and fusion calculations are performed layer by layer; the final anomaly prediction result is generated through the discriminative output of the fusion branch.

[0196] experiment:

[0197] This invention uses the publicly available UNSW-NB15 and NSL-KDD datasets for experimental verification. The UNSW-NB15 dataset, constructed by the UNSW Canberra Labs, covers normal traffic and nine typical attack behaviors, realistically simulating modern network environments. The NSL-KDD dataset is an improved version of the KDD CUP99 dataset, removing redundant samples and including attack types such as DoS, Probe, R2L, and U2R, as well as normal traffic, resulting in a more balanced data distribution.

[0198] Tables 1-4 present some of the experimental results, where the four evaluation metrics—Accuracy, Recall, Precision, and F1-Score—are used, with higher values ​​indicating better performance. Experimental results and performance verification: Experiments on the UNSW-NB15 dataset show that in binary classification tasks, the accuracy of the proposed method reaches 90.91%, higher than baseline methods such as Random Forest (86.97%), Multilayer Perceptron (85.78%), and GRU (84.42%). In multi-class classification tasks, the accuracy reaches 78.87%, better than models such as CNN-BiLSTM (78.18%) and LSTM+Attention (75.05%). Experiments on the NSL-KDD dataset show that in binary classification tasks, the accuracy reaches 83.89%, significantly better than traditional methods and existing deep learning models. In multi-class classification tasks, the accuracy is 79.57%, with a macro-average F1 score of 0.615, demonstrating high detection accuracy in large-scale attack categories such as DoS and Probe.

[0199] Experimental results show that the multi-branch deep learning model of this invention exhibits excellent performance in both binary and multi-class classification tasks, especially achieving high precision and recall in mainstream attack categories (such as DoS, Generic, and Exploits). However, in minority class attacks (such as R2L, U2R, and Backdoor), the recall rate is relatively low due to the scarcity of samples. To address this, future research can further improve the detection capability for minority class attacks by introducing data augmentation, resampling, and cost-sensitive learning mechanisms.

[0200] Table 1. Binary classification results of UNSW_NB15

[0201]

[0202]

[0203] Table 2 UNSW_NB 15 Multiclass Classification Results

[0204] Model Accuracy Recall Precision F1-Score RandomForest 0.7537 0.5309 0.4811 0.4543 MLP 0.7528 0.4787 0.4882 0.4493 GRU 0.6264 0.4194 0.5127 0.3278 LSTM+Attention 0.7505 0.4266 0.3904 0.3612 CNN+BiLSTM 0.7818 0.4437 0.4752 0.4335 ProposedModel 0.7887 0.5122 0.4804 0.4608

[0205] Table 3 NSL-KDD binary classification results

[0206] Model Accuracy Recall Precision F1-Score RandomForest 0.7598 0.8062 0.7855 0.7584 MLP 0.8109 0.8358 0.8303 0.8108 GRU 0.8184 0.8408 0.8369 0.8184 LSTM+Attention 0.8216 0.8426 0.8396 0.8216 CNN+BiLSTM 0.8068 0.8262 0.8320 0.8066 ProposedModel 0.8389 0.8546 0.8536 0.8389

[0207] Table 4. NSL-KDD Multiclass Classification Results

[0208] Model Accuracy Recall Precision F1-Score RandomForest 0.7378 0.8202 0.4720 0.4711 MLP 0.7717 0.7965 0.5690 0.5877 GRU 0.7897 0.5331 0.6761 0.5298 LSTM+Attention 0.7851 0.5312 0.8204 0.5267 CNN+BiLSTM 0.7974 0.6212 0.6613 0.5648 ProposedMethod 0.7957 0.5983 0.8345 0.6150

Claims

1. A method for identifying multi-branch deep anomaly flow based on a disturbance response control mechanism, characterized in that, Includes the following steps: A. Data preprocessing and perturbation-compatible coding modeling a. Raw traffic collection and feature initialization: Collect multiple types of traffic samples. Each sample contains multiple category features and numerical features. All samples are used to form an initial raw feature set X0, where each sample contains multi-dimensional information with fixed dimensions. b. Perturbation-compatible one-hot encoding for categorical fields: Perturbation-compatible one-hot encoding is performed on all categorical fields. The specific encoding process includes constructing a standard category set C and introducing an additional category c. unk Representing unknown or noisy inputs, each category field is encoded as a sparse vector of dimension |C|+1. After multi-field encoding, a tensor X is generated. cat ; c. Bit-level expansion mapping for highly dynamic numerical fields: For numerical fields with drastically fluctuating value ranges, a perturbation decomposition-based bit expansion strategy is introduced. The maximum number of digits *d* for each field is calculated, and each numerical value is mapped to a *d*-bit integer array, forming a high-dimensional sparse tensor. Multiple fields are expanded and then merged to form tensor X. digit ; d. Multi-source feature fusion and structural splicing: The original numerical tensor X is fused and structured. num Encoding tensor X cat With position expansion tensor X digit Perform a concatenation operation to generate an input tensor with uniform structural features: X=Concat(X num ,X cat ,X digit ),X∈R N×D Where N represents the number of samples, and D is the dimension after fusion; e. Input tensor normalization and tuning: Constructing the normalization control operator T norm This is used to perform standard normalization on the input tensor, ultimately yielding a normalized tensor. The normalization processing strategy is defined as follows: Where, μ∈R 1×D Let σ represent the feature mean calculated along the sample dimension, where σ∈R 1×D Let X be the standard deviation, ∈ be a numerically stable constant, and the output tensor X retains dimensional information to obtain a normalized tensor. B. Initialization and multi-path parallel construction of heterogeneous representation subnetworks: Normalized feature tensors... As a unified input interface, three types of functional sub-network structures are injected simultaneously to construct a multi-path response space. The sub-network structures include Φ MLSE Branch, Φ TEC Branches and Φ BFC Branching, output feature Z k The output feature Z k The dimensions are unified to 128, which facilitates structurally consistent fusion and subsequent joint training. The Φ MLSE The branch is constructed using a multi-layer nonlinear activation mapping, a channel saliency enhancement mechanism, and a local context interaction module structure, with an input dimension of R. N×D The final output is mapped to R. N×128 , denoted as: The Φ TEC The branch is constructed based on pseudo-structure modeling and cross-embedding perception mechanism, firstly by... Project to L∈E dimensions, then rearrange into a structure tensor R. N×L×E After being nested with attention structures, it is compressed to R N×128 , denoted as: The Φ BFC The branch adopts a two-layer structure compression + sparse modulation mechanism, wherein Φ BFC Branches are based on the architecture-aware path, directly... Process into low-dimensional discriminant vectors: C.Φ MLSE Branched multilayer nonlinear expansion and explicit response channel modulation mechanism: a. First-layer structure expansion and response control: First, the normalized input tensor... Mapping to an intermediate higher-order space via linear transformation Let D1 = 512: Subsequently, a self-stabilizing nonlinear activation ψ is introduced. swish With the perturbation-weighted mechanism, the response regulation tensor H1 is constructed as follows: in For Dropout operations, ⊙ indicates element-wise weighting by channel dimension; b. Explicit Disturbance Response Weighting Module definition The module calculates the channel disturbance response value per channel, assuming the input... Let represent the intermediate feature vector of the i-th sample, and its channel control vector is defined as follows: For each channel c: This represents a channel-specific perturbation mapping network; η is the perturbation amplitude hyperparameter; α controls the compression ratio; φ c (·) is a nonlinear transderivative function based on variance adjustment; c. Second-layer feature expansion and re-modulation: H1 is further processed and projected onto... Let D2 = 256, and implement perturbation response control symmetrical to the first layer: d. Enhanced contextual interaction and structural convergence further reduce the dimensionality of the feature map H2 to Set D3 = 128 and send it to the local context interaction module: A lightweight local attention mechanism is employed to uncover local structural consistency and improve discriminative power; D.Φ TEC Branch-local perception and global structure coupling mapping mechanism: To model the inter-layer dependencies of input features in a pseudo-structured space, a table structure enhancement coupling module is proposed. Through structure-aware embedding rearrangement, multi-head interactive aggregation, perturbation compression, and nested feedback mechanisms, the original tensor is processed... Perform multi-granularity structure modeling, and the entire process output is as follows: a. Subspace structure activation embedding via linear embedding Map the input to an L·E-dimensional pseudo-structure space, where L=6 is the structure step size and E=6 is the embedding dimension per step: Where ψ ReLU (·) represents the ReLU activation function. This means reorganizing the flattened vector into L structural sub-units, each sub-unit having dimension E; b. Structurally Coupled Interaction Module Employing a multi-head coupled attention mechanism With Feedforward Network Where H=4 represents the number of structure-sensing heads, each... All are defined as: Note the weight matrix. It can be trained independently, enhancing the ability to model inter-channel responses; Then through a structural feedback feedforward network Input structure delay and hierarchical feature interaction: in For a linear mapping, ψ GELU It is a higher-order activation function; c. Disturbance response compression and regularized suppression Introducing the disturbance control compression mechanism DERI to H ff Random gating and structural response modulation are applied to each structural unit: ∈ i ~Bernoulli(1-α i ), Where: α i The dynamic sparse control law can be adjusted through training or statistical structural complexity. For shared perturbation projection networks; output maintains tensor shape invariance, H deri ∈R N×L×E ; d. Unify the dimensional projection output, flattening the structure tensor to R. N×(L·E) And compressed into a unified fusion dimension: The compression operation ensures consistency with Φ MLSE ,Φ BFC The consistent output dimensions of the branches facilitate subsequent branch merging. E.Φ BFC Branched dual-channel perturbation compression mapping network: a. Basic structure projection and disturbance gating on normalized feature input Apply initial structure mapping And via the bootstrap activation function Adjust response range: in: For trainable fully connected linear mappings It is a sigmoid-like monotonic response function with adjustable curvature, wherein γ1>0; For perturbation compression unit; b. Compressed mapping and sparsity control further map H1 to a low-dimensional space and introduce structural-level L1 sparsity constraints to control invalid responses: in: Activation curve The curvature is slightly less than γ1 for smoother post-response control; a term λ is added to H2 during training. sp ‖H2‖1, where λ sp For sparsity regularization weights; c. Unified Projection and Output Representation: The final output vector will be compressed to a dimension R consistent with other branches. 128 As a basic structural guidance pathway: This feature will participate in multi-branch weighted fusion, providing a benchmark for higher-order representations; F.ΨAGIF multi-branch guided response weighted fusion mechanism: Design an adaptive guided weight fusion mechanism to integrate the three heterogeneous substructure paths. The system dynamically learns the importance weights of each element to achieve structure-aware feature fusion. Set the output of each sub-path For each branch k, the guiding weight ω k ∈R N×128 The generation method is as follows: in: This represents the local nested structure compression mapping of each branch; ρ(·) represents the normalized perceptron function, such as LayerNorm or a custom statistical normalization function; σ(∈) is the Sigmoid activation function, which implements element-level gating; The trainable guiding matrix is ​​mapped back to the weight dimension; The final fused feature output is: G.Ψ DEC Nonlinear discriminant mapping and normalized response output mechanism: a. Discriminative feature mapping path construction: First, nonlinear response and perturbation suppression processing is applied to the fused features to enhance robustness to edge samples and low-confidence discriminative regions. Where: ψ γ To guide the nonlinear response function, such as To map a linear projection to 128 dimensions; This is a disturbance response suppression mechanism; b. Multi-class probability normalization output: The final output is mapped to the probability space R through the normalized softmax function NormSoftmax with a temperature parameter. N×C Where C is the number of categories: in: For classification projection head; τ∈(0,1] is the temperature adjustment parameter used to adjust the entropy sensitivity of the probability distribution, usually taken as τ=0.5; output This represents the prediction confidence level of each sample for each category; H. Modeling of impaired multi-class perceptual loss: We propose a weighted sparse distribution-aware loss function that integrates class weights, output sparsity constraints, and numerical stability adjustment, defined as follows: Where: N is the number of samples, and C is the number of categories; y represents the predicted probability distribution of sample i; i ∈{1,…,C} are the true labels; For category y i The weighting factor is usually defined as: f y Represents the sample frequency of category y; ∈ = 10 -7 The first term is a stability constant to avoid log(0); the second term is an l1 sparsity regularization to suppress overly averaged prediction tendencies; μ∈R + As a hyperparameter, it adjusts the sparsity constraint strength; I. Disturbance-aware training temperature control scheduling and multimodal self-stabilizing update mechanism: Multi-strategy controller with disturbance response, self-stabilizing adjustment and modal coordination capabilities Its core comprises the following three innovative sub-modules: a. Disturbance temperature control controller This module dynamically monitors the error change rate during training and calculates the perturbation temperature index T accordingly. dyn (t) is used to reflect the current training stability. When the perturbation amplitude is large, the system will automatically reduce the parameter update rate γ. t ; b. Multi-scale convergence prediction and structural backflow module This module constructs a sliding window observation sequence. c. Construct a modal jump cooling mechanism Specifically, this includes: detecting the current submodal. The training state is temporarily frozen; parameter updates for this mode are temporarily frozen, and the training objective is switched to another sub-mode. like If the error trend recovers and improves, then the freeze is lifted and the training path is restored; J. Model Parameter Freezing and Deployment Export Construction: After completing the full-cycle multi-round training scheduling, the system freezes the current optimal discriminant structure and its weight set to construct the final deployable discriminant model.

2. The multi-branch deep anomaly flow discrimination method based on disturbance response control mechanism according to claim 1, characterized in that: The multi-dimensional information includes communication direction, protocol identifier, access status, and number of bytes.