Malicious Command Line Detection Method Based on Multi-Dimensional Feature Learning and Feature Focusing

By adopting CharBERT model and multi-dimensional feature learning in malicious command line detection, combining attention masking mechanism and multi-layer perceptron, multi-level features and fine-grained features of command line parameters are solved, and the existing detection methods are delayed and limited in identifying new threats is achieved, achieving more efficient and accurate malicious command line detection.

CN119918012BActive Publication Date: 2025-06-17ZHEJIANG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510346012.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-06-17
Estimated Expiration
2045-03-24

AI Technical Summary

Technical Problem

Existing malicious command line detection methods have delays and limitations in identifying new threats, making it difficult to effectively deal with the ever-changing new malicious command line variants.

Method used

The CharBERT model is used to output high-dimensional features, combined with attention masking mechanism and multi-layer perceptron, and through multi-dimensional feature learning and regional focus attention mechanism, multi-level features and fine-grained features of command line parameters are extracted.

Benefits of technology

It significantly improves the accuracy and robustness of malicious command line detection, and can achieve excellent detection performance in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119918012B_ABST
    Figure CN119918012B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical fields of network security and machine learning, and discloses a malicious command line detection method based on multi-dimensional feature learning and feature focusing, including outputting high-dimensional features of the parameters of the command line to be detected; respectively inputting the basic features into a plurality of dilated convolutions with dynamic dilation rates to obtain a plurality of branch features; fusing the branch features with the basic features to obtain preliminary fused features; using an activation function root to obtain hierarchical progressive fused features; performing an adaptive average pooling operation on the hierarchical progressive fused features; using a multi-layer perceptron to extract attention weights, and fusing the attention weights with the hierarchical progressive fused features to obtain focused features; performing weighted fusion on all the focused features to obtain the final command line features, and using a detection model to learn the command line features and output a detection result, where the detection result is a malicious command line or a normal command line. The present invention comprehensively and effectively extracts the features of command line parameters and improves the detection accuracy of command lines.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of network security and machine learning, and particularly relates to a malicious command line detection method based on multi-dimensional feature learning and feature focusing. Background Art

[0002] The attack targets of malicious command lines include the integrity of the system, user privacy, and the stability of network infrastructure. Traditional detection methods (such as blacklist-based detection, heuristic detection, and custom rule detection, etc.) have delays and limitations in identifying new threats. This is because these methods rely on known command line features or manual updates and are difficult to effectively cope with continuously changing new malicious command line variants.

[0003] With the progress of machine learning technology, machine learning-based malicious command line detection methods have gradually received attention. Machine learning technologies used for malicious command line detection include traditional shallow learning technologies and deep learning technologies. In particular, deep learning shows high accuracy and generalization ability in automatically learning complex and non-linear hidden features. Currently, deep learning models used for malicious command line detection include CNN (Convolutional Neural Network), RNN (Recurrent Neural Network), and Transformer, etc. However, the standard Transformer faces some specific challenges in malicious command line detection: (1) The token-based input mechanism limits its ability to capture character-level information, which is crucial for detecting subtle changes in command lines; (2) The Transformer is not as effective as the CNN in capturing local patterns, which is key for identifying malicious substructures in command lines; (3) The Transformer lacks the ability to recognize the inherent hierarchical structure in command line instructions. Therefore, existing methods cannot comprehensively and effectively extract the features of command line parameters, resulting in limited detection accuracy of the classification model. Summary of the Invention

[0004] The purpose of the present invention is to provide a malicious command line detection method based on multi-dimensional feature learning and feature focusing, which can comprehensively and effectively extract the features of command line parameters and improve the detection accuracy of command lines.

[0005] To achieve the above purpose, the technical solution adopted by the present invention is as follows:

[0006] A malicious command line detection method based on multi-dimensional feature learning and feature focusing, the malicious command line detection method based on multi-dimensional feature learning and feature focusing includes:

[0007] Using the CharBERT model to output high-dimensional features for the parameters of the command line to be detected;

[0008] Based on the attention mask mechanism, basic features are obtained from high-dimensional features. The basic features are respectively input into multiple dilated convolutions with dynamic dilation rates to obtain multiple branch features;

[0009] For each branch feature, the branch feature is fused with the basic feature to obtain a preliminary fusion feature;

[0010] An activation function is used to output dynamic weight coefficients according to the preliminary fusion features, and all preliminary fusion features are weighted and fused based on the dynamic weight coefficients to obtain hierarchical progressive fusion features;

[0011] An adaptive average pooling operation is performed on the hierarchical progressive fusion features to generate pooling features of different dimensions;

[0012] For each dimension of the pooling features, a multi-layer perceptron is used to extract attention weights, and the attention weights are fused with the hierarchical progressive fusion features to obtain focused features of each dimension;

[0013] All dimensions of the focused features are weighted and fused to obtain the final command line features. The detection model learns the command line features and outputs the detection result, where the detection result is a malicious command line or a normal command line.

[0014] The following also provides several optional methods, which are not additional limitations to the above overall solution, but only further supplements or optimizations. Without technical or logical contradictions, each optional method can be combined with the above overall solution alone, or multiple optional methods can be combined with each other.

[0015] Preferably, the method for determining the dynamic dilation rate is as follows:

[0016] ;

[0017] Wherein, represents the dynamic dilation rate of the th dilated convolution, is the maximum operation, represents the learning parameter of the th dilated convolution, represents the high-dimensional feature, represents the Frobenius norm, represents the floor operation.

[0018] Preferably, the obtaining of the basic features from the high-dimensional features based on the attention mask mechanism includes:

[0019] The high-dimensional features are input into a depthwise separable convolution to obtain depth convolution features;

[0020] Convert the high-dimensional features into an attention mask using an activation function;

[0021] Take the multiplication result of the depth convolution features and the attention mask as the basic features.

[0022] Preferably, the fusing the branch features and the basic features to obtain a preliminary fused feature includes: reading a fusion coefficient and applying it to the basic features, and then adding the result to the branch features as the preliminary fused feature.

[0023] Preferably, the adopting an activation function to output a dynamic weight coefficient according to the preliminary fused feature includes:

[0024] If the activation function is the softmax normalization function, the dynamic weight coefficient is calculated as follows:

[0025] ;

[0026] In the formula, represents the dynamic weight coefficient corresponding to the th preliminary fused feature, represents the number of preliminary fused features, represents the exponential function with the natural constant as the base;

[0027] When or , represents the learning parameter of the basic features, represents the basic features; otherwise represents the learning parameter of the th preliminary fused feature, represents the th preliminary fused feature, represents the learning parameter of the th preliminary fused feature, represents the learning parameter of the th preliminary fused feature.

[0028] Preferably, the weighted fusion of all preliminary fused features based on the dynamic weight coefficient to obtain a hierarchical progressive fused feature includes:

[0029] ;

[0030] In the formula, is the hierarchical progressive fused feature, represents the number of preliminary fused features, represents the dynamic weight coefficient corresponding to the th preliminary fused feature, represents the residual ratio parameter, represents the high-dimensional features; when When it represents a basic feature; otherwise it represents the th preliminary fusion feature.

[0031] Preferably, the attention weights are fused with the hierarchical progressive fusion features to obtain the focused features for each dimension, including:

[0032] ;

[0033] In the formula, represents the focused feature of the th dimension, is the attention weight fusion ratio, is the hierarchical progressive fusion feature, represents the attention weight of the th dimension.

[0034] Preferably, the weighted fusion of the focused features for all dimensions includes:

[0035] Performing an upsampling operation on the focused feature for each dimension to obtain an upsampled feature;

[0036] Assigning a weighted weight to the upsampled feature to obtain a weighted feature;

[0037] Using a summation function to sum the weighted features for all dimensions and output a command line feature.

[0038] A malicious command line detection method based on multi-dimensional feature learning and feature focusing provided by the present invention detects malicious command lines based on command line parameters. By combining character-level embedding and multi-level feature extraction, it can more comprehensively capture the fine-grained features and context semantic relationships in the command line, thereby improving the detection ability for complex malicious command lines. Using multi-dimensional feature learning and regional focusing attention mechanism, local and global features are captured through multi-dimensional convolution with different dilation rates, and important region features are dynamically weighted through the regional focusing attention mechanism, realizing the comprehensive and effective extraction of the features of command line parameters, thus significantly improving the detection accuracy and robustness, and being able to achieve excellent detection performance in complex scenarios such as unbalanced, multi-classification, cross-dataset testing, and adversarial sample attacks. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 is a flowchart of a malicious command line detection method based on multi-dimensional feature learning and feature focusing of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0041] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments, and are not intended to limit the present invention.

[0042] As Figure 1 shown, this embodiment provides a malicious command line detection method based on multi-dimensional feature learning and feature focusing, including the following steps:

[0043] (1) Feature extraction based on command line parameters: Use the CharBERT model (character-aware pre-trained model) to output high-dimensional features for the parameters of the command line to be detected.

[0044] This embodiment uses a pre-trained CharBERT model to encode the command line argument (Command Line Arguments, abbreviated as CLA) sequence to generate high-dimensional features. Specifically, the CharBERT model extends character-aware embeddings on the basis of the BERT model, and can capture important context information from the characters and sub-words in the command line arguments. Each command line argument is decomposed into multiple character-level or sub-word-level tokens, and these tokens are embedded and represented in a two-channel manner. And the output is extracted from the multi-layer Transformer encoder of the CharBERT model. The output of each layer represents multi-dimensional information, from low-level character features to high-level context semantics. A weighted strategy based on the attention mechanism is used to dynamically weight and fuse the outputs of each layer to generate a multi-layer feature matrix, thereby forming a comprehensive representation of the command line arguments and capturing fine-grained and global malicious patterns in the command line arguments. Finally, the embedded representation and the multi-layer feature matrix are fused to obtain high-dimensional features.

[0045] (2) Multi-dimensional feature learning: Capture multi-dimensional local details and structural relationships through multi-dimensional feature learning.

[0046] (2.1) Adaptive multi-dimensional convolution: Based on the attention mask mechanism, obtain the basic features according to the high-dimensional features, and input the basic features into multiple dilated convolutions with dynamic dilation rates respectively to obtain multiple branch features.

[0047] Depthwise separable convolution (DSConv) and dilated convolution with dynamic dilation rate are used to capture multi-dimensional feature information. It introduces a dynamic dilation rate mechanism that automatically adjusts the dilation rate according to the complexity of the input features, effectively capturing multi-scale command line features. The dynamic dilation rate is determined by the following formula:

[0048] ;

[0049] where, represents the dynamic dilation rate of the -th dilated convolution, is the maximum operation, represents the learnable parameter of the -th dilated convolution. The initial value of the learnable parameter is randomly initialized and continuously optimized during the training process. represents the high-dimensional feature, represents the Frobenius norm, which is used to measure the overall strength of the input features. represents the floor operation, , is the number of dilated convolutions. This formula ensures that the dynamic dilation rate is at least 1 and is adaptively adjusted according to the strength of the input features, enabling the model to flexibly focus on command line features at different scales.

[0050] Formally, the high-dimensional feature is represented as , where is the number of channels, is the height, is the width. This process first applies the DSConv operation to , that is, inputs the high-dimensional feature into the depthwise separable convolution to obtain the depth convolution feature , represents the DSConv operation; then an attention mask mechanism is introduced to enable the model to adaptively focus on the key character features in the command line, improving the sensitivity to malicious features. In this embodiment, the activation function (such as the Sigmoid function) is used to convert the high-dimensional feature into an attention mask, and finally the multiplication result of the depth convolution feature and the attention mask is taken as the base feature. The formula is as follows:

[0051] ;

[0052] In the formula, represents the base feature, represents the attention mask, represents the activation function, is the learnable parameter of the activation function.

[0053] After obtaining the base feature After that, convolution operations with a dynamic dilation rate are used to process these features, which are expressed by the following formula:

[0054] ;

[0055] In the formula, is the th branch feature, , is the dilated convolution operation, which is a special convolution operation that enlarges the receptive field by inserting "holes" between the elements of the convolution kernel without increasing the number of parameters. is the dynamically determined dilation rate that controls the size of the holes. Different from the traditional dilated convolution that uses a fixed dilation rate, the dilation rate in the present invention is dynamically adjusted according to the complexity of the input features, enabling the model to adaptively focus on command-line features at different scales and enhancing the ability to capture complex malicious command-line patterns.

[0056] (2.2) Hierarchical progressive feature fusion: Multidimensional features are hierarchically progressively fused through 1×1 convolution with dynamic weights to generate a feature representation with rich semantic information. Feature fusion adopts weighted non-linear transformation to enhance the model's ability to recognize subtle features of malicious command lines. Context information from multiple dimensions is integrated through residual connections with an attention mechanism.

[0057] (2.2.1) For each branch feature, the branch feature is fused with the base feature to obtain a preliminary fused feature. In this embodiment, the fusion coefficient is read and applied to the base feature, and then added to the branch feature as the preliminary fused feature, which is expressed by the following formula:

[0058] ;

[0059] In the formula, represents the th preliminary fused feature, , is the fusion coefficient corresponding to the th preliminary fused feature and is a learnable parameter, serves as the fusion coefficient for each branch feature and the base feature to achieve adaptive feature fusion and enhance the model's ability to distinguish different types of malicious command lines.

[0060] (2.2.2) An activation function is used to output dynamic weight coefficients based on the preliminary fused features, and all preliminary fused features are weighted and fused based on the dynamic weight coefficients to obtain hierarchical progressive fused features.

[0061] In this embodiment, the softmax normalization function is introduced as the activation function to achieve the adaptive weight allocation of the features of each branch, enabling the dynamic adjustment of the importance of each branch according to the characteristics of the input command line. The dynamic weight coefficient is calculated as follows:

[0062] ;

[0063] In the formula, represents the dynamic weight coefficient corresponding to the th preliminary fusion feature, represents the number of preliminary fusion features, represents the exponential function with the natural constant as the base.

[0064] When or , represents the learnable parameter of the base feature, represents the base feature; otherwise represents the learnable parameter of the th preliminary fusion feature, represents the th preliminary fusion feature, represents the th learnable parameter of the preliminary fusion feature, represents the th preliminary fusion feature.

[0065] In the hierarchical progressive fusion, this embodiment uses the weighted sum with the weight coefficient to replace the traditional simple sum, and adds the parameter to control the residual contribution of the original feature , further enhancing the model's learning ability for complex command line features. The formula is expressed as follows:

[0066] ;

[0067] In the formula, is the hierarchical progressive fusion feature, represents the number of preliminary fusion features, represents the th dynamic weight coefficient corresponding to the preliminary fusion feature, represents the residual ratio parameter, which is set according to experience or experiments. It should be noted that, represents that there are branches, each branch has a dilated convolution and obtains the corresponding branch feature and preliminary fusion feature, and there are corresponding learning parameters, dynamic weight coefficients, etc. All can be understood as dilated convolutions, or can also be understood as Branch feature

[0068] In the multi-dimensional feature learning of this embodiment, by introducing an adaptive weight allocation, a dynamic dilation rate, and an attention mask mechanism, the detection ability of the model for complex malicious command lines is significantly improved, especially showing stronger robustness and precision when dealing with highly obfuscated and deformed command lines. Compared with traditional methods, our hierarchical progressive fusion strategy can more effectively retain and integrate feature information at different levels, reduce information loss, and enhance the generalization ability of the model.

[0069] (3) Region-focus attention mechanism: The spatial region-focus attention mechanism is adopted to differentially weight different regions of the features, highlighting the segments with rich information.

[0070] (3.1) Adaptive average pooling: An adaptive average pooling operation is performed on the hierarchical progressive fusion features to generate pooling features of different dimensions. For example, pooling features of sizes 1×1, 2×2, and 4×4 can be generated. These pooling features are used for subsequent attention calculation to help the model assign weights to different regions of the features.

[0071] (3.2) Attention mechanism: For each dimension of the pooling features, a multi-layer perceptron (MLP) is used to extract the attention weights, and the attention weights are fused with the hierarchical progressive fusion features to obtain the focused features of each dimension. The formula is as follows:

[0072] ;

[0073] In the formula, represents the attention weight of the th dimension, represents the pooling feature of the th dimension, , are the weight parameters of the multi-layer perceptron, which are learnable parameters, , are the bias terms of the multi-layer perceptron, which are learnable parameters, is the activation function, such as the ReLU activation function, is the Sigmoid function, which maps the output to between 0 and 1.

[0074] ;

[0075] In the formula, represents the focused feature of the th dimension, is the attention weight fusion ratio, which controls the fusion ratio of the original feature and the attention-weighted feature and is set according to experience or experiments.

[0076] (3.3) is also performing fusion and weighting: weighting and fusing the focused features of all dimensions to obtain the final command line features, using the detection model to learn the command line features, and outputting the detection result, which is either a malicious command line or a normal command line.

[0077] When weighting and fusing the focused features of all dimensions, first perform an upsampling operation on the focused features of each dimension to obtain the upsampled features; then assign weighting weights to the upsampled features to obtain the weighted features; finally, use the summation function to sum the weighted features of all dimensions and output the command line features. The specific formula is as follows:

[0078] ;

[0079] In the formula, is the command line feature, is the number of pooling scales, that is, the number of generated pooling features, is the weighted weight coefficient of the th pooling feature, satisfying that the sum of the weights is 1,

[0080] is the upsampling operation, which upsamples the features to the original size. In the regional focus attention mechanism, attention weights are adopted to extract key information from the pooling features through a multi-layer perceptron. Apply the attention weights to the original features and control the feature fusion ratio through the parameter

[0081] Integrate the pooling features of different scales, provide a multi-scale perspective, enhance the model's comprehensive understanding ability of the command line, improve the recognition accuracy of malicious command lines, and achieve efficient and accurate malicious command line detection.

[0082] This operation ensures that the most important parts in the feature map are enhanced while weakening irrelevant or noisy information, thereby providing a more accurate feature representation for the final detection model. The detection model refers to a classifier model used for malicious command line detection, such as a Random Forest classifier, a Support Vector Machine (SVM), etc. After obtaining the final feature representation, it is input into the detection model to complete the malicious command line detection, and the output result is "malicious" (malicious command line) or "normal" (normal command line).

[0083] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0084] The above-described embodiments only represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the appended claims.

Claims

1. A malicious command line detection method based on multi-dimensional feature learning and feature focusing, characterized in that: The malicious command line detection method based on multi-dimensional feature learning and feature focusing includes: The CharBERT model is used to output high-dimensional features of the parameters of the command line to be detected; Based on the attention mask mechanism, basic features are obtained according to high-dimensional features, and the basic features are input into multiple dilated convolutions with dynamic dilation rates to obtain multiple branch features; For each branch feature, the branch feature is fused with the basic feature to obtain a preliminary fusion feature; The activation function is used to output dynamic weight coefficients according to the preliminary fusion features, and all preliminary fusion features are weighted fused based on the dynamic weight coefficients to obtain hierarchical progressive fusion features; Perform adaptive average pooling operation on the hierarchical progressive fusion features to generate pooling features of different dimensions; For the pooled features of each dimension, a multi-layer perceptron is used to extract the attention weight, and the attention weight is fused with the hierarchical progressive fusion features to obtain the focus features of each dimension; Perform weighted fusion on the focused features of all dimensions to obtain the final command line features, use the detection model to learn the command line features, and output the detection result, which is a malicious command line or a normal command line; The method of obtaining basic features based on high-dimensional features based on the attention mask mechanism includes: Input high-dimensional features into the depth-separable convolution to obtain deep convolution features; Use activation functions to convert high-dimensional features into attention masks; Take the multiplication result of the deep convolution feature and the attention mask as the basic feature; The attention weight is fused with the hierarchical progressive fusion feature to obtain the focus feature of each dimension, including: ; In the formula, Indicates The focus features of the dimensions, Attention weight fusion ratio, It is a hierarchical progressive fusion feature. Indicates The attention weights of the dimensions.

2. The malicious command line detection method based on multi-dimensional feature learning and feature focusing according to claim 1 is characterized in that: The dynamic expansion rate is determined as follows: ; in, Indicates The dynamic dilation rate of the dilated convolution, To obtain the maximum value, Indicates The learning parameters of the dilated convolution, Represents high-dimensional features, represents the Frobenius norm, Indicates a floor operation.

3. The malicious command line detection method based on multi-dimensional feature learning and feature focusing according to claim 1 is characterized in that: The step of fusing the branch feature with the basic feature to obtain the preliminary fused feature includes: reading the fusion coefficient and applying it to the basic feature, and then adding the fusion coefficient to the branch feature as the preliminary fused feature.

4. The malicious command line detection method based on multi-dimensional feature learning and feature focusing according to claim 1 is characterized in that: The activation function is used to output a dynamic weight coefficient according to the preliminary fusion feature, including: The activation function is a softmax normalization function, and the dynamic weight coefficient is calculated as follows: ; In the formula, Indicates The dynamic weight coefficient corresponding to the initial fusion feature, represents the number of initial fusion features, Represents a natural constant An exponential function with base ; when or hour, represents the learning parameters of the basic features, Indicates the basic feature; otherwise Indicates The learning parameters of the initial fusion features, Indicates Initial fusion features, Indicates The learning parameters of the initial fusion features, Indicates A preliminary fusion feature.

5. The malicious command line detection method based on multi-dimensional feature learning and feature focusing according to claim 1 is characterized in that: The step of weighted fusion of all preliminary fusion features based on dynamic weight coefficients to obtain hierarchical progressive fusion features includes: ; In the formula, It is a hierarchical progressive fusion feature. represents the number of initial fusion features, Indicates The dynamic weight coefficient corresponding to the initial fusion feature, represents the residual scale parameter, Represents high-dimensional features; when hour, Indicates the basic feature; otherwise Indicates A preliminary fusion feature.

6. The malicious command line detection method based on multi-dimensional feature learning and feature focusing according to claim 1 is characterized in that: The weighted fusion of the focus features of all dimensions includes: Perform upsampling operation on the focused features of each dimension to obtain upsampled features; Assign weighted weights to the upsampled features to obtain weighted features; Use the sum function to sum the weighted features of all dimensions and output the command line features.

Citation Information

Patent Citations

  • Target detection method and device based on receptive field enhancement and attention guide aggregation

    CN116580274A

  • Detection model training method and device, code detection method and device and related equipment

    CN118839721A