Malware Family Classification Method and System Based on Dynamic Features and Ensemble Learning
Through a method based on dynamic features and ensemble learning, Bi-LSTM, TextCNN and CNN+LSTM models are used to feature learning on the dynamic API call sequence of malicious code, and an integrated learning model is built, which solves the problem that a single model is difficult to fully learn the inherent information of features and improves the accuracy of malware family classification.
Patent Information
- Application Number
- CN202410576940.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-10
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-05-10
AI Technical Summary
The existing malware family classification model is a single model, and it is difficult to fully learn the inherent information of features, resulting in certain limitations in the classification effect.
Using a method based on dynamic features and ensemble learning, dynamic API call sequences of malicious code are extracted as features, and feature learning is performed through three basic classifiers, namely, long and short-term memory neural network (Bi-LSTM), text convolutional neural network (TextCNN), and convolutional neural network + long and short-term memory neural network (CNN+LSTM), and integrated learning model is constructed in combination with extreme gradient enhancement algorithm.
Classifying the malware family through integrated learning models improves the accuracy of classification and avoids the impact of code obfuscation on feature extraction.
Smart Images

Figure CN119312326B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of malware family classification, and specifically to a malware family classification method and system based on dynamic features and ensemble learning. Background Art
[0002] In recent years, with the continuous progress of information and network technologies, the number and types of malware have increased at an alarming rate. Its attack targets have gradually evolved from individual users in the early stage to enterprises and even countries, and the scope of influence has gradually expanded from traditional personal computers to various networked intelligent terminals including smartphones. Although some existing research work on malicious code detection and classification has made positive progress, the rapid growth of the number of malware variants still has a greater impact on the classification effect of traditional methods. The research on malicious code classification based on artificial intelligence technology has become a new hot spot, and researching efficient malware classification methods has become a very important and challenging task at present.
[0003] Among the newly added malicious codes, most are variants of known malicious codes rather than completely newly generated malicious code files. Among these variants, a large part are family variants of known malicious codes, that is, new types of malware generated after the evolution of functions or anti-detection technologies. Existing work on malware family classification generally extracts its feature data and code fragments, analyzes their homology with known samples, and then infers the family of suspicious malicious samples. In the existing intelligent classification work implemented by means of artificial intelligence technology, the feature extraction is mainly divided into two categories. One is the static features obtained through static analysis methods. For example, the patent named "A Malicious Code Detection Method Based on Static Features" (CN115859290A) completes classification through a fully connected layer and a Softmax layer after alternately processing a static feature array with a first separable convolution and a second separable convolution. However, with the development and application of technologies such as code obfuscation, the role of static features has been greatly limited, and the generalization ability of classifiers based on static features has also decreased sharply. The other is the dynamic features during program runtime obtained through dynamic analysis methods. The features extracted by dynamic analysis pay more attention to characterizing the behavior of the program.
[0004] In addition to the differences in features, most of the existing work still uses a single model in model construction and focuses on improving its performance and effect. However, due to the different performances of different features on different classifiers, it is difficult to comprehensively learn the internal information of features using a single machine learning classification model, which makes the classification effect have certain limitations. Summary of the Invention
[0005] To solve the problem that the existing malware family classification model is a single model, which is difficult to comprehensively learn the intrinsic information of features and has certain limitations, the present invention proposes a malware family classification method and system based on dynamic features and ensemble learning. First, the dynamic API call sequence of malicious code, which is a dynamic feature, is extracted, and then three base classifiers, namely a bidirectional long short-term memory neural network model, a text convolutional neural network model, and a convolutional neural network + long short-term memory neural network model, are set up to perform feature learning on the extracted dynamic features. After that, the extreme gradient boosting algorithm is used to complete the construction of the ensemble learning model to achieve the classification of malware families.
[0006] To achieve the above object, the technical solution of the present invention is as follows:
[0007] The first aspect of the present invention proposes a malware family classification method based on dynamic features and ensemble learning, including the following steps:
[0008] Step 1: Extract the dynamic application programming interface (API) call sequence of each malicious code in the training malware, and summarize the dynamic API call sequences as the original features;
[0009] Specifically, the dynamic API call sequence is extracted through sandbox and BSA sandbox analysis tools to facilitate the subsequent learning of the implicit information in the dynamic API call sequence.
[0010] Step 2: Construct a bidirectional long short-term memory neural network (Bi-LSTM) model as the first base classifier to make full use of context information;
[0011] Step 3: Construct a text convolutional neural network (TextCNN) model as the second base classifier to more efficiently self-learn and extract local features in the text through convolution and pooling operations;
[0012] Step 4: Construct a convolutional neural network (CNN) + long short-term memory neural network (LSTM) model as the third base classifier to extract the spatial features contained in the dynamic API sequence through the convolutional block and capture the long-term dependence information in the dynamic API call sequence through the LSTM layer;
[0013] Step 5: Use the original features as the inputs of the three base classifiers respectively, and iteratively train the three base classifiers through the extreme gradient boosting algorithm. After the training is completed, an ensemble learning classification model is obtained as the malware family classification model;
[0014] Specifically, in this step, K classification and regression trees are established in sequence, and the new decision tree in each iterative learning is added to the ensemble model and weighted averaged to improve the overall prediction performance and achieve a higher prediction accuracy, so as to facilitate the classification of the family of malicious code.
[0015] Step 6: Classify all malicious codes in the target malware through the malicious software family classification model, which provides an important basis for the control and removal of the target malware.
[0016] Further, in Step 2, an attention layer is introduced into the Bi-LSTM model to construct a neural network model of Bi-LSTM and attention mechanism as the first base classifier.
[0017] Specifically, an attention Attention layer is introduced into the Bi-LSTM model, and the output of the Bi-LSTM layer in the Bi-LSTM model is used as the input of the attention layer to construct a neural network model of Bi-LSTM and attention mechanism as the first base classifier, which is convenient for making full use of context information and further focusing on important information.
[0018] Further, in Step 3, the convolutional layer in the TextCNN model includes four convolutional kernels of different sizes, and a softmax layer of normalized exponential function is added after the fully connected layer.
[0019] Specifically, the convolutional layer in the TextCNN model includes four convolutional kernels of different sizes, which is convenient for fully abstracting local features of different views in the sequence. A softmax layer is added after the fully connected layer to convert the output vector of the pooling layer into the required prediction result.
[0020] Further, in Step 4, the third base classifier sequentially includes an embedding layer, two groups of convolutional blocks, a long short-term memory neural network LSTM layer, and a fully connected layer. Each group of convolutional blocks includes a convolutional layer and a pooling layer. The convolutional layer selects one-dimensional convolution, and the pooling layer uses max pooling operation. A softmax is added after the fully connected layer for classification.
[0021] Specifically, the embedding layer performs word embedding processing on the dynamic API call sequence of the input malware sample and converts it into a vector representation for learning family features in the sample file.
[0022] Two groups of convolutional blocks are used for dimensionality reduction of the extracted spatial features.
[0023] The LSTM layer is used to maintain long-term dependence information and capture semantic relationships in the sequence.
[0024] The fully connected layer converts the outputs of the previous layers.
[0025] Further, in Step 5, during training, the multi-classification cross-entropy is selected as the loss function, which is convenient for calculation and can make the ensemble learning classification model have better classification ability.
[0026] Furthermore, in the fifth step, the ensemble learning classification model is optimized by an optimization algorithm, and the optimization algorithm adopts any one of the stochastic gradient descent optimization algorithm, the momentum optimization algorithm, and the adaptive learning rate optimization algorithm, which is convenient for improving the classification accuracy.
[0027] The second aspect of the present invention proposes a malware family classification system based on dynamic features and ensemble learning, including:
[0028] A feature extraction module, which is used to extract the dynamic application programming interface (API) call sequences of each malicious code in the malware, and summarize the dynamic API call sequences as the original features;
[0029] Specifically, the dynamic API call sequences are extracted through sandbox and BSA sandbox analysis tools, which is convenient for learning the implicit information in the dynamic API call sequences next.
[0030] The first base classifier module is used to construct a bidirectional long short-term memory neural network (Bi-LSTM) model as the first base classifier, which is convenient for making full use of context information;
[0031] The second base classifier module is used to construct a text convolutional neural network (TextCNN) model as the second base classifier, which is convenient for more efficiently self-learning and extracting local features in the text through convolution and pooling operations;
[0032] The third base classifier module is used to construct a convolutional neural network (CNN) + long short-term memory neural network (LSTM) model as the third base classifier, which is convenient for extracting spatial features contained in the dynamic API sequence through convolutional blocks and capturing long-term dependency information in the dynamic API call sequence through the LSTM layer;
[0033] The ensemble learning classifier module is used to take the original features as the inputs of the three base classifiers respectively, and iteratively train the three base classifiers through the extreme gradient boosting algorithm. After the training is completed, an ensemble learning classification model is obtained as the malware family classification model;
[0034] Specifically, in this step, 3 classification and regression trees are established in sequence, and the new decision trees in each iterative learning are added to the ensemble model and weighted averaged to improve the overall prediction performance and achieve a higher prediction accuracy, which is convenient for realizing the family classification of malicious codes;
[0035] The classification module is used to perform malware family classification on all malicious codes in the target malware through the malware family classification model, which provides an important basis for the control and removal of the target malware.
[0036] Further, in the first base classifier module, an attention layer is introduced into the Bi-LSTM model, and a neural network model combining Bi-LSTM and the attention mechanism is constructed as the first base classifier;
[0037] Specifically, an attention layer is introduced into the Bi-LSTM model. The output of the Bi-LSTM layer in the Bi-LSTM model is used as the input of the attention layer, and a neural network model combining Bi-LSTM and the attention mechanism is constructed as the first base classifier, which is convenient for making full use of context information and further focusing on important information.
[0038] Further, in the second base classifier module, the convolutional layer in the TextCNN model includes four convolutional kernels of different sizes, and a softmax layer of the normalized exponential function is added after the fully connected layer;
[0039] Specifically, the convolutional layer in the TextCNN model includes four convolutional kernels of different sizes, which is convenient for fully abstracting local features of different views in the sequence. A softmax layer is added after the fully connected layer to convert the output vector of the pooling layer into the required prediction result.
[0040] Further, in the third base classifier module, the third base classifier sequentially includes an embedding layer, two groups of convolutional blocks, a long short-term memory neural network (LSTM) layer, and a fully connected layer. Each group of convolutional blocks includes a convolutional layer and a pooling layer. The convolutional layer selects one-dimensional convolution, and the pooling layer uses max pooling operation. A softmax is added after the fully connected layer for classification;
[0041] Specifically, the embedding layer performs word embedding processing on the dynamic API call sequence of the input malware sample, converting it into a vector representation for learning the family features in the sample file;
[0042] Two groups of convolutional blocks are used for dimensionality reduction of the extracted spatial features;
[0043] The LSTM layer is used to maintain long-term dependence information and capture semantic relationships in the sequence;
[0044] The fully connected layer converts the outputs of the previous layers.
[0045] Further, in the ensemble learning classifier module, during training, the multi-class cross-entropy is selected as the loss function, which is convenient for calculation and can make the ensemble learning classification model have better classification ability.
[0046] Furthermore, in the integrated learning classifier module, an optimization algorithm is used to optimize the integrated learning classification model. The optimization algorithm can be any one of the stochastic gradient descent optimization algorithm, the momentum optimization algorithm, and the adaptive learning rate optimization algorithm, which is convenient for improving the classification accuracy.
[0047] Through the above technical solutions, the beneficial effects of the present invention are as follows:
[0048] The present invention extracts feature data through a dynamic analysis method, avoiding the influence of technologies such as code obfuscation on feature extraction. By setting three base classifiers, namely the Bi-LSTM model, the TextCNN model, and the CNN+LSTM model, comprehensive learning is carried out on the malicious software API sequence features dynamically extracted. Through the extreme gradient boosting algorithm, the three base classifiers are iteratively trained and merged to complete the construction of the integrated learning classification model. By classifying malicious software families through integrated learning, the classification accuracy is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 is a flowchart of the method for classifying malicious software families based on dynamic features and integrated learning of the present invention.
[0050] Figure 2 is a flowchart of the extraction of malicious code API sequences in the method for classifying malicious software families based on dynamic features and integrated learning of the present invention.
[0051] Figure 3 is a schematic diagram of the structure of the base classifier based on Bi-LSTM+Attention in the method for classifying malicious software families based on dynamic features and integrated learning of the present invention.
[0052] Figure 4 is a schematic diagram of the structure of the base classifier based on TextCNN in the method for classifying malicious software families based on dynamic features and integrated learning of the present invention.
[0053] Figure 5 is a schematic diagram of the structure of the base classifier based on CNN+LSTM in the method for classifying malicious software families based on dynamic features and integrated learning of the present invention.
[0054] Figure 6 is a schematic diagram of the learning curves of three algorithms in the method for classifying malicious software families based on dynamic features and integrated learning of the present invention.
[0055] Figure 7 is an architecture diagram of the system for classifying malicious software families based on dynamic features and integrated learning of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0056] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments:
[0057] Embodiment 1
[0058] As Figure 1 shown, a malware family classification method based on dynamic features and ensemble learning includes the following steps:
[0059] S1: Extract the dynamic application programming interface (API) call sequence of each malicious code in the training malware, and summarize the dynamic API call sequences as the original features.
[0060] Specifically, dynamic features generally include registry information, dynamic API and system call information, dynamic opcode sequences, etc. The inventor found through research that the API calls generated during the code running process can better reflect its true behavior and intention. Therefore, in this embodiment, the dynamic API call sequence, a dynamic feature, is selected to characterize malicious code, and the classification of malicious code, especially unknown family malicious code, is achieved by learning the implicit information contained therein.
[0061] As an implementable manner, existing machine learning methods (such as K-Nearest Neighbor (KNN), Support Vector Machine (SVM), Decision Tree (DT) and Random Forests (RF), Naive Bayesian Model (NBM) and Bayesian Network, etc.) can be used to extract the dynamic API call sequence.
[0062] As another implementable manner, first use sandbox and BSA sandbox analysis tools to analyze the malware samples, and then use a regularization method to extract the API call sequence of each malicious code from the analysis results and save it to a json file. The extraction process is as Figure 2 shown and includes the following steps:
[0063] <1> Start the sandbox.
[0064] <2> Start the sandbox analysis tool.
[0065] <3> Run the malware in the sandbox.
[0066] <4> Determine whether the analysis of the malware by the sandbox analysis tool is completed. If not, return to <2>. If completed, proceed to the next step.
[0067] <5> Obtain the analysis results.
[0068] <6>Close the sandbox.
[0069] By running malicious code using the Sandbox and analyzing the running process of the malicious code with the BSA sandbox analysis tool, all API log files of the dynamic execution process of the malicious code can be obtained, including API sequences, executed operations, modified files, executed files, etc. In this paper, the API call sequences of the malicious code are selected as the original feature inputs. By writing a Python script, regular matching is performed on the API log files, the corresponding API call names are selected, and the required sequences are formed in the call order. Finally, all API call sequences are aggregated into a JSON file.
[0070] S2: Construct a bidirectional long short-term memory neural network Bi-LSTM model as the first base classifier.
[0071] S3: Construct a text convolutional neural network TextCNN model as the second base classifier.
[0072] S4: Construct a convolutional neural network CNN + long short-term memory neural network LSTM model as the third base classifier.
[0073] S5: Respectively use the original features as the inputs of the three base classifiers, and iteratively train the three base classifiers through the extreme gradient boosting algorithm. After training is completed, an ensemble learning classification model is obtained as the malware family classification model.
[0074] S6: Classify the malware family of all malicious codes in the target malware through the malware family classification model.
[0075] The malware family classification method based on dynamic features and ensemble learning provided by the embodiments of the present invention trains multiple base learners and integrates them according to the extreme gradient boosting algorithm, and finally obtains an ensemble learning model instead of a single model classification method, which can comprehensively learn the internal information of the features, so the classification accuracy can be significantly improved.
[0076] Embodiment 2
[0077] Based on the above embodiments, the embodiments of the present invention provide a construction process of the first base learner, which is specifically as follows:
[0078] The structure of the first base classifier constructed by the embodiments of the present invention is as Figure 3 shown, including an input layer, an embedding layer, a Bi-LSTM layer, an attention layer, a fully connected layer, and an output layer.
[0079] Specifically, traditional RNNs can theoretically solve the context correlation problem. However, since the neurons in the model only retain the output of the adjacent level during specific calculations, the long-term dependence problem in the input data cannot be well represented and learned. Therefore, in this embodiment, LSTM units with long-term memory are selected as the hidden layer units of the recurrent neural network. To make full use of context information rather than simply associating with the previous context, a bidirectional LSTM network is further selected to learn features, such as Figure 3 the Bi-LSTM layer in
[0080] Specifically, Bi-LSTM is an improvement of RNN, which is mainly composed of a forward LSTM and a backward LSTM combined. Each time step contains an LSTM unit to selectively remember, forget, and output information. Suppose a malware API sequence data X = [x 1 , x 2 ,..., x T with T elements is given, where i = 1, 2,..., T, and x i is the i-th element in the malware API sequence data. Each x i is transformed into a D-dimensional real vector e i through the embedding layer, resulting in a vector matrix E composed of the D-dimensional embeddings of all elements in X, E = [e 1 , e 2 ,..., e T , and e i is the i-th element in the sequence. After this T×D-dimensional vector matrix E is input into the Bi-LSTM model, there are bidirectional hidden state outputs at time step t. Among them, the hidden state output of the forward LSTM is denoted as and the hidden state output of the backward LSTM is denoted as
[0081]
[0082] where is the hidden state output of the forward LSTM at time step t - 1, is the hidden state output of the backward LSTM at time step t + 1, and e i is a D-dimensional real vector transformed by the embedding layer.
[0083] The hidden state output of Bi-LSTM at time step t is composed of the combination and splicing of the bidirectional hidden states, denoted as Suppose the number of nodes in the LSTM hidden layer is n here. After the sequence training is completed, a T×2n-dimensional hidden state set H = [h 1 , h 2 ,…, hT , i = 1, 2, …, T, h i is the hidden state output at time i.
[0084] Furthermore, when using Bi-LSTM for data learning, a large amount of data often needs to be faced. However, at different times, in fact, only a small number of individual data are useful and need to be noted. Therefore, in order to better learn effective data, the embodiment of the present invention further introduces an attention mechanism into the above Bi-LSTM model.
[0085] Specifically, an attention layer is introduced into the first base classifier. Taking the hidden state set H of Bi-LSTM as the input, the attention mechanism is used to calculate the important parts among the T hidden states, that is, an attention weight is assigned to each of the T hidden states. First, we perform a non-linear transformation using the activation function Tanh to complete the dimension transformation; then we use the softmax function to normalize it to obtain the output attention vector α. This vector is the attention distribution of the query vector in the attention mechanism on H, where each item represents a probability corresponding to an element of the original input, and the sum of all items is 1. The weight matrix obtained by the Attention layer is as follows.
[0086] M = tanh(H) (3)
[0087] α = softmax(w T M) (4)
[0088] r = Hα T (5)
[0089] h * = tanh(r) (6)
[0090] Among them, H is the hidden state set, M is the result of the hidden state set H after being transformed by the tanh function, T corresponds to the input length of H ∈ R D×T , D is the dimension of the word vector, α is the output attention vector, α T is the transpose of α, w T is the transpose of a parameter vector obtained by training and learning, r is the vector representation of the API sequence, with dimension D, and h * is the representation of the API sequence finally used for classification, with dimension D.
[0091] When h * is sent into the fully connected (Dense) layer as the output of the attention layer, a Dropout process is also designed. First, overfitting is alleviated for it, and then it is sent into the fully connected layer. Finally, a softmax classifier is used to detect the detected samples and output the detection results.
[0092] Example 3
[0093] Based on the above embodiments, an embodiment of the present invention provides a construction process of a second base classifier.
[0094] The currently most commonly used deep learning models mainly focus on two categories: RNN and CNN. These two types of models are respectively good at different scenarios in the application of the natural language processing field. When long-term dependencies need to be processed, such as understanding through long sentences, RNN often performs better than CNN; on the contrary, if local features or key phrases need to be extracted and processed, CNN performs better than RNN.
[0095] Since the CNN model is usually more suitable for processing image problems, the second base classifier selects TextCNN specifically for text classification. TextCNN is a neural network model based on convolutional neural network designed specifically for text classification problems. It can more efficiently self-learn and extract local features in the text through convolutional and pooling operations, and map the text representation to the classification label through a fully connected layer. Compared with traditional text classification methods, TextCNN can process texts of different lengths and shows good performance in some text classification tasks.
[0096] The structure of the second base classifier is as Figure 4 shown, mainly including an input layer, a convolutional layer, a pooling layer, and a fully connected layer. Among them, the task of the input layer is to convert the API sequence into a corresponding vector representation through a word embedding mechanism. In order to fully abstract the local features of different views in the sequence, the convolutional layer of the original TextCNN is improved. Convolution kernels with window sizes of 2, 3, 4, and 5 are respectively selected to perform convolution processing on the vector representation after word embedding of the API call sequence, and automatically extract feature vectors; in the pooling layer, max pooling operation is adopted to process the different convolution outputs of the convolutional layer respectively, further extract and reduce the dimension of them, and cascade the outputs of the pooling layer to splice and obtain a 256-dimensional one-dimensional feature vector as the final output of the pooling layer. In order to convert the output vector of the pooling layer into the required prediction result, a softmax layer is added for flattening, and finally the predicted classification label is output. This model only contains one layer of convolution and one layer of pooling, with a simple structure, which helps to improve the training speed and classification performance.
[0097] Specifically, different from graph convolution, graph convolution is a two-dimensional convolution kernel that slides from left to right and from top to bottom. However, for the API sequence, it is meaningless to perform convolution operations on multiple APIs from left to right. Therefore, one-dimensional convolution is adopted here to extract the features in the sequence. The sequence is uniformly represented as a k-dimensional word vector, and the sequence of length n is represented according to the following formula:
[0098] x 1:n = [x 1 , x 2 , …, x n (7)
[0099] where x 1:n is a sequence of length n, and x i ∈ R k represents the i-th API in the sequence.
[0100] The convolutional kernel w ∈ R hk is used to perform the convolution operation. In this paper, 4 convolutions are designed, and 2, 3, 4, and 5 are taken as the window values respectively. The general representation of the features is as follows:
[0101] c i = f(w · x i:i+h-1 + b) (8)
[0102] where x i:i+h-1 is the convolution window, h is the size of the convolution window, c i represents the feature generated by the convolution window x i:i+h-1 , b is the bias term, f represents the non-linear function, and w is the convolutional kernel and w ∈ R hk .
[0103] The convolutional kernel w with window size h performs convolution processing from top to bottom on the sequence vector generated by the word embedding. The width of the convolution is the same as the dimension of the word vector. Such a convolution operation can take into account both semantics and word order.
[0104] The convolution process can be described by the following formula:
[0105] {X 1:h , X 2:h+1 , X 3:h+2 ,..., X h+1:n} (9)
[0106] where X h+1:n is the convolution window.
[0107] After such a convolution operation, a feature vector of the following form can be obtained:
[0108] c = [c 1 , c 2 , c 3 ,..., c n-h+1 (10)
[0109] where c is the feature vector after convolution, and c n-h+1 is the (n - h + 1)-th element in the feature vector.
[0110] The pooling layer uses max pooling operation. Through the 1-Max-pooling operation, each feature vector obtained by convolution is pooled into a single value. After pooling, all the pooled values are concatenated to obtain the final output of the pooling layer. To prevent overfitting, dropout with two different ratios is added to sparsify the fully connected layer, and finally the fully connected layer outputs the classification labels.
[0111] Example 4
[0112] Based on the above embodiments, an embodiment of the present invention provides a construction process of a third base classifier.
[0113] Instead of selecting a ready-made model, the third base classifier constructs a malware family classification model based on CNN+LSTM by combining neural network models of CNN and LSTM to achieve family classification of malware samples. According to the task requirements, the model designs multiple layers with different functions, including an embedding layer, a convolutional layer, a pooling layer, an LSTM layer, and a fully connected layer, and uses Softmax for classification. The structure of the third base classifier is as Figure 5 shown.
[0114] Among them, the embedding layer performs word embedding processing on the input dynamic API call sequence of malware samples, converting it into a vector representation for learning the family features in the sample file; the convolutional layer and max pooling are used to extract the spatial features contained in the dynamic API sequence. Two sets of such processes are designed for this base classifier, using convolutional kernels of different sizes to better capture features of different sizes and achieve dimensionality reduction of the extracted spatial features; the LSTM layer is used to capture the long-term dependence information in the dynamic API call sequence to learn the temporal features and semantic information contained in the sample; the fully connected layer maps the output of the LSTM layer to the target category (i.e., malware family) and uses Softmax for classification.
[0115] Here, word embedding is selected instead of the one-hot encoding commonly used in existing methods because in one-hot encoding, the status of all words is equal, which will cause the dependency relationship between words to be ignored, and this is very important information in the dynamically collected API call sequence. Therefore, choosing one-hot encoding will cause feature loss in the preprocessing stage and will inevitably affect the classification result; on the other hand, the encoding mechanism of one-hot encoding will cause the formed input vector to have a very large dimension, and too many 0 values make the features too sparse and redundant, thus affecting the final classification effect.
[0116] The convolutional layer in the base classifier selects one-dimensional convolution because the features here are sequences rather than images, and one-dimensional convolution is a convolution operation specifically designed for processing sequence data. Different from two-dimensional convolution, one-dimensional convolution performs convolution operations in only one direction and is usually used to extract local features in sequences. Here, the order relationship of behaviors reflected by the API sequence is extracted. According to the size of the convolution kernel, the range of consecutive operations to be concerned can be adjusted. The recurrent neural network layer selects the LSTM layer, which can effectively maintain long-term dependence information and capture semantic relationships in the sequence. Here, the long-term dependence call relationship of the API sequence is correspondingly extracted, which can reflect the behavioral association logic of the samples to a certain extent.
[0117] The pooling layer selects max pooling because max pooling retains the most significant original feature information while reducing the dimension, which can make the model have better robustness and can adapt to changes such as small offsets and noises in the input data.
[0118] The fully connected layer linearly combines the outputs of the previous layers and performs a non-linear transformation through an activation function. Its principle is to expand the input data into a one-dimensional vector, perform a linear transformation through a weight matrix, and then use the activation function to process to obtain the final output. The activation function of this base classifier selects Softmax because it is more suitable for multi-class classification tasks and meets the research needs of malware family classification. The principle of Softmax is to first calculate the scores of each class (i.e., the original outputs in exponential form), and then normalize all the scores to obtain the probability estimates of each class. Specifically, for a multi-class classification problem with n classes, Softmax outputs an n-dimensional vector, and each element represents the probability estimate of the corresponding class, and the sum of all elements is 1. Its specific definition is shown by the following formula:
[0119]
[0120] where x i is the output value of the i-th node, C is the number of output nodes, that is, the number of classification categories, is the sum of c e xc s. Through the Softmax function, the output values of multi-class classification can be converted into a probability distribution with a range of [0, 1] and a sum of 1.
[0121] Example 5
[0122] Based on the above embodiments, the embodiment of the present invention provides a model integration optimization method based on XGBoost.
[0123] XGBoost is an improved tree boosting machine learning algorithm, which is an improvement based on the Gradient Boosting Decision Tree (GBDT). By using optimization techniques such as sparse matrices and cache blocks, it further enhances the model performance and is an efficient ensemble learning method.
[0124] XGBoost adopts the Gradient Boosting algorithm to improve the performance of the ensemble model. By sequentially building K classification and regression trees, the prediction results are made to be as close as possible to the true values and have as strong a generalization ability as possible. The Gradient Boosting algorithm iteratively trains the base classifiers to minimize the error between the prediction results of each base classifier and the true labels, thereby constructing a new and relatively more powerful ensemble learning classifier through training. Different from the voting method, the ensemble of XGBoost is serial. Each base classifier learns the error value between the result of the previous base classifier and the actual value. In the process of each iterative learning, a new decision tree is added to the ensemble model and weighted averaged to improve the overall prediction performance. Through the learning of multiple models, the error between the model value and the actual value can be continuously reduced, that is, the predicted value is made closer to the actual value, achieving a higher prediction accuracy. Through the training of one base classifier after another, new trees are continuously generated, and the final output is shown in the following formula, that is, the results of all trees are accumulated to obtain the predicted value of a sample.
[0125] The working principle of XGBoost can be expressed by the following formula:
[0126]
[0127]
[0128] Among them, is the model prediction result after the nth round of iterative learning, yi is the actual value, and f k (x i ) is the prediction result of the kth base classifier.
[0129] Each step of the process of selecting and generating trees in the above training process is guided by its specific objective function. The composition of XGBoost includes three parts: the objective function, the regularization term, and the decision tree structure. The objective function is used to define the optimization objective of the model, including the loss function and the regularization term. The decision tree structure can search for the best split point through the greedy algorithm and output a predicted value at each leaf node. XGBoost upgrades the loss function from the first order to the second order Taylor expansion on the basis of GBDT, so it converges faster; a regularization term is introduced into the loss function to limit the complexity of the tree, which also helps to avoid overfitting.
[0130] To obtain a better model in XGBoost ensemble learning, it is necessary to first select an appropriate loss function and optimizer according to the training objective. In this invention, the multi-classification cross-entropy is selected as the loss function, and the softmax is selected as the activation function for use in combination to optimize the family classification model of malware.
[0131] The optimization algorithms for the model mainly include the stochastic gradient descent optimization algorithm, the momentum optimization algorithm, and the adaptive learning rate optimization algorithm. In this paper, the most suitable optimization method is selected according to the training accuracy value to optimize the model.
[0132] Among them, gradient descent is the most basic and commonly used optimization algorithm. Its core idea is to update the parameters in the negative gradient direction in each iteration to minimize the loss function. Specifically, the gradient descent algorithm can be expressed by the following formula:
[0133]
[0134] where L(θ t ) is the current loss function, θ t represents the parameters of the loss function at the t-th iteration, is the gradient of the loss function, and α is the learning rate.
[0135] The momentum optimization algorithm (Momentum) is an improvement based on the gradient descent algorithm. Its main idea is to introduce the influence of historical gradients when updating parameters to accelerate convergence and improve the generalization performance of the model. Specifically, the momentum optimization algorithm can be expressed by the following formula:
[0136]
[0137] where is the gradient of the loss function, α is the learning rate, β is the momentum decay coefficient, and v t is the momentum at the t-th iteration.
[0138] The adaptive learning rate optimization algorithm is also an optimization algorithm that adaptively adjusts the learning rate according to gradient information. Common ones include AdaGrad, RMSProp, and Adam, etc. These algorithms adaptively adjust the learning rate in different ways to better adapt to the characteristics of different parameters and different stages of the training process. In this paper, three optimization algorithms, RMSprop, Adam, and SGD, are selected for training respectively to select the most suitable optimization algorithm to optimize the ensemble model, and the learning curve is as Figure 6 shown.
[0139] Through comparative experiments, the RMSProp optimizer, which is the most stable in the corresponding dataset samples of this paper, is finally selected. From Figure 6From the training results, it can be seen that the RMSProp optimizer has higher accuracy and better stability compared to the other two optimizers. Therefore, in the optimization of the integrated model, the RMSProp optimization algorithm is selected to optimize the integrated learning classification model.
[0140] Example 6
[0141] Corresponding to the above method, as Figure 7 shown, an embodiment of the present invention provides a malware family classification system based on dynamic features and integrated learning, including:
[0142] A feature extraction module, configured to extract the dynamic application programming interface (API) call sequence of each malicious code in the training malware, and summarize the dynamic API call sequences as original features.
[0143] The first base classifier module is configured to construct a bidirectional long short-term memory neural network (Bi-LSTM) model as the first base classifier.
[0144] The second base classifier module is configured to construct a text convolutional neural network (TextCNN) model as the second base classifier.
[0145] The third base classifier module is configured to construct a convolutional neural network (CNN) + long short-term memory neural network (LSTM) model as the third base classifier.
[0146] An integrated learning classifier module is configured to use the original features as the inputs of the three base classifiers respectively, and iteratively train the three base classifiers through the extreme gradient boosting algorithm. After the training is completed, an integrated learning classification model is obtained as the malware family classification model.
[0147] A classification module is configured to classify the malware family of all malicious codes in the target malware through the malware family classification model.
[0148] It should be noted that the malware family classification system based on dynamic features and integrated learning provided by the embodiment of the present invention is to implement the above-mentioned malware family classification method based on dynamic features and integrated learning. Its functions can be specifically referred to the above method embodiments, and will not be elaborated here.
[0149] In summary, the present invention extracts feature data through a dynamic analysis method, avoiding the influence of technologies such as code obfuscation on feature extraction. By setting three base classifiers, namely the Bi-LSTM model, the TextCNN model, and the CNN+LSTM model, to comprehensively learn the malicious software API sequence features extracted dynamically. The three base classifiers are iteratively trained through the extreme gradient boosting algorithm, and the three base classifiers are merged to complete the construction of the integrated learning classification model. The malicious software families are classified through integrated learning, improving the accuracy of classification.
Claims
1. Malware family classification method based on dynamic features and ensemble learning, characterized by: The following steps are involved: Step 1: Extract the dynamic application programming interface API call sequence of each malicious code in the training malware, and summarize the dynamic API call sequence as the original feature; Step 2: Build a bidirectional long short-term memory neural network Bi-LSTM model as the first base classifier; Step 3: Build a text convolutional neural network TextCNN model as the second base classifier; Step 4: Build a convolutional neural network CNN + long short-term memory neural network LSTM model as the third base classifier; Step 5: Use the original features as the input of the three base classifiers respectively, and iteratively train the three base classifiers through the extreme gradient boosting algorithm. After the training is completed, an integrated learning classification model is obtained as the malware family classification model; Step 6: Classify all malicious codes in the target malware into malware families using the malware family classification model; In the step 2, an attention layer is introduced into the Bi-LSTM model, and a neural network model of Bi-LSTM and attention mechanism is constructed as the first base classifier; In the step 3, the convolution layer in the TextCNN model includes four convolution kernels of different sizes, and a normalized exponential function softmax layer is added after the fully connected layer; In the step 4, the third base classifier includes an embedding layer, two groups of convolution blocks, a long short-term memory neural network LSTM layer and a fully connected layer in sequence, wherein each group of convolution blocks includes a convolution layer and a pooling layer, the convolution layer selects one-dimensional convolution, the pooling layer adopts a maximum pooling operation, and a normalized exponential function softmax layer is added after the fully connected layer for classification.
2. The malware family classification method based on dynamic features and ensemble learning according to claim 1 is characterized in that: In the step 5, during training, multi-classification cross entropy is selected as the loss function.
3. The malware family classification method based on dynamic features and ensemble learning according to claim 1 is characterized in that: In the step five, the ensemble learning classification model is optimized by an optimization algorithm, and the optimization algorithm adopts any one of a stochastic gradient descent optimization algorithm, a momentum optimization algorithm, and an adaptive learning rate optimization algorithm.
4. Malware family classification system based on dynamic features and ensemble learning, characterized by: include: A feature extraction module is used to extract a dynamic application programming interface (API) call sequence of each malicious code in the malware, and summarize the dynamic API call sequence as an original feature; The first base classifier module is used to build a bidirectional long short-term memory neural network Bi-LSTM model as the first base classifier; The second base classifier module is used to build a text convolutional neural network TextCNN model as the second base classifier; The third base classifier module is used to build a convolutional neural network CNN + long short-term memory neural network LSTM model as the third base classifier; The ensemble learning classifier module is used to use the original features as the input of the three base classifiers respectively, and iteratively train the three base classifiers through the extreme gradient boosting algorithm. After the training is completed, an ensemble learning classification model is obtained as the malware family classification model; A classification module, configured to classify all malicious codes in the target malware into malware families using the malware family classification model; In the first base classifier module, an attention layer is introduced into the Bi-LSTM model, and a neural network model of Bi-LSTM and attention mechanism is constructed as the first base classifier; In the second base classifier module, the convolution layer in the TextCNN model includes four convolution kernels of different sizes, and a normalized exponential function softmax layer is added after the fully connected layer; In the third base classifier module, the third base classifier includes an embedding layer, two groups of convolution blocks, an LSTM layer and a fully connected layer in sequence, wherein each group of convolution blocks includes a convolution layer and a pooling layer, the convolution layer selects one-dimensional convolution, the pooling layer adopts the maximum pooling operation, and a normalized exponential function softmax layer is added after the fully connected layer for classification.
Citation Information
Patent Citations
Malicious code dynamic behavior knowledge graph construction method and system and storage medium
CN114707137A
Malicious code detection method based on static characteristics and storage medium
CN115859290A