A malicious code recognition method, device, storage medium, and terminal device

By converting the PE file dataset into grayscale images and extracting features based on image texture features, a multi-layer neural network model is established, which solves the problem of poor malicious code recognition in the existing technology, and achieves efficient and accurate malicious code recognition.

CN116127449BActive Publication Date: 2025-05-30XIAN UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211090115.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-07
Publication Date
2025-05-30
Estimated Expiration
2042-09-07

AI Technical Summary

Technical Problem

The prior art has poor recognition effect when identifying unknown malicious codes, with high false alarm rates and missed alarm rates, and it is difficult to achieve real-time analysis and accurate identification.

Method used

By loading the PE file dataset, it is divided into a training set and a verification set, the data is converted into a grayscale image set by using file pixelation, and feature extraction is performed based on image texture features, a multi-layer neural network model is established for training, and finally a malicious code training model is obtained for detecting PE file types.

Benefits of technology

It improves the recognition efficiency of malicious code, reduces the false alarm rate and missed alarm rate, realizes accurate identification of unknown malicious code, and is suitable for real-time analysis in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116127449B_ABST
    Figure CN116127449B_ABST
Patent Text Reader

Abstract

The present invention relates to malicious code recognition, and specifically relates to a method and device for malicious code recognition, a storage medium, and a terminal device. The purpose is to solve the technical problem that the prior art has poor effect in recognizing unknown malicious codes. The present invention provides a method and device for malicious code recognition, a storage medium, and a terminal device. The method includes the following steps: loading a PE file dataset and converting it into a grayscale image set; performing data enhancement processing on the grayscale image set by using the Gamma algorithm to obtain a grayscale enhanced image set; extracting features from the obtained grayscale enhanced image set based on image texture features; establishing a malicious code model, and using the extracted image texture features and the grayscale enhanced image set as inputs to train the model to obtain a malicious code training model; loading the PE file to be recognized; using the above malicious code training model for detection; and outputting the file type of the PE file to be recognized according to the detection result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to malicious code recognition, and in particular to a method and device for malicious code recognition, a storage medium, and a terminal device. Background Art

[0002] With the development of information technology, the attack frequency of malicious code has increased geometrically, and malicious code variants have become an important threat to information security. Among them, malicious codes such as Trojans, viruses, and worms represented by the PE (Portable Executable) file type have the most extensive impact. In the prior art, static detection or dynamic monitoring technology is usually used to detect malicious code. However, traditional static detection technology has insufficient detection capabilities for unknown codes and malicious code variants. Although dynamic detection can effectively detect unknown codes and malicious code variants and has good generalization capabilities, its disadvantages are high false positive rates and high false negative rates. In addition, machine learning is also commonly used for malicious code recognition, usually using random forest algorithms or support vector algorithms. Although the above machine learning analysis methods can distinguish between benign files and malicious code files, and detect known or unknown malicious codes and their variants, the detection accuracy of these methods is low, the real-time performance is poor, and the detection effect depends on the experience of analysts. Moreover, in the face of an increasingly large and rapidly growing malicious code library, neither static detection nor dynamic monitoring can meet the needs of real-time analysis, and it is difficult to achieve precise identification of unknown malicious codes. Summary of the Invention

[0003] The object of the present invention is to solve the technical problem of poor recognition effect of the prior art in identifying unknown malicious codes, and to provide a method and device for malicious code recognition, a storage medium, and a terminal device.

[0004] To solve the above technical problem, the technical solution provided by the present invention is as follows:

[0005] A method for malicious code recognition, characterized in that it includes the following steps:

[0006] Step 1, load the PE file dataset and divide it into a training set and a validation set according to a certain ratio, wherein the PE file dataset includes benign files and malicious code files;

[0007] Step 2, respectively convert the training set and the validation set into a first grayscale image set and a second grayscale image set through file pixelization;

[0008] Alternatively, the training set and the validation set are respectively converted into a first grayscale image set and a second grayscale image set through file pixelization, and the obtained first grayscale image set and second grayscale image set are subjected to data augmentation processing through the Gamma algorithm to obtain a first grayscale enhanced image set and a second grayscale enhanced image set;

[0009] Step 3: Feature extraction is performed on the obtained first grayscale image set or first grayscale enhanced image set, and the obtained second grayscale image set or second grayscale enhanced image set based on image texture features to obtain a first feature vector and a second feature vector;

[0010] Step 4: Establish a malicious code model, where the malicious code model includes a multi-layer neural network model;

[0011] Step 5: The malicious code model is trained by the first feature vector, the first grayscale image set or the first grayscale enhanced image set, and the second feature vector, the second grayscale image set or the second grayscale enhanced image set to obtain a malicious code training model;

[0012] Step 6: Load the PE file to be recognized, and convert the PE file to be recognized into a grayscale image to be recognized through file pixelization;

[0013] Step 7: Detect the grayscale image to be recognized according to the obtained malicious code training model to obtain a detection result;

[0014] Step 8: Output the file type of the PE file to be recognized according to the obtained detection result. If the PE file to be recognized is a benign file, output the benign file class number. If the PE file to be recognized is a malicious code file, output the malicious code file class number.

[0015] Further, the multi-layer neural network model includes:

[0016] Input layer: Input the first feature vector, and the first grayscale image set or the first grayscale enhanced image set;

[0017] Input the second feature vector, and the second grayscale image set or the second grayscale enhanced image set;

[0018] Convolution layer: Includes a first convolution layer, a second convolution layer, and a third convolution layer, all using a 3×3 convolution kernel;

[0019] Pooling layer: Includes a first pooling layer, a second pooling layer, and a third pooling layer, all using the maximum pooling method, and the pooling width is set to 3×3;

[0020] Dropout layer: Includes a first Dropout layer and a second Dropout layer;

[0021] Fully connected layer: including a first fully connected layer, a second fully connected layer, and a third fully connected layer, all of which adopt a Sigmoid classifier;

[0022] Output layer: output the file type;

[0023] Among them, the input layer, the first convolutional layer, the first pooling layer, the second convolutional layer, the second pooling layer, the third convolutional layer, the third pooling layer, the first Dropout layer, the first fully connected layer, the second Dropout layer, the second fully connected layer, the third fully connected layer, and the output layer are connected in sequence for data.

[0024] Furthermore, in step 2, the training set and the validation set are respectively converted into a first grayscale image set and a second grayscale image set through file pixelization, including the following steps:

[0025] Step 2.1, convert the byte sequence of each PE file in the training set into a first vector group, and convert each first vector group into a corresponding first grayscale value group;

[0026] Convert the byte sequence of each PE file in the validation set into a second vector group, and convert each second vector group into a corresponding second grayscale value group;

[0027] Step 2.2, generate a first grayscale image set corresponding one-to-one to the training set according to the obtained first grayscale value group;

[0028] Generate a second grayscale image set corresponding one-to-one to the validation set according to the obtained second grayscale value group.

[0029] Furthermore, in step 3, if there are PE files of different sizes in the training set or the validation set, then perform a stepped division on the file size range in the training set or the validation set, divide the PE files of different sizes into different categories according to the range where their file sizes are located, and perform corresponding image width constraints on the PE files of different categories.

[0030] Furthermore, in step 3, the SURF algorithm is selected for feature extraction.

[0031] Meanwhile, the present invention provides a malicious code recognition device for implementing the above recognition method, which is characterized in that: it includes a training unit, and a loading unit, an identification unit, and an output unit connected in sequence;

[0032] The training unit includes an acquisition module, an image module, a feature extraction module, and a model training module connected in sequence:

[0033] The acquisition module is used to acquire a PE file data set and divide it into a training set and a validation set according to a certain ratio;

[0034] The image module is used to pixelize the training set and the validation set through files respectively to convert them into a first grayscale image set and a second grayscale image set;

[0035] Alternatively, the training set and the validation set are pixelized through files respectively to convert them into a first grayscale image set and a second grayscale image set, and the obtained first grayscale image set and second grayscale image set are subjected to data augmentation processing through the Gamma algorithm to obtain a first grayscale augmented image set and a second grayscale augmented image set;

[0036] The feature extraction module is used to extract features from the obtained first grayscale image set or first grayscale augmented image set, and the obtained second grayscale image set or second grayscale augmented image set based on image texture features to obtain a first feature vector and a second feature vector;

[0037] The model training module is used to establish a malicious code model, and train the malicious code model through the first feature vector, the first grayscale image set or the first grayscale augmented image set, and the second feature vector, the second grayscale image set or the second grayscale augmented image set to obtain a malicious code training model;

[0038] The loading unit is used to load the PE file to be recognized, and pixelize the PE file to be recognized through a file to convert it into a grayscale image to be recognized;

[0039] The recognition unit is used to detect the grayscale image to be recognized according to the malicious code training model to obtain a detection result;

[0040] The output unit is used to output the file type of the PE file to be recognized according to the obtained detection result. If the PE file to be recognized is a benign file, the benign file class number is output. If the PE file to be recognized is a malicious code file, the malicious code file class number is output.

[0041] The present invention provides a computer-readable storage medium, on which a computer program is stored. The special feature is that when the program is executed by a processor, the steps of the above recognition method are implemented.

[0042] The present invention provides a terminal device, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. The special feature is that when the processor executes the computer program, the steps of the above recognition method are implemented.

[0043] The beneficial effects of the present invention compared with the prior art:

[0044] 1. A malicious code recognition method provided by the present invention divides the loaded PE file dataset into a training set and a validation set. Using the data processing method of file pixelization, the training set and the validation set are respectively transformed into a first grayscale image set and a second grayscale image set through file pixelization. Feature extraction is performed on the obtained first grayscale image set and second grayscale image set based on image texture features. Then, the first grayscale image set, the second grayscale image set, and the extracted image texture features are input into the established malicious code model. The present invention performs feature learning and training tests on the malicious code model through the first grayscale image set, the second grayscale image set, and the obtained image texture features, and finally obtains a malicious code training model that can be used to detect PE file types. This model can overall improve the recognition efficiency of malicious codes and solve the technical problems of low detection effect and incomplete feature extraction in the prior art.

[0045] 2. A malicious code recognition method provided by the present invention converts the byte sequence in each PE file into image pixel values that can represent its features, so that the malicious code recognition problem is transformed into a binary classification problem of images, and then image technology can be applied to analyze PE files. The advantage of this method is that it can ignore the complex formatting analysis of PE files and simplify the process.

[0046] 3. A malicious code recognition method provided by the present invention uses the pixel values of images as research features. On the one hand, this measure improves the singularity problem caused by using APIs, opcodes, system calls, etc. in executable files as features in the prior art, and also avoids the harmful consequences brought by directly running malicious codes. On the other hand, through the combination of PE file visualization processing and image texture feature extraction, the classification effect is improved.

[0047] 4. A malicious code recognition method provided by the present invention, if there are PE files of different sizes in the obtained first grayscale image set, the file size range can be divided to perform corresponding width constraints on the files in different ranges, so that PE files of different sizes can retain as many pixel points as possible, ensuring the balance of file sizes.

[0048] 5. A malicious code recognition method provided by the present invention can also perform data enhancement processing on the first grayscale image set and the second grayscale image set through the Gamma algorithm before feature extraction to avoid overfitting phenomena during the model training process.

[0049] 6. A malicious code recognition method provided by the present invention, in order to ensure that the distance between different species of points in the lowest-dimensional feature space is greater than the distance between the same species of points, feature extraction algorithms based on texture features such as SITF, SURF, HOG, and LBP can be selected to reduce the dimensions in the complex feature space, thereby realizing the dimension compression of the complex feature space.

[0050] 7. A malicious code recognition method provided by the present invention. The established multi-layer neural network model includes three convolutional layers, three pooling layers, two Dropout layers and three fully connected layers. And in the embodiments of the present invention, preferred embodiments are given, which can achieve a better classification effect and improve the ability to capture important features of malicious code in PE files to a certain extent. The malicious code model includes a multi-layer neural network model, which has been verified by practice and can achieve the expected accuracy and stability. It is not only applicable to the fast and accurate recognition of malicious code of PE file type in complex scenarios, but also applicable to the detection of unknown malicious code.

[0051] 8. A malicious code recognition method provided by the present invention. Since the smallest space that can capture eight-neighborhood information in pixels is 3×3, in this embodiment, multiple convolutional layers all use 3×3 convolutional kernels. Moreover, multiple convolutional layers using 3×3 convolutional kernels have more non-linear relationships compared with a convolutional layer using a large-size convolutional kernel, making the connections between the extracted image texture features closer.

[0052] 9. A malicious code recognition method provided by the present invention. Since it is necessary to retain local texture features, the present invention adopts the max pooling method and sets the pooling width to 3×3 to reduce the dimension, reduce the calculation amount, prevent overfitting, and at the same time reduce the image scale and improve the calculation speed.

[0053] 10. A malicious code recognition method provided by the present invention. Using the malicious code training model as the recognition basis, the recognition effect has a high accuracy. Those skilled in the art can make adaptive adjustments to the quantity, scale and training complexity of the malicious code model according to actual recognition needs before training. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 is a flowchart of an embodiment of a malicious code recognition method provided by the present invention;

[0055] Figure 2 is a schematic diagram of the process of file pixelization in an embodiment of a malicious code recognition method provided by the present invention;

[0056] Figure 3 is a set of grayscale images generated in an embodiment of a malicious code recognition method provided by the present invention, where Figure 3 (a) in Figure 3 and (b) in Figure 3 are both grayscale images generated from malicious code files, Figure 3 and (c) in

[0057] Figure 4 and (d) in Figure 3 are both grayscale images generated from benign files;Schematic diagram of a multi-layer neural network model in an embodiment of a malicious code recognition method provided by the present invention;

[0058] Explanation of reference numerals:

[0059] I - Input layer, C1 - First convolutional layer, C2 - Second convolutional layer, C3 - Third convolutional layer, MP1 - First pooling layer, MP2 - Second pooling layer, MP3 - Third pooling layer, D1 - First Dropout layer, D2 - Second Dropout layer, F1 - First fully connected layer, F2 - Second fully connected layer, F3 - Third fully connected layer, O - Output layer. Detailed implementation manners

[0060] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments.

[0061] Refer to Figure 1 , the present invention provides a malicious code recognition method, including the following steps:

[0062] Step 1, load the PE file dataset, and divide the PE file dataset into a training set and a validation set according to a certain ratio, where the PE file dataset includes benign files and malicious code files;

[0063] In this embodiment, the PE file dataset is the model sample dataset. Among them, the benign files are executable files obtained by traversing under a pure Windows system using web crawler means, and the malicious code files come from Trojans, viruses or other malicious codes collected in typical application environments. In the PE file dataset of this embodiment, the ratio of benign files to malicious code files is 1:1, and the benign files and malicious code files in the PE file dataset are randomly allocated according to a ratio of 8:2 to generate a training set and a validation set. Of course, in other embodiments, the ratio of the training set to the validation set can be set according to requirements.

[0064] Step 2, respectively convert the training set and the validation set into a first grayscale image set and a second grayscale image set through file pixelation, and perform data augmentation processing on the obtained first grayscale image set and second grayscale image set through the Gamma algorithm to obtain a first grayscale enhanced image set and a second grayscale enhanced image set;

[0065] Refer to Figure 2 , in this embodiment, the method of file pixelation is used to convert the PE file dataset into a grayscale image set, that is, the training set and the validation set in the PE file dataset are respectively converted into the first grayscale image set and the second grayscale image set in the grayscale image set, specifically including the following steps:

[0066] Step 2.1: Convert the byte sequence of each PE file in the training set into a first vector group, and then convert each first vector group into a corresponding first grayscale value group;

[0067] Convert the byte sequence of each PE file in the validation set into a second vector group, and then convert each second vector group into a corresponding second grayscale value group;

[0068] Step 2.2: Generate a first grayscale image set corresponding one-to-one to the training set according to the obtained first grayscale value group;

[0069] Generate a second grayscale image set corresponding one-to-one to the validation set according to the obtained second grayscale value group.

[0070] PE files are usually in bytecode format. In this embodiment, the byte sequence of each PE file is converted into a vector group through system commands, and then the vector group is converted into a grayscale value group, and a corresponding grayscale image is generated according to the grayscale value group. As Figure 3 shown, (a) in Figure (3) is a grayscale image generated from a file of the VB.AT malicious code family, Figure 3 and (b) is a grayscale image generated from a file of the Dontovo.A malicious code family, Figure 3 both (c) and (d) are grayscale images generated from startup files under the Windows system, that is, grayscale images generated from benign files. It can generally be observed by the naked eye that there are significant image differences between the grayscale images generated from malicious code files and those generated from benign files, and this kind of image difference has a high discrimination degree in the neural network.

[0071] Moreover, by using the Gamma algorithm to perform data augmentation processing on the first grayscale image set and the second grayscale image set, the brightness contrast of the grayscale images can be increased, the dark details can be better displayed, the requirements for extracting sensitive features in the dark parts of the grayscale images can be guaranteed, and the overfitting phenomenon during the training process can be effectively avoided.

[0072] Of course, in other embodiments, it is also possible not to perform data augmentation processing on the obtained first grayscale image set. However, compared with the preferred embodiment in this embodiment, the implementation effect is slightly worse.

[0073] Step 3: Extract features from the obtained first grayscale enhanced image set and second grayscale enhanced image set based on image texture features to obtain a first feature vector and a second feature vector;

[0074] In this embodiment, if there are PE files of different sizes in the training set or the validation set, it is difficult to unify the sizes of the grayscale images after file pixelation. In response to this situation, this embodiment will divide the file size range in the training set or the validation set in a stepped manner, classify PE files of different sizes according to the range where their file sizes are located, and perform corresponding image width constraints on PE files in different ranges. The corresponding relationship between the PE file size and the corresponding grayscale image width is shown in Table 1 below:

[0075] File size Width ≤50KB 32 50KB - 100KB 64 100KB - 300KB 128 300KB - 600KB 256 600KB - 1000KB 512 ≥1000KB 1024

[0076] Table 1

[0077] According to the corresponding relationship shown in Table 1 above, this embodiment performs corresponding width constraints on the grayscale images generated after pixelation of PE files of different sizes, so as to retain as many pixel points as possible. Moreover, the grayscale images generated by multiple PE files within the same range have the same width constraint value, and grayscale images with the same pixel value can be formed through batch processing, enabling this embodiment to achieve a better classification effect on the basis of ensuring the balance of the grayscale image sizes. This embodiment selects the SURF algorithm to extract features from the obtained grayscale image set for dimensionality reduction. Of course, in other embodiments, algorithms such as SITF, HOG, and LBP can also be used to extract image texture features from the obtained first grayscale image set and the obtained second grayscale image set to obtain the first feature vector and the second feature vector.

[0078] Step 4, establish a malicious code model. The malicious code model includes a multi-layer neural network model. The multi-layer neural network model includes an input layer I, a first convolutional layer C1, a first pooling layer MP1, a second convolutional layer C2, a second pooling layer MP2, a third convolutional layer C3, a third pooling layer MP3, a first Dropout layer D1, a first fully connected layer F1, a second Dropout layer D2, a second fully connected layer F2, a third fully connected layer F3, and an output layer O, and data connections are made in sequence.

[0079] The parameters of the multi-layer neural network model in this embodiment are shown in Table 2 below:

[0080]

[0081] Table 2

[0082] In the table, Input represents the input layer I of the multi-layer neural network model, Output represents the output layer O, Size is the scale of the current layer of the neural network, Kernel Size is the size of the convolution kernel, Padding indicates padding with 0, MaxPoolSize represents the pooling width setting, Activation indicates the function activation used in this layer, and Dropout represents the probability of neuron inactivation.

[0083] Referring to Figure 4 , the multi-layer neural network model in this embodiment includes:

[0084] Input layer I: Input the first feature vector, and the first grayscale image set or the first grayscale enhanced image set;

[0085] Input the second feature vector, and the second grayscale image set or the second grayscale enhanced image set;

[0086] Convolution layer: It includes the first convolution layer C1, the second convolution layer C2 and the third convolution layer C3, all using a 3×3 convolution kernel;

[0087] In a convolutional neural network, the size of the convolution kernel usually depends on the distribution and discrimination degree of the extracted features. The size of the extracted features should correspond to the convolution kernel. If the size of the extracted features is small, the corresponding convolution kernel should also be small. If the convolution kernel is too large, some local features will be lost. If multiple smaller-sized convolution kernels are used in the convolution layer, on the one hand, it can reduce the network parameters, and on the other hand, it is equivalent to adding more non-linear parameters, which can improve the network expression ability. Since the smallest space that can capture eight-neighborhood information in pixels is 3×3, the multiple convolution layers in this embodiment all use convolution kernels with a size of 3×3. And multiple convolution layers with 3×3 convolution kernels have more non-linear relationships than a convolution layer with a large-sized convolution kernel, making the connection between image features closer. The convolution operation of grayscale images can be completed using a convolution kernel, which is used to convert linear operations into non-linear operations. The calculation formula is shown as follows:

[0088]

[0089] Among them, f represents the non-linear function, i represents the number of input data of the convolution layer, M represents the number of neurons input to the convolution layer, w ij represents the weight between the i-th input data and the j-th neuron in the convolution layer, x i represents the i-th input data, b j represents the deviation value of the j-th neuron, and y represents the output data of the convolution layer.

[0090] The convolutional layer in this embodiment uses a non - linear function, which can implement complex analog operations. By reading the movement of the preset window of the image, the first grayscale image set or the first grayscale enhanced image set, and the image features in the second grayscale image set or the second grayscale enhanced image set are obtained.

[0091] Pooling layer: It includes the first pooling layer MP1, the second pooling layer MP2 and the third pooling layer MP3, all of which adopt the max - pooling method, and the pooling width is set to 3×3.

[0092] The multiple pooling layers in this embodiment are used to take the feature extraction results obtained by the convolutional layer as input, and further extract features from the feature extraction results to obtain deeper - level features. Since max - pooling is generally used to retain texture features, and this embodiment needs to retain more local features, the width of max - pooling is set to 3×3. The purpose is to have fewer dimensions, reduce the computational workload, and make it not easy to over - fit. It can also reduce the scale of the image, thereby improving the calculation speed.

[0093] Dropout layer: It includes the first Dropout layer D1 and the second Dropout layer D2.

[0094] In this embodiment, through the activation function adopted in each pooling layer and the two Dropout layers, the occurrence of over - fitting phenomena in the connection process can be reduced. The activation functions in this embodiment all adopt the cross - entropy loss function, which can make the predicted value approach the true value as much as possible and optimize the model. The function formula is shown as follows:

[0095] Loss=-∑ i q i lgα i

[0096] Among them, i represents the number of the PE file type to be recognized. For example, i = 1 means the PE file to be recognized is a benign file, and i = 0 means the PE file to be recognized is a malicious code file. α i is the confidence that the model predicts the PE file type to be recognized as the number i, and q i represents the PE file with the true type of the number i, and Loss represents the degree of difference between the predicted type and the actual type.

[0097] The activation function in this embodiment adopts the Rectified Linear Units (ReLU for short). The operation method of ReLU is more simplified, which is beneficial to the training and convergence of the neural network. Its mathematical expression is:

[0098]

[0099] Among them, the function returns the maximum value between 0 and the given number x. When the value is greater than or equal to 0, it returns the maximum value, that is, the value passed when it is greater than or equal to 0. The value of x in this embodiment is set to 1.

[0100] During the process of model training, if the number of hidden layers is large, it will lead to an increase in the amount of calculation, and then overfitting will occur. Therefore, in order to avoid the occurrence of overfitting, in this embodiment, two Dropout layers are used to randomly discard the number of neurons, weaken the joint adaptability between neurons, and improve the generalization ability of the network. Among them, the mathematical expression of the Dropout function is:

[0101] v ~ Bernoulli(q)

[0102] λ'(m) = v × λ(m)

[0103] Among them, the × here represents the Hadamard product, that is, the element-wise multiplication operation. v is a random vector that independently and identically appears in the Bernoulli distribution with a probability of q. As one of the hyperparameters, the typical value of the probability of q is between 0.5 and 0.8. If the value of q is too large, the training time will be extended; otherwise, it may not be able to effectively avoid the overfitting phenomenon. The output of the current layer is λ(m), and the random vector v will be sampled and element-wise multiplied with the output λ(m) to obtain the refined output λ'(m).

[0104] Fully connected layer: including the first fully connected layer F1, the second fully connected layer F1, and the third fully connected layer F1, all using the Sigmoid classifier;

[0105] The two Dropout layers in this embodiment are used to flatten the input data of this layer or convert it into a one-dimensional vector, and then input this one-dimensional vector into the first fully connected layer. After testing, the feature learning effect obtained here is better. Since the objects captured by the convolutional layer in this embodiment are local features, the fully connected layer in this embodiment only combines the previous local overall features again through the weight matrix to generate a new graph. In this embodiment, a large number of features obtained from multiple previous convolutional layers and multiple pooling layers are connected through the fully connected layer, so that the feature nodes of each layer are connected to the feature nodes of the previous layer, thereby realizing the retention of more possible image features. Since the classification of malicious code in this embodiment is a binary classification problem, and the features of different malicious codes are mutually exclusive, the Sigmoid classifier is used for classification.

[0106] Output layer O: Output the file category;

[0107] As shown above, the file types in this embodiment are represented by numbers. For example, output 1 indicates that the PE file to be recognized is a benign file, and output 0 indicates that the PE file to be recognized is a malicious code file.

[0108] Step 5: Train the malicious code model with the first feature vector, the first grayscale image set or the first grayscale enhanced image set, and the second feature vector and the second grayscale image set or the second grayscale enhanced image set to obtain a malicious code training model.

[0109] In this embodiment, the scale of the multi-layer convolutional neural network model can be adjusted according to requirements. As the scale increases, more complex malicious code recognition and classification can be achieved. In other embodiments, according to the actual classification intensity, those skilled in the art can adjust the number and scale of each layer while ensuring the classification effect. When adjusting, it is necessary to try to select network parameters with fewer numbers, a moderate network size, and not too deep network layers to prevent overfitting or gradient dispersion phenomena and reduce the computational complexity.

[0110] Step 6: Load the PE file to be recognized and convert the PE file to be recognized into a grayscale image to be recognized through file pixelization.

[0111] Step 7: Detect the grayscale image to be recognized according to the obtained malicious code training model to obtain a detection result.

[0112] Step 8: Output the file type of the PE file to be recognized according to the obtained detection result. If the PE file to be recognized is a benign file, output the benign file class number. If the PE file to be recognized is a malicious code file, output the malicious code file class number.

[0113] In this embodiment, when recognizing a PE file to be recognized, it is detected by the malicious code training model and its file class number is output. If there are enough types of malicious code files in the PE file dataset, when the malicious code training model is recognizing, if the PE file to be recognized is a malicious code file, it can accurately output the class number of the family to which the malicious code file type belongs, rather than simply outputting the number 0 indicating that the PE file to be recognized is a malicious code file.

[0114] Based on the above method, the present invention provides a malicious code recognition device, including: a training unit, and a loading unit, an identification unit, and an output unit connected in sequence, wherein the training unit includes an acquisition module, an image module, a feature extraction module, and a model training module connected in sequence:

[0115] The acquisition module is used to acquire the PE file dataset and divide it into a training set and a validation set according to a certain ratio.

[0116] An image module, configured to pixelize the training set and the validation set through files respectively to convert them into a first grayscale image set and a second grayscale image set;

[0117] Alternatively, pixelize the training set and the validation set through files respectively to convert them into a first grayscale image set and a second grayscale image set, and perform data enhancement processing on the obtained first grayscale image set and second grayscale image set through the Gamma algorithm to obtain a first grayscale enhanced image set and a second grayscale enhanced image set;

[0118] A feature extraction module, configured to extract features from the obtained first grayscale image set or first grayscale enhanced image set, and the obtained second grayscale image set or second grayscale enhanced image set based on image texture features to obtain a first feature vector and a second feature vector;

[0119] A model training module, configured to establish a malicious code model, and train the malicious code model through the first feature vector, the first grayscale image set or the first grayscale enhanced image set, and the second feature vector, the second grayscale image set or the second grayscale enhanced image set to obtain a malicious code training model.

[0120] A loading unit, configured to load a PE file to be recognized, and pixelize the PE file to be recognized through a file to convert it into a grayscale image to be recognized;

[0121] An identification unit, configured to detect the grayscale image to be recognized according to the malicious code model to obtain a detection result;

[0122] An output unit, configured to output the file type of the PE file to be recognized according to the obtained detection result. If the PE file to be recognized is a benign file, output a benign file class number. If the PE file to be recognized is a malicious code file, output a malicious code file class number.

[0123] The malicious code model training method provided by the present invention can be applied in a computer-readable storage medium. The computer-readable storage medium stores a computer program. The above recognition method can be stored in the computer-readable storage medium as a computer program. When the computer program is executed by a processor, each step of the above training method is implemented.

[0124] In addition, the recognition method provided by the present invention can also be applied to a terminal device. The terminal device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the recognition method of the present invention are implemented. Here, the terminal device can be a computer, a notebook, a palm computer, and various cloud servers and other computing devices. The processor can be a general-purpose processor, a digital signal processor, an application-specific integrated circuit, or other programmable logic devices, etc.

[0125] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. For those of ordinary skill in the art, the specific technical solutions described in the foregoing embodiments can be modified, or some of the technical features can be equivalently replaced, and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions protected by the present invention.

Claims

1. A method for identifying malicious code, characterized in that: It includes the following steps: Step 1, load the PE file dataset and divide it into a training set and a validation set according to a certain ratio. Among them, the PE file dataset includes benign files and malicious code files; Step 2, respectively convert the training set and the validation set into a first grayscale image set and a second grayscale image set through file pixelization; Or, respectively convert the training set and the validation set into a first grayscale image set and a second grayscale image set through file pixelization, and perform data augmentation processing on the obtained first grayscale image set and second grayscale image set through the Gamma algorithm to obtain a first grayscale enhanced image set and a second grayscale enhanced image set; Among them, respectively converting the training set and the validation set into a first grayscale image set and a second grayscale image set through file pixelization includes the following steps: Step 2.1, convert the byte sequence of each PE file in the training set into a first vector group, and convert each first vector group into a corresponding first grayscale value group; Convert the byte sequence of each PE file in the validation set into a second vector group, and convert each second vector group into a corresponding second grayscale value group; Step 2.2, generate a first grayscale image set corresponding one-to-one with the training set according to the obtained first grayscale value group; Generate a second grayscale image set corresponding one-to-one with the validation set according to the obtained second grayscale value group; Step 3, perform feature extraction on the obtained first grayscale image set or first grayscale enhanced image set, and the obtained second grayscale image set or second grayscale enhanced image set based on image texture features to obtain a first feature vector and a second feature vector; If there are PE files of different sizes in the training set or the validation set, then perform a stepped division on the file size range in the training set or the validation set, divide the PE files of different sizes into different categories according to the range where their file sizes are located, and perform corresponding image width constraints on the PE files of different categories; Step 4, establish a malicious code model, and the malicious code model includes a multi-layer neural network model; Step 5, train the malicious code model through the first feature vector, the first grayscale image set or the first grayscale enhanced image set, and the second feature vector, the second grayscale image set or the second grayscale enhanced image set to obtain a malicious code training model; Step 6, load the PE file to be identified and convert the PE file to be identified into a grayscale image to be identified through file pixelization; Step 7, detect the grayscale image to be identified according to the obtained malicious code training model to obtain a detection result; Step 8, output the file type of the PE file to be identified according to the obtained detection result. If the PE file to be identified is a benign file, output the benign file class number. If the PE file to be identified is a malicious code file, output the malicious code file class number.

2. A method for identifying malicious code according to claim 1, characterized in that: The multi-layer neural network model includes: Input layer I: Input the first feature vector, and the first grayscale image set or the first grayscale enhanced image set; Input the second feature vector, and the second grayscale image set or the second grayscale enhanced image set; Convolutional layer: including a first convolutional layer C1, a second convolutional layer C2, and a third convolutional layer C3, all using a 3×3 convolutional kernel; Pooling layer: including a first pooling layer MP1, a second pooling layer MP2, and a third pooling layer MP3, all using the max pooling method, and the pooling width is set to 3×3; Dropout layer: including a first Dropout layer D1 and a second Dropout layer D2; Fully connected layer: including a first fully connected layer F1, a second fully connected layer F1, and a third fully connected layer F1, all using a Sigmoid classifier; Output layer O: output file type; Among them, the input layer I, the first convolutional layer C1, the first pooling layer MP1, the second convolutional layer C2, the second pooling layer MP2, the third convolutional layer C3, the third pooling layer MP3, the first Dropout layer D1, the first fully connected layer F1, the second Dropout layer D2, the second fully connected layer F2, the third fully connected layer F3, and the output layer O are connected in sequence for data connection.

3. A malicious code recognition method according to claim 2, characterized in that: In step 3, the SURF algorithm is selected for feature extraction.

4. A malicious code recognition device for implementing a malicious code recognition method according to any one of claims 1-3, characterized in that: It includes a training unit, and a loading unit, an identification unit, and an output unit that are connected in sequence; The training unit includes an acquisition module, an image module, a feature extraction module, and a model training module that are connected in sequence: The acquisition module is used to acquire a PE file dataset and divide it into a training set and a validation set according to a certain ratio; The image module is used to pixelize the training set and the validation set into a first grayscale image set and a second grayscale image set respectively through file pixelization; Alternatively, the training set and the validation set are respectively pixelized through file pixelization into a first grayscale image set and a second grayscale image set, and the obtained first grayscale image set and second grayscale image set are subjected to data enhancement processing through the Gamma algorithm to obtain a first grayscale enhanced image set and a second grayscale enhanced image set; The feature extraction module is used to extract features from the obtained first grayscale image set or first grayscale enhanced image set, and the obtained second grayscale image set or second grayscale enhanced image set based on image texture features to obtain a first feature vector and a second feature vector; The model training module is used to establish a malicious code model, and train the malicious code model through the first feature vector, the first grayscale image set or the first grayscale enhanced image set, and the second feature vector, the second grayscale image set or the second grayscale enhanced image set to obtain a malicious code training model; The loading unit is used to load the PE file to be recognized and pixelize the PE file to be recognized into a grayscale image to be recognized; The identification unit is used to detect the grayscale image to be recognized according to the malicious code training model to obtain a detection result; The output unit is configured to output the file type of the PE file to be recognized according to the obtained detection result. If the PE file to be recognized is a benign file, the benign file class number is output. If the PE file to be recognized is a malicious code file, the malicious code file class number is output.

5. A computer-readable storage medium, on which a computer program is stored, characterized in that: when the program is executed by a processor, the steps of a malicious code recognition method as described in any one of claims 1-3 are implemented.

6. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that: when the processor executes the computer program, the steps of a malicious code recognition method as described in any one of claims 1-3 are implemented.

Citation Information

Patent Citations

  • Training and detecting method and device of malicious code family

    CN107392019A

  • Malicious code detection method based on machine learning

    CN114692148A