Contrastive learning malware detection model based on multi-feature fusion
By adopting a comparative learning model of multi-feature fusion in malware detection, fusing API names, call relationships and parameter information, the problems of insufficient detection accuracy and incompetent code confusion in the existing technology are solved, and higher detection accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202310376848.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-10
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2043-04-10
AI Technical Summary
Existing malware detection technologies have problems such as insufficient detection accuracy, inconsistency in code and invariant mutation, especially the problems of insufficient single-dimensional information and one-sided feature information.
A contrast learning model based on multi-feature fusion is adopted, by constructing multi-dimensional features, fusing API names, API call relationship diagrams and API parameter information, using contrast learning to extract features and generate an encoder, and finally guiding the classifier to perform malware classification.
It improves the accuracy and robustness of malware detection, can more effectively resist malware mutations and code obfuscation, and provides more comprehensive feature information to support model training.
Smart Images

Figure CN118779874B_ABST
Abstract
Description
Technical Field
[0001] This patent relates to the field of malware detection technology and designs a comparative learning model based on multi-feature fusion to improve the accuracy of malware detection. Background Art
[0002] In the past few decades, with the rapid development of computer technology, malware has also been constantly evolving and mutating in various ways, discovering vulnerabilities and attacking them. These software seriously threaten user privacy and system security, and may cause irreparable losses. Therefore, it is very important to detect the existence of malware in a timely manner, especially some mutating malware-like malware. Generally speaking, the technology related to this patent can be divided into three aspects. First, the API sequence called by the software during execution is used to extract the information that may be contained by natural language processing or image processing technology as a method for detecting malware; second, the software behavior is considered from the perspective of the call graph, and the calling function and the called function are used as nodes of the graph, and the calling relationship is used as an edge. The graph matching algorithm is used to determine whether the graph conforms to the behavior pattern of the malware; third, the API call parameters are considered, because the API itself does not have any malicious intent, it may be that the call sequence or the input parameters are different from the normal situation, which leads to its malicious behavior, so the information contained in the API parameters cannot be ignored. Many studies have also conducted information extraction and learning on the huge parameter space of the API, and used it as a basis for determining malware.
[0003] All three methods of detecting malware have achieved a detection accuracy of about 97%, but each method has certain drawbacks. First, detection based solely on the API sequence is easily fooled by code obfuscation and packaging technology, resulting in poor detection results. As for the method of building a call graph, the problem is that not all call graphs can be accurately matched, and the design of the graph matching algorithm is not a simple matter, which leads to a lot of time spent on algorithm verification and additional processing of call graphs that cannot be matched. This obviously has certain technical barriers and reduces the accuracy. The problem with the last method of extracting information from API parameters is that the parameter space of the API is very large and it is difficult to classify and organize it uniformly. Whether the parameters are malicious and their maliciousness requires manual evaluation. First, an evaluation system for all parameters needs to be established, and then the parameters need to be classified and organized, and then machine learning is used to classify malware. These methods provide the classifier with relatively one-sided feature information. As malware continues to evolve, many models will have concept drift problems and cannot resist code obfuscation well. Summary of the invention
[0004] In order to overcome the shortcomings of the above-mentioned prior art, this patent proposes a new multi-feature fusion malware classification model based on contrastive learning, which effectively integrates features from different angles and relies on the powerful feature learning ability of contrastive learning to achieve higher accuracy and robustness.
[0005] The technical route of this patent is as follows:
[0006] Dynamic API call sequences are considered to be the most accurate features that can reflect and capture malicious behaviors. Many scholars have conducted extensive research in this area and found that the timing of call sequences, the networking of call graphs, and the parameters of API calls all play a certain role in identifying malware. However, due to different methods, the information entropy obtained may be inconsistent, so the training data given to the model will be different. In order to better mine the features, this patent integrates the information of these three dimensions by adding dimensions for analysis based on the idea of multi-feature fusion. As shown in the figure, first construct a multi-dimensional feature, extract the required API sequence name based on the behavior file of each malware, and encode the API call relationship graph and parameter information in the order in which the API appears to form a three-dimensional graph.
[0007] Secondly, contrastive learning has unlimited potential in the field of unsupervised representation learning. Compared with generative learning, contrastive learning does not need to pay attention to the tedious details of the instance, but only needs to learn to distinguish the data in the feature space at the abstract semantic level. Therefore, the model and its optimization become simpler and have stronger generalization ability. As shown in the second part of the figure, contrastive learning is used to extract the previously generated image features to generate an encoder to convert each malware into a vector representation.
[0008] Finally, the learned encoder parameters are used to guide the generation of an accurate classifier.
[0009] Compared with the existing technology, the excellent effect of this patent is that it proposes a feature construction method that integrates three dimensions of information: API name, API centrality, and API parameters, solving the problem of insufficient single-dimensional information and providing multiple levels of information for subsequent feature learning as much as possible. It also proposes solutions to possible malware mutations and code obfuscation problems, and uses the feature abstraction ability of contrastive learning to resist such attacks, which can improve the robustness of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The figure is a system structure diagram of the model proposed in this patent, showing the overall framework of this patent; DETAILED DESCRIPTION
[0011] The present patent is further described below in conjunction with the accompanying drawings.
[0012] Step 1: Collect the PE behavior files obtained through the sandbox and perform data analysis on them. Finally, we decide to use the API name, the process ID that calls the API, the original API parameter value, and the additional parameter value to complete the construction of multiple features. The API name sequence can provide the order in which the malware calls the API. The call graph will be generated by the API name and the process ID that calls the API. The centrality of each API is calculated based on the generated call graph. Finally, the information is supplemented based on the parameters used by the API at different calling periods.
[0013] Step 2, use Word2Vec as a language model to encode the API sequence. Word2Vec takes into account the context of words, learns semantic and grammatical information, and the word vector dimension obtained is small. Treat each malware API sequence as a piece of text with contextual connection, and use Word2Vec to encode each API into a word vector of length 4 to form the first channel of the image data. Here, the first 4096 APIs will be selected to represent a malware. Previous research techniques have confirmed that the malicious behavior of malware will be discovered in the first two minutes of software execution, and events will be frequently registered during this period. Therefore, in order to reduce unnecessary calculations, it is more appropriate to take the first 4096 APIs as the main calculation object. Based on this consideration, the data format of each channel of this patent takes the first 4096 APIs of the entire dynamic execution file for calculation. The encoding form of an API is 4*1. Finally, each channel forms a size of 4*4096, and then it is converted to 128*128 format for easy calculation.
[0014] Step 3, extract the API name and process call ID from the behavior file of the malware to obtain the API call graph. This graph is a weighted directed graph, and the weight is the frequency of occurrence of the call relationship. Each node is the API function or caller ID that appears. Calculate the centrality of each node based on the graph. Centrality is a concept used in social network analysis to measure the degree to which a point or a person in the network is close to the center of the entire network. This degree is expressed in numbers and is called centrality. Here, the centrality of an API is represented by calculating degree centrality, closeness centrality, Katz centrality, and harmonic centrality, which can improve the details of the graph to the greatest extent. The calculation formulas for the four centralities are as follows:
[0015] Degree centrality: The greater the node degree of a node, the greater the degree centrality of the node, that is, it is in a key position in the entire network and connects many nodes. The formula is as follows:
[0016]
[0017] where ki It represents the number of existing edges connected to node i, and N-1 represents the number of edges connecting node i to other nodes.
[0018] Closeness centrality: It represents the average value of the shortest distance from a node to other points in its connected components. It can help find the points in the network that are close to each other, that is, the key social nodes. The formula is as follows:
[0019]
[0020] Where y is any node that x can reach except itself, and the direction can be either inbound or outbound. d(x,y) is the shortest distance from x to y, and k-1 refers to the number of ys.
[0021] Katz centrality: Calculates the relative influence of a node in the entire network by measuring the number of directly adjacent nodes and all other nodes that can be connected to the network through other nodes. Nodes that are not directly adjacent will have a decay factor ∝ to re-proportion the distance. The node weight calculation formula for each node is as follows:
[0022]
[0023] k represents the nodes of the entire graph that can be accessed k times and included in the calculation. If there is no restriction, all the nodes of the entire graph are considered. n represents the number of nodes in the calculation range. A k Represents the total number of connections involved in the k-times between j and i.
[0024] Harmonic centrality: A variation of proximity centrality, the sum of the reciprocals of all the shortest distances from x to other nodes of a node. The higher the value, the higher the centrality. The formula is as follows:
[0025]
[0026] Where x represents a compute node, i is its serial number, j is all nodes in the graph except i, n-1 is the number of j, and d i,j is the shortest distance from i to j. When there is no edge connecting i and j, d(i,j)=+∞, then d(i,j)=0.
[0027] The centrality of each API is represented by a 4*1 vector, and the data format is formed as described in Step 2 to form the second channel data.
[0028] Step 4, based on previous work, select the most important parameter strings as considerations, including file path, dll file address, registry address and buffer. The file path is a parameter that needs to be considered because it may point to important files. All program executions of the dynamic link library need to be linked to the dll file, and there will also be some code and data shared, which can be easily exploited by malicious people. The registry refers to some strings that begin with "HKEY", which are used to store system and application settings. The string that begins with "MZ" usually points to the buffer of the entire PE file, and often appears in malicious PE files to overwrite existing data. The number of times the parameter appears is counted by matching the API parameter string to obtain the last channel data, and the data format is as described in Step 2.
[0029] Step 5: Combine the data of the above three channels into a 128*128*3 data format to provide data for subsequent model learning.
[0030] Step 6: Use supervised contrastive learning to train the encoder, treating all samples belonging to the same category as positive samples and samples of different categories as negative samples. The encoder can obtain the vector representation corresponding to each malware, which can separate malicious samples belonging to different categories.
[0031] Supervised contrastive learning consists of a data augmentation module, an encoder network, and a projection network. Data augmentation is the process of converting the input image x into a randomly enhanced image. For each image, two randomly enhanced sub-images are generated to represent different fields of view of the original image. The enhanced images are trained by the encoder network. Mapping to the representation space finally obtains the representation vector to represent the original image. In order to train and obtain the feature vector of the final image, it is necessary to calculate the difference in the projection distance of the two randomly enhanced sub-images generated above through the projection network, that is, the loss formula, and minimize the distance by continuously adjusting the encoding network parameters, and finally map the representation vector into a final vector z. In supervised contrastive learning, since all data have labels, taking positive samples as an example, supervised contrastive learning will expand the number of positive pairs in each x according to the label, that is, all sub-data generated by data enhancement with the same label are regarded as positive samples.
[0032] ResNet50 is used as the image encoder in the encoder network, which can extract image features well. A 128*128*3 data format is used in the image data construction stage. In order to make ResNet50 more suitable for malware classification tasks, this patent modifies the fully connected layer of the network and specifies its output as a unified 512-dimensional representation of a malware.
[0033] Step 7, when the classification is finally performed, the encoded 512-dimensional feature vector and its label are output to the classifier as a data set. The input of the ResNet50 classifier is set to the size of the encoder output dimension, that is, 512. The encoder-related learned model is used as the initial model of the ResNet50 classifier through transfer learning, and its related parameters are set to be consistent with the encoder.
[0034] In summary, this patent is used in the field of malware detection and proposes a new solution to the problems of incomplete feature information and inability to resist attacks such as code obfuscation in previous related studies.
Claims
1. A contrastive learning malware detection model based on multi-feature fusion, characterized in that: Includes the following modules: (1) Multi-feature extraction module: The multi-feature extraction module is used to select the first 4096 API calls from the malware as representatives, and extract each API name sequence feature, API centrality feature, and API parameter feature, including: API name sequence feature extraction: Use the Word2Vec model to encode each API into a word vector of length 4*1; API centrality feature extraction: construct a weighted directed API call graph based on the API name and process call ID, where the edge weight is the frequency of occurrence of the call relationship and the node is the API function or caller ID; calculate the degree centrality, closeness centrality, Katz centrality and harmonic centrality of the weighted directed graph to form a centrality feature vector with a length of 4*1; API parameter feature extraction: file path, dynamic link library (DLL) file address, registry address and buffer are used as API parameter features. The number of parameter occurrences is counted by matching the API parameter string to form a parameter feature vector with a length of 4*1. (2) Feature fusion module: The feature fusion module is used to fuse the API name sequence features, API centrality features, and API parameter features generated by the multi-feature extraction module, and specifically includes: For each API, the length of each feature is 4*1. After calculation for 4096 APIs, each type of feature forms a feature vector in 4*4096 data format; The 4*4096 feature vector is converted into a 128*128 feature vector, and finally a feature vector in a 128*128*3 data format is formed, wherein the three channels correspond to the API name sequence feature, the API centrality feature, and the API parameter feature respectively; (3) Supervised contrastive learning module: The supervised contrastive learning module includes data enhancement, encoder network and projection network, specifically including: Data enhancement: Randomly enhance the input 128*128*3 format feature vector to generate two sub-images to represent different fields of view of the original image; Encoder network: A modified ResNet50 model is used, and the output of its last fully connected layer is adjusted to 512 dimensions to map the enhanced sub-image to the representation space and obtain the representation vector; Projection network: The loss function is used to calculate the distance difference between two randomly enhanced sub-images in the projection space, and the parameters of the encoder network are continuously adjusted to minimize the projection distance between the two sub-images, and finally the representation vector is mapped into a unified final vector; (4) Classification module: The classification module uses a transfer learning method to take the encoder model trained in the supervised contrastive learning module as the initial model, specifically including: Set the relevant parameters of the classifier model to be consistent with the encoder network; The 512-dimensional feature vectors generated by the encoder network and their corresponding labels are input into the classifier as a data set for training, ultimately achieving classification detection of malware.
Citation Information
Patent Citations
Malicious software behavior detection and classification system based on deep learning
CN113961922A