A Malware Analysis Method and Device Based on the PE File Format

Through a clustering algorithm based on PE file format, PE file information is obtained and parsed, feature vectors are determined and clustered, the problem of low malware analysis is solved, and fast and effective malware analysis is achieved.

CN116192462BActive Publication Date: 2025-07-22NSFOCUS INFORMATION TECHNOLOGY CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211732234.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-30
Publication Date
2025-07-22
Estimated Expiration
2042-12-30

AI Technical Summary

Technical Problem

In the prior art, malware analysis is inefficient, relies on manual analysis and is difficult to quickly process large amounts of malware samples.

Method used

A clustering algorithm based on PE file format is adopted to obtain and parse PE file information, determine the feature vector, perform clustering and quadratic clustering, and form a target clustering set, and display the target clustering to assist in analysis.

Benefits of technology

It improves the efficiency of malware analysis, shortens analysis time, simplifies the user analysis process, and improves the interpretability of analysis results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116192462B_ABST
    Figure CN116192462B_ABST
Patent Text Reader

Abstract

This application relates to the field of network information security technology, and in particular, to a method and device for malicious software analysis based on the PE file format. In this method, multiple PE files are obtained. The multiple PE files are parsed to obtain file information. According to the file information, the feature vectors corresponding to the multiple PE files are determined. Each feature vector in the feature vector set is clustered to obtain a cluster set. Among them, the similarity between any feature vectors included in any cluster in the cluster set is less than or equal to the first threshold. Each cluster in the cluster set is clustered to obtain a target cluster set. Among them, the similarity between any clusters included in any target cluster in the target cluster set is less than or equal to the second threshold, and the second threshold is greater than the first threshold. The target cluster set is displayed. The above solution uses a clustering method to analyze malicious software in the PE file format, improving the efficiency of malicious software analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of network security technologies, and in particular, to a method and apparatus for analyzing malware based on the PE file format. Background Art

[0002] A large part of network attacks in the network are realized through malware. Therefore, the analysis of malware samples is an important part of network threat detection. Users can discover threats in a timely manner by analyzing malware samples and establish defense mechanisms.

[0003] However, the number of malware is extremely large, and currently the analysis of malware mostly relies on manual analysis, with low efficiency. Summary of the Invention

[0004] Embodiments of this application provide a method and apparatus for analyzing malware based on the PE file format, which are used to detect malware in a timely manner and improve the efficiency of analyzing malware.

[0005] In a first aspect, embodiments of this application provide a method for analyzing malware based on the PE file format, including: obtaining a plurality of PE files; parsing the plurality of PE files to obtain file information, where the file information includes format information and general information of the plurality of PE files; determining characteristic vectors corresponding to the plurality of PE files according to the file information; clustering each characteristic vector in the set of characteristic vectors to obtain a set of clusters, where any cluster in the set of clusters includes characteristic vectors with a similarity less than or equal to a first threshold, and the set of characteristic vectors includes characteristic vectors corresponding to each of the plurality of PE files; clustering each cluster in the set of clusters to obtain a set of target clusters, where any target cluster in the set of target clusters includes clusters with a similarity less than or equal to a second threshold, and the second threshold is greater than the first threshold; and displaying the set of target clusters.

[0006] In the above method, by clustering the characteristic vectors, the division of a large number of malware samples in the PE file format can be realized in a timely manner to assist users in analyzing malware. In this application, by clustering the set of clusters again to obtain a set of target clusters, the number of clusters can be reduced, the missing clusters can be easily determined, and it is convenient for users to analyze malware. At the same time, the method for analyzing malware based on the PE file format in this application is a clustering algorithm with a low time complexity, which can improve the efficiency of analyzing malware.

[0007] Optionally, determining the characteristic vectors corresponding to the plurality of PE files according to the file information specifically includes: in the case where the file information includes numerical information and text information, vectorizing the text information to obtain a text vector; and combining the values of the text vector with the values of the numerical information to obtain a characteristic vector.

[0008] In the above method, by vectorizing text-based information, text vectors are obtained. Then, the values of the text vectors are combined with the values of numerical information to obtain feature vectors, which facilitates subsequent measurement of the similarity between all feature vectors for clustering.

[0009] Optionally, according to the file information, determining the feature vectors corresponding to multiple PE files further includes: in the case where the file information is numerical information, using the numerical information as the feature vector.

[0010] In the above method, by using numerical information as the feature vector, it is convenient to measure the similarity between all feature vectors for subsequent clustering.

[0011] Optionally, before clustering each feature vector in the feature vector set, the method further includes: scaling multiple feature vectors using a scaling function to obtain scaled multiple feature vectors, where, when the value at the i-th position of any one of the multiple feature vectors is greater than zero, the value at the i-th position of the scaled feature vector is less than the value at the i-th position of the feature vector before scaling.

[0012] In the above method, by scaling multiple feature vectors using a scaling function, scaled multiple feature vectors are obtained. The range difference of the values of each feature vector can be reduced, while ensuring that the characteristics of the samples can be reflected. The range of the eigenvalues in each feature vector is restricted to ensure that the eigenvalues are within a reasonable interval, preventing a large difference in the value ranges of each feature vector in different dimensions.

[0013] Optionally, the scaling function satisfies the following formula:

[0014]

[0015] where, v i is the value at the i-th position of the feature vector.

[0016] Optionally, clustering each feature vector in the feature vector set to obtain a cluster set specifically includes: using any one feature vector in the feature vector set as the starting vector, clustering other feature vectors in the feature vector set whose similarity to the starting vector is less than or equal to the first threshold to form a cluster. Using any feature vector in the feature vector set whose similarity to the starting vector is greater than the first threshold as a new starting point, and returning to execute clustering other feature vectors in the feature vector set whose similarity to the starting point is less than or equal to the first threshold to form a cluster, to obtain the cluster set.

[0017] In the above method, a cluster set is obtained by clustering each feature vector in the feature vector set. Each feature vector can be preliminarily partitioned and assigned to the initial cluster set, which facilitates subsequent clustering to determine the target cluster set.

[0018] Optionally, the cluster set contains M clusters, where M is a positive integer. Clustering each cluster in the cluster set to obtain a target cluster set specifically includes: using any one feature vector in the Nth cluster as the starting vector, where N is an integer greater than 0 and less than or equal to M. Determining the similarity between the starting vector and the feature vectors included in the (N + K)th cluster, where K is an integer greater than 0 and less than or equal to M - N. When the similarity is less than or equal to the second threshold, the Nth cluster and the (N + k)th cluster are used as target clusters to obtain the target cluster set.

[0019] In the above method, by clustering the cluster set again to obtain the target cluster set, the number of clusters can be reduced and it is convenient to determine the missing clusters.

[0020] In a second aspect, an embodiment of the present application provides a malicious software analysis device based on the PE file format, including:

[0021] An acquisition module for acquiring a plurality of PE files;

[0022] An analysis module for analyzing a plurality of PE files to obtain file information, where the file information includes the format information and general information of the plurality of PE files;

[0023] The analysis module is further configured to determine the feature vectors corresponding to the plurality of PE files according to the file information;

[0024] A processing module for clustering each feature vector in the feature vector set to obtain a cluster set, where the similarity between any feature vectors included in any cluster in the cluster set is less than or equal to the first threshold, and the feature vector set includes the feature vectors corresponding to each of the plurality of PE files;

[0025] The processing module is further configured to cluster each cluster in the cluster set to obtain a target cluster set, where the similarity between any clusters included in any target cluster among the plurality of target clusters is less than or equal to the second threshold, and the second threshold is greater than the first threshold;

[0026] A display module for displaying the target cluster set.

[0027] In a third aspect, an embodiment of the present application further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, the processor implements any one of the malicious software analysis methods based on the PE file format in the first aspect above.

[0028] Fourthly, an embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, any of the malware analysis methods based on the PE file format in the first aspect is implemented.

[0029] Fifthly, an embodiment of the present application further provides a computer program product, including a computer program, and the computer program is executed by a processor to implement any of the malware analysis methods based on the PE file format in any item of the above first aspect.

[0030] For the technical effects brought by any implementation manner in the second aspect to the fifth aspect, reference may be made to the technical effects brought by the corresponding implementation manner in the first aspect, which will not be elaborated herein. Description of the Drawings

[0031] Figure 1 It is a schematic diagram of an application scenario of a malware analysis method based on the PE file format provided by an embodiment of the present application;

[0032] Figure 2 It is a flowchart of a malware analysis method based on the PE file format provided by an embodiment of the present application;

[0033] Figure 3 It is a schematic diagram of a cluster set provided by an embodiment of the present application;

[0034] Figure 4 It is a schematic diagram of a target cluster set provided by an embodiment of the present application;

[0035] Figure 5 It is an exemplary flowchart of a malware analysis method based on the PE file format provided by an embodiment of the present application;

[0036] Figure 6 It is a schematic diagram of a device for malware analysis based on the PE file format provided by an embodiment of the present application;

[0037] Figure 7 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed Embodiments

[0038] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are some, rather than all, of the embodiments of the present application. Based on the embodiments recorded in this application document, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the technical solutions of the present application.

[0039] It should be noted that in the description of this application, "a plurality of" is understood as "at least two". "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. The connection between A and B can represent two situations: A is directly connected to B and A is connected to B through C. In addition, in the description of this application, terms such as "first" and "second" are only used for the purpose of distinguishing descriptions, and cannot be understood as indicating or implying relative importance, nor can they be understood as indicating or implying order.

[0040] Most network attacks in the network are achieved through malware. Therefore, the analysis of malware samples is an important part of network threat detection. However, the number of malware is extremely large, and the analysis of malware currently mostly relies on manual analysis, with low efficiency. Therefore, it is impossible to analyze each malware one by one. In addition, a large number of malicious samples are usually derived from a small number of malicious samples. Therefore, by analyzing a small number of samples, the characteristics of the same type of samples can be understood without repeated analysis.

[0041] Therefore, the primary task of sample analysis is to focus on the analysis of high-value unknown samples. During the analysis process, it is usually necessary to screen the collected samples, remove known samples and low-value samples, and focus on the analysis of high-value samples.

[0042] However, the number of high-value samples after screening is still large. In order to process these samples faster, this application uses a clustering method to analyze the above samples. Since the samples belonging to the same cluster are similar after clustering, only by analyzing some samples in each cluster can the information of all samples in the entire cluster be mastered, thereby achieving the rapid analysis of a large number of unknown samples, discovering new families, new variants, etc., and saving analysis time.

[0043] The prior art mainly analyzes malware in the following ways.

[0044] In an existing technical method, the file is encoded by using a file hash algorithm. Then, when retrieving similar files, the files with the same hash value are regarded as similar files, thereby completing clustering. However, in order to avoid hash collisions, the hash algorithm is usually designed to be an algorithm with a large degree of randomness. Therefore, even if the files are not very different, completely different hash results may be generated, with extremely high randomness and difficult to measure the distance.

[0045] In another existing technical approach, the use of machine learning, especially deep learning methods for file representation and clustering, has been increasing. However, deep learning-based methods often require model training and are difficult to handle unknown samples. Models with good performance usually have complex structures and high computational costs when used. When faced with a large number of samples, there are performance bottlenecks. Moreover, the feature vectors or final results generated by such methods are usually high-dimensional vectors that are difficult for users to understand, resulting in the non-interpretability of the results. For users, such results are difficult to provide good analysis and reference.

[0046] It can be seen that due to the complex and large number of changes in malware itself, how to quickly analyze a large number of malware in the PE file format is an urgent problem to be solved currently.

[0047] In view of this, in the embodiments of the present application, in order to analyze malware based on the PE file format in real time, a malware analysis method based on the PE file format is proposed, including: obtaining a plurality of PE files; parsing the plurality of PE files to obtain file information, where the file information includes format information, byte information, and string information of the plurality of PE files; determining feature vectors corresponding to the plurality of PE files according to the file information; clustering each feature vector in the feature vector set to obtain a cluster set, where the similarity between any two feature vectors included in any cluster in the cluster set is less than or equal to a first threshold, and the feature vector set includes feature vectors corresponding to each of the plurality of PE files; clustering each cluster in the cluster set to obtain a target cluster set, where the similarity between any two clusters included in any target cluster in the target cluster set is less than or equal to a second threshold, and the second threshold is greater than the first threshold; and displaying the target cluster set.

[0048] Some terms related to the embodiments of the present application are introduced below:

[0049] 1. A Portable Executable (PE) file is a program file on the Microsoft Windows operating system that can run on the Windows system and perform specific functions. Common executable program (EXE) files, dynamic link library (DLL) files, system (SYS) files, Component Object Model (COM) files, etc. are all PE files. And such files usually need to follow a specific file format to run properly on the system.

[0050] 2. Malware refers to software that performs malicious acts in a computer system. Such malicious acts usually include infecting files, damaging the system, stealing data, etc. In the usage scenario of this application, malware refers to a PE file with malicious behavior. Unless otherwise specified hereinafter, malware, malicious samples, malicious code, and PE files all refer to such software.

[0051] 3. The clustering algorithm refers to the process of classifying a certain number of individual samples according to their respective characteristics and determining similar samples as the same class of samples. A class of samples after clustering can be called a cluster, and the samples within a cluster should be samples that the clustering algorithm considers to belong to the same type.

[0052] 4. The hash algorithm refers to a mapping method, usually a mathematical function operation, which can convert the input data into a mapping result within a specific range. This result is called a hash value. Different inputs may produce the same result after the hash algorithm, which is called a hash collision. The hash algorithm is generally very sensitive to the input data. With simple modification of the original data, the final hash value will change greatly. Usually, it is difficult to deduce the original data from the hash value through the hash algorithm.

[0053] In particular, the preferred embodiments of this application are described below in conjunction with the accompanying drawings of the specification. It should be understood that the preferred embodiments described herein are only used to illustrate and explain this application and are not used to limit this application. And without conflict, the embodiments of this application and the features in the embodiments can be combined with each other.

[0054] Figure 1 FIG. shows a schematic diagram of the application scenario of an optional malware analysis method based on the PE file format of this application. This scenario includes a server 100 and a terminal 101. The server 100 and the terminal 101 can be communicatively connected through a network to implement the malware analysis method based on the PE file format of this application.

[0055] Users can use the server 100 to interact with the terminal 101 through the network, such as receiving or sending messages, etc. Various client applications can be installed on the terminal 101, such as program writing applications, web browser applications, search applications, etc. The terminal 101 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, desktop computers, etc. The server 100 can be implemented as an independent server or a server cluster composed of multiple servers.

[0056] The server 100 is used to obtain multiple PE files, parse the multiple PE files to obtain file information, where the file information includes format information, byte information, and string information of the multiple PE files, determine the feature vectors corresponding to the multiple PE files according to the file information, cluster each feature vector in the feature vector set to obtain a cluster set, where any cluster in the cluster set includes feature vectors with a similarity less than or equal to a first threshold, and the feature vector set includes the feature vectors corresponding to each PE file in the multiple PE files, cluster each cluster in the cluster set to obtain a target cluster set, where any target cluster in the target cluster set includes clusters with a similarity less than or equal to a second threshold, the second threshold is greater than the first threshold, and display the target cluster set.

[0057] It can be understood that the malware analysis method based on the PE file format provided in the embodiments of the present application can be executed by the server 100 or by the terminal 101.

[0058] As Figure 2 shown, the flowchart of a malware analysis method based on the PE file format provided in the embodiments of the present application may specifically include the following operations. Hereinafter, the server is taken as an example of the execution subject for description.

[0059] S201: Obtain multiple PE files.

[0060] In a possible embodiment, Windows is the most widely used desktop operating system at present, and its executable files adopt the PE file format. The server can obtain the PE files of the Windows operating system.

[0061] In another possible embodiment, the server can log in to a website that can obtain PE files to obtain multiple PE files.

[0062] S202: Parse the multiple PE files to obtain file information.

[0063] Among them, the file information includes format information and general information of the multiple PE files.

[0064] In a possible embodiment, the PE file is a data stream organized in a linear structure. The server can parse it according to the basic structure of the PE file, starting from the MS-DOS header in sequence, followed by the PE header, the section table, and then all section entities to obtain file information.

[0065] Next, the file information will be introduced in detail.

[0066] Format information refers to information related to the PE format. Since PE format files need to follow a certain format for the system to load them. In the format information, there is numerical information as well as information in the form of strings. The format information can be summarized into the following five categories:

[0067] 1) File basic information: File basic information is the basic information related to the overall file. This type of information includes the actual length of the file, the virtual space length of the file, the number of imported and exported functions, whether it contains debug mode, thread-related information, redirection, the number of special flags, etc.

[0068] 2) File header information: File header information refers to the basic information contained in the file header. This type of information is mainly the information that the system needs to process when loading the file. These information include the file header timestamp, the target machine, the target system, the Dynamic Link Library (DLL), features (strings), the file linker version, the subsystem version, the length of the code block, the length of the header, etc.

[0069] 3) Imported functions: Imported functions are a list of other necessary functions that need to be imported before the program runs. Here, the present invention extracts the set of important imported functions, and represents an imported function by a function name string.

[0070] 4) Exported functions: Exported functions refer to the functions that a PE file can export. Similar to the imported functions, here the set of exported functions is selected, and an exported function is represented by a function name string.

[0071] 5) Segment information: The data composition of a PE file can be divided into different segments such as data segments, code segments, and stack segments. Segment information refers to the information related to these segments. These information include segment names, lengths, entropy, virtual lengths, representative string lists, etc.

[0072] General information refers to data that has no direct connection with the PE file information, mainly involving some information at the binary level and the coding level. General information can be summarized into the following three categories:

[0073] 1) Byte histogram: The byte histogram is generated by counting byte information. Regarding the file as a byte stream of 256 types of bytes, counting the number of each type of byte can generate a byte histogram. The byte histogram represents the distribution of 256 types of bytes.

[0074] 2) Byte entropy histogram: The byte entropy histogram is also generated by counting byte information. The byte entropy histogram uses a fixed-length window and samples bytes by moving at a specified step. Each time it moves, it calculates the entropy value with base 2 logarithm of the current window, and performs histogram statistics on the entropy values of all windows, and a 256-dimensional feature vector can be obtained.

[0075] 3) Related information of strings: The related information of strings refers to the set of all byte strings composed of printable characters (encoding range from 0x20 to 0x7f) in the file and with a length greater than 5. The number of these strings themselves is huge, so the strings themselves are not used as information. This application counts the number of strings, the average length, the histogram of printable characters, and the character entropy of printable characters. In addition, counting the number of occurrences of file paths, uniform resource locators (URLs), registry keys, and specific strings (such as Mark Zbikowski, MZ) through regular matching is also regarded as important information to be extracted.

[0076] S203: Determine the feature vectors corresponding to multiple PE files according to the file information.

[0077] In a possible embodiment, when the file information is numerical information, the server uses the numerical information as the feature vector. For example, assume that the file information of the PE file is numerical information. The basic file information is (131, 167…153) and the byte histogram is (24, 79…62). Then the feature vector is obtained by combining the numerical information of the basic file information and the byte histogram, resulting in (131, 167…153, 24, 79…62). It can be understood that when multiple types of information included in the file information are numerical information, the server combines the numerical information as the feature vector. This application does not make specific limitations on the order of combining numerical information.

[0078] In another possible embodiment, some of the information directly extracted from the PE file by the server exists in text form. To measure the similarity of all information subsequently, this application vectorizes the text information to obtain text vectors. And the obtained text vectors are combined with the numerical information as the feature vector of the PE file.

[0079] For example, the server of this application can vectorize different types of text information such as imported functions, exported functions, and full-text strings respectively to obtain text vectors. Finally, all the text vectors corresponding to the PE file and the numerical information are combined to obtain the feature vector.

[0080] It can be understood that since the order of combining the text vectors and the numerical information has no impact on subsequent clustering. Therefore, this application does not make specific limitations on the order of combining the text vectors and the numerical information.

[0081] In a possible embodiment, in order to vectorize text-based information, the server may use the Feature Hash method to vectorize the text-based information. Among them, the Feature Hash method can encode character-based information into a fixed-length vector.

[0082] The Feature Hash method can be formally described as the following process. The string set satisfies the following formula:

[0083] S = {(s i , n i ) | i = 1, 2, 3,..., l}

[0084] where s i represents the i-th string, and n i represents the number of times the string s i appears, and l represents the set length.

[0085] By inputting a string through the hash algorithm , the output is an integer obtained after the string undergoes a hash operation. And the algorithm output is i.e., an integer vector of length m.

[0086] In the above method, since the text-based information is vectorized through the Feature Hash method, the vectors of two similar texts will be very similar. Therefore, in order to facilitate subsequent clustering based on the feature vectors, this application can vectorize the text-based information through the Feature Hash method to obtain text vectors.

[0087] For example, the hash algorithm pseudocode is as follows:

[0088]

[0089] In another possible embodiment, the server of this application can also use methods such as the Bag of Word model and the Neural Network Language Model (NNLM) to vectorize the text-based information.

[0090] S204: Cluster each feature vector in the feature vector set to obtain a cluster set.

[0091] Among them, the similarity between any feature vectors included in a cluster in the cluster set is less than or equal to the first threshold. The feature vector set includes the feature vectors corresponding to each PE file in multiple PE files.

[0092] Since the ranges of the eigenvectors in the eigenvector set vary greatly in different dimensions. For example, the file length often has values in the tens of thousands or even millions, while the entropy usually has values less than 1. Without processing, the range differences in different dimensions will greatly affect the importance of different eigenvectors when calculating distances. For example, eigenvectors with smaller values will be severely ignored when calculating distances, while the actual meaning corresponding to this eigenvector may be relatively important.

[0093] Therefore, this application uses a feature scaling method to process the values of the eigenvectors in each dimension, adjusts the range differences of each feature, and at the same time ensures that the features of the samples can be reflected.

[0094] Suppose the eigenvector is V = (v1, v2, v3,..., v m ), and the feature scaling function satisfies the following formula:

[0095]

[0096] where, v i is the value at the i-th position of the eigenvector.

[0097] In a possible case, the server can use the scaling function to scale multiple eigenvectors to obtain multiple scaled eigenvectors. Among them, when the value at the i-th position of any eigenvector in the multiple eigenvectors is greater than zero, the value at the i-th position of the scaled eigenvector is less than the value at the i-th position of the eigenvector before scaling.

[0098] For example, the eigenvector before scaling is (12, 45, 79, 67). Then the eigenvector after scaling using the scaling function is (log 13, log46, log80, log68). Another example, the eigenvector before scaling is (-5, -123, -45, -20). Then the eigenvector after scaling using the scaling function is (-log6, -log124, -log46, -log21).

[0099] In order to cluster similar PE files, this application uses a clustering algorithm based on the greedy idea to cluster the eigenvectors. The clustering algorithm based on the greedy idea can complete the clustering task with a lower time complexity and at the same time ensure a certain degree of effectiveness credibility.

[0100] It should be noted that the embodiments of this application do not limit the clustering algorithm. The server can also use other clustering algorithms, such as the mean shift clustering algorithm, hierarchical clustering algorithm, etc. to cluster the eigenvectors.

[0101] In an alternative embodiment, the server may use any one of the feature vectors in the feature vector set as the starting vector, and cluster the other feature vectors in the feature vector set whose similarity to the starting vector is less than or equal to the first threshold to form clusters. Any feature vector in the feature vector set whose similarity to the starting vector is greater than the first threshold is used as a new starting point, and the process of clustering the other feature vectors in the feature vector set whose similarity to the starting point is less than or equal to the first threshold is repeated to obtain a set of clusters.

[0102] It can be understood that the first threshold can be an empirical value such as 5, 10, etc. preset by those skilled in the art, and can be reasonably set according to the specific application scenario. The similarity can also be a distance such as Euclidean distance, Manhattan distance, etc., which is not specifically limited in this application.

[0103] For example, the server may use a distance metric algorithm to determine the distance between feature vectors. Suppose the first threshold is 2. The server may use any one of the feature vectors in the feature vector set as the starting vector. The server obtains each of the other elements in the sample set, that is, the feature vectors. The distance between the other feature vectors and the starting vector is determined. If the distance between the feature vector and the starting vector is less than the first threshold, the feature vector and the starting vector are in the same cluster.

[0104] For example, the pseudocode of the clustering algorithm is as follows:

[0105]

[0106]

[0107] Among them, E = {e i | i = 1, 2, 3,..., k} is the sample set. e i is the i-th sample among them. The distance metric algorithm F dis (e i , e j ) outputs the distance between the samples e i and e j . The distance threshold ∈ s s is a real number. The set of clusters C = {c i} The samples are the set of clustering results, where c i represents the i-th cluster in the set of clusters. c i may contain multiple samples in the sample set E.

[0108] Optionally, the server in this application may use distance calculation as the basic operation of the clustering algorithm with the greedy idea.

[0109] Assume that the number of samples in the sample set is n. Then the server can determine that the highest time complexity of the clustering algorithm based on the greedy idea is O(n 2 ), which represents the time complexity required to run the algorithm when all feature vectors form their own clusters.

[0110] The optimal time complexity is O(n), which represents the time complexity required to run the algorithm when all feature vectors belong to one cluster.

[0111] The average time complexity is O(n log n), which represents the time complexity between the highest time complexity and the optimal time complexity. It can represent the time complexity required when, during the operation of the algorithm with most probabilities, there are both clusters formed by multiple feature vectors and clusters formed by single feature vectors.

[0112] Through the above method, the user can timely determine the time complexity of the clustering algorithm based on the greedy idea, which is convenient for the user to subsequently analyze malicious software in the PE file format in a timely manner.

[0113] S205: Cluster each cluster in the cluster set to obtain a target cluster set.

[0114] Among them, the similarity between any two clusters included in any target cluster in the target cluster set is less than or equal to a second threshold. The second threshold is greater than the first threshold.

[0115] It can be understood that the second threshold can be empirical values such as 10, 15, etc. preset by those skilled in the art and can be reasonably set according to specific application scenarios.

[0116] In order to reduce the number of clusters and discover missing clusters, perform a secondary analysis on similar samples that have not been merged into the same cluster in the cluster set and integrate the formed clusters. The server can cluster each cluster in the cluster set to obtain a target cluster set.

[0117] In an optional embodiment, the cluster set contains M clusters. Among them, M is a positive integer. The server uses any feature vector in the Nth cluster as the starting vector. Among them, N is an integer greater than 0 and less than or equal to M. The server determines the similarity between the starting vector and the feature vectors included in the (N + K)th cluster. Among them, K is an integer greater than 0 and less than or equal to M - N. When the similarity is less than or equal to the second threshold, the Nth cluster and the (N + k)th cluster are used as target clusters to obtain a target cluster set.

[0118] For example, the server can use a distance metric algorithm to confirm the distance between clusters. Assume that the second threshold is 5. The server can use any eigenvector in the cluster set as the starting vector. The server obtains all other elements in the sample set, that is, the eigenvectors. Determine the distance between the other eigenvectors and the starting vector. If the distances between all the eigenvectors in the Kth cluster and the starting vector are less than the second threshold, the Kth cluster and the cluster where the starting vector is located are in the same target cluster.

[0119] For example, the pseudo code of the clustering algorithm is as follows:

[0120]

[0121] Among them, C = {c i | i = 1, 2, 3,..., p} is the cluster set. c i is the ith sample among them. The distance metric algorithm F dis (e i , e j ) outputs the distance between the samples e i and e j . The distance threshold ∈ c is a real number. The cluster set C = {c i} samples are the set of clustering results, where c i represents the ith cluster in the cluster set. c i can contain multiple clusters in the cluster set C.

[0122] It can be understood that the method for clustering each cluster in the cluster set and determining the time complexity of the target cluster set in this application is the same as the method for clustering each eigenvector in the eigenvector set and determining the time complexity of the cluster set in the above text, which will not be elaborated here.

[0123] In the above method, the greedy clustering algorithm is used for clustering, which can perform fast clustering, discover associated samples, and analyze malware based on the PE file format.

[0124] As Figure 3 shown, the server clusters each eigenvector in the eigenvector set to obtain a cluster set. Among them, the cluster set includes cluster A, cluster B, cluster C, cluster D, and cluster E. As Figure 4 shown, the server clusters each cluster in the cluster set to obtain a target cluster set. Among them, the target cluster set includes target cluster AB, target cluster C, target cluster D, and target cluster E.

[0125] S206: Display the target cluster set.

[0126] The server sends the target cluster set to the terminal. After receiving the above target cluster set, the terminal displays the target cluster set on the electronic screen for the user to view.

[0127] In the above method, by displaying the target cluster set on the electronic screen for the user to view, the user experience is improved, enabling the user to analyze malware in a timely manner. It is convenient for the user to understand all the information of the malware in the entire target cluster through the target cluster set. Thus, rapid analysis of a large number of unknown malware is achieved, saving analysis time.

[0128] As Figure 5 shown, this application provides an exemplary flowchart for malware analysis based on the PE file format.

[0129] S501. Obtain multiple PE files;

[0130] S502. Parse the multiple PE files to obtain file information;

[0131] S503. In the case where the file information includes numerical information and text information, vectorize the text information to obtain text vectors;

[0132] S504. Combine the values of the text vectors with the values of the numerical information to obtain feature vectors;

[0133] S505. Scale the multiple feature vectors using a scaling function to obtain the scaled multiple feature vectors;

[0134] S506. Use any one feature vector in the feature vector set as the starting vector, and cluster other feature vectors in the feature vector set whose similarity to the starting vector is less than or equal to the first threshold to form clusters;

[0135] S507. Use any feature vector in the feature vector set whose similarity to the starting vector is greater than the first threshold as a new starting point, and return to execute clustering other feature vectors in the feature vector set whose similarity to the starting point is less than or equal to the first threshold to form clusters, to obtain a cluster set;

[0136] S508. Use any one feature vector in the Nth cluster as the starting vector, where N is an integer greater than 0 and less than or equal to M;

[0137] S509. Determine the similarity between the starting vector and the feature vectors included in the (N + K)th cluster, where K is an integer greater than 0 and less than or equal to M - N;

[0138] S510. In the case where the similarity is less than or equal to the second threshold, use the Nth cluster and the (N + k)th cluster as target clusters to obtain a target cluster set;

[0139] S511. Display the set of target clusters.

[0140] Furthermore, based on the same technical concept, an embodiment of the present application also provides a malware analysis device in the PE file format, which is used to implement the above malware analysis method process for the PE file format in the embodiment of the present application. Refer to Figure 6 As shown, the malware analysis device in the PE file format includes: an acquisition module 601, a parsing module 602, a processing module 603, and a display module 604. Among them:

[0141] The acquisition module 601 is used to acquire multiple PE files;

[0142] The parsing module 602 is used to parse multiple PE files to obtain file information, where the file information includes the format information and general information of multiple PE files;

[0143] The parsing module 602 is further used to determine the feature vectors corresponding to multiple PE files according to the file information;

[0144] The processing module 603 is used to cluster each feature vector in the feature vector set to obtain a cluster set, where the similarity between any feature vectors included in any cluster in the cluster set is less than or equal to a first threshold, and the feature vector set includes the feature vectors corresponding to each PE file in multiple PE files;

[0145] The processing module 603 is further used to cluster each cluster in the cluster set to obtain a target cluster set, where the similarity between any clusters included in any target cluster in the target cluster set is less than or equal to a second threshold, and the second threshold is greater than the first threshold;

[0146] The display module 604 is used to display the target cluster set.

[0147] Optionally, according to the file information, to determine the feature vectors corresponding to multiple PE files, the processing module 603 specifically is used for:

[0148] In the case that the file information includes numerical information and text information, vectorize the text information to obtain a text vector;

[0149] Combine the values of the text vector with the values of the numerical information to obtain a feature vector.

[0150] Optionally, according to the file information, to determine the feature vectors corresponding to multiple PE files, the processing module 603 is further used for:

[0151] In the case that the file information is numerical information, use the numerical information as the feature vector.

[0152] Optionally, before clustering each feature vector in the feature vector set, the processing module 603 is further configured to:

[0153] Scale multiple feature vectors using a scaling function to obtain scaled multiple feature vectors, where, when the value at the i-th position of any one of the multiple feature vectors is greater than zero, the value at the i-th position of the scaled feature vector is less than the value at the i-th position of the feature vector before scaling.

[0154] Optionally, the scaling function satisfies the following formula:

[0155]

[0156] where v i is the value at the i-th position of the feature vector.

[0157] Optionally, clustering each feature vector in the feature vector set to obtain a cluster set, the processing module 603 is specifically configured to:

[0158] Use any one feature vector in the feature vector set as the starting vector, and cluster other feature vectors in the feature vector set whose similarity to the starting vector is less than or equal to the first threshold to form a cluster;

[0159] Use any feature vector in the feature vector set whose similarity to the starting vector is greater than the first threshold as a new starting point, and return to execute clustering other feature vectors in the feature vector set whose similarity to the starting point is less than or equal to the first threshold to form a cluster to obtain a cluster set.

[0160] Optionally, the cluster set includes M clusters, where M is a positive integer. Clustering each cluster in the cluster set to obtain a target cluster set, the processing module 603 is specifically configured to:

[0161] Use any one feature vector in the N-th cluster as the starting vector, where N is an integer greater than 0 and less than or equal to M;

[0162] Determine the similarity between the starting vector and the feature vectors included in the (N + K)-th cluster, where K is an integer greater than 0 and less than or equal to M - N;

[0163] When the similarity is less than or equal to the second threshold, use the N-th cluster and the (N + k)-th cluster as the target clusters to obtain a target cluster set.

[0164] Based on the same inventive concept, an embodiment of the present application further provides an electronic device, which can implement the process of the malware analysis method for the PE file format provided in the above embodiments of the present application. In one embodiment, the electronic device may be a server, or a terminal device or other electronic devices. As Figure 7 shown, the electronic device may include:

[0165] At least one processor 701, and a memory 702 connected to the at least one processor 701. In the embodiments of the present application, the specific connection medium between the processor 701 and the memory 702 is not limited. Figure 7 In the example, it is assumed that the processor 701 and the memory 702 are connected through a bus 700. The bus 700 is Figure 7 shown by a thick line in the figure. The connection manners between other components are only for illustrative purposes and are not to be construed as limiting. The bus 700 can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 7 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus. Alternatively, the processor 701 may also be referred to as a controller, and the name is not limited.

[0166] In the embodiments of the present application, the memory 702 stores instructions executable by the at least one processor 701. By executing the instructions stored in the memory 702, the at least one processor 701 can execute a malware analysis method for a PE file format described above. The processor 701 can implement Figure 5 the functions of each module in the device shown in the figure.

[0167] Among them, the processor 701 is the control center of the device, and can connect various parts of the entire control device through various interfaces and lines. By running or executing the instructions stored in the memory 702 and calling the data stored in the memory 702, various functions of the device and process data, so as to monitor the device as a whole.

[0168] In a possible design, the processor 701 may include one or more processing units. The processor 701 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communications. It can be understood that the above modem processor may not be integrated into the processor 701. In some embodiments, the processor 701 and the memory 702 may be implemented on the same chip, and in some embodiments, they may also be implemented on separate chips independently.

[0169] The processor 701 may be a general-purpose processor, such as a CPU, a digital signal processor, an application-specific integrated circuit, a field-programmable gate array, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of a malicious software analysis method for a PE file format disclosed in combination with the embodiments of the present application may be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.

[0170] The memory 702, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. The memory 702 may include at least one type of storage medium, for example, it may include flash memory, a hard disk, a multimedia card, a card-type memory, a random access memory (RAM), a static random access memory (SRAM), a programmable read-only memory (PROM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic memory, a magnetic disk, an optical disk, and so on. The memory 702 is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 702 in the embodiments of the present application may also be a circuit or any other device capable of implementing a storage function, for storing program instructions and / or data.

[0171] By programming the design of the processor 701, the code corresponding to the malicious software analysis method for a PE file format introduced in the foregoing embodiments can be solidified into the chip, so that the chip can execute Figure 2 the steps of the malicious software analysis method for a PE file format shown in the embodiments when running. How to program the design of the processor 701 is a well-known technology to those skilled in the art and will not be elaborated here.

[0172] Based on the same inventive concept, the embodiments of the present application also provide a storage medium that stores computer instructions. When the computer instructions run on a computer, the computer is made to execute a malicious software analysis method for a PE file format discussed above.

[0173] In some possible embodiments, aspects of a malware analysis method for the PE file format provided by this application can also be implemented in the form of a program product, which includes program code. When the program product runs on a device, the program code is used to cause the control device to execute the steps in a malware analysis method for the PE file format according to various exemplary embodiments of this application described above in this specification.

[0174] Those skilled in the art should understand that the embodiments of this application can be provided as a method, a system, or a computer program product. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0175] This application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of this application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or multiple flows and / or blocks

[0176] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, such that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable devices provide steps for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or multiple flows and / or blocks

[0177] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these changes and modifications.

Claims

1. A malware analysis method based on the PE file format, characterized in that, The method includes: Obtaining a plurality of PE files; Parsing the plurality of PE files to obtain file information, where the file information includes the format information and general information of the plurality of PE files; Determining, according to the file information, the feature vectors corresponding to the plurality of PE files; Scale multiple feature vectors using a scaling function to obtain the scaled multiple feature vectors. Among them, when the value at the -th position of any one of the multiple feature vectors is greater than zero, the value at the -th position of the scaled feature vector is less than the value at the -th position of the feature vector before scaling. The scaling function satisfies the following formula: wherein, is the value at the -th position of the feature vector, is a set of feature vectors; Clustering each feature vector in the feature vector set to obtain a cluster set, where the similarity between any feature vectors included in any cluster in the cluster set is less than or equal to a first threshold, and the feature vector set includes the feature vectors corresponding to each PE file in the plurality of PE files; Clustering each cluster in the cluster set to obtain a target cluster set, where the similarity between any clusters included in any target cluster in the target cluster set is less than or equal to a second threshold, and the second threshold is greater than the first threshold; Displaying the target cluster set.

2. The method according to claim 1, wherein The determining, according to the file information, the feature vectors corresponding to the plurality of PE files specifically includes: In the case where the file information includes numerical information and text information, vectorizing the text information to obtain a text vector; Combining the values of the text vector with the values of the numerical information to obtain the feature vector.

3. The method according to claim 1, characterized in that The determining, according to the file information, the feature vectors corresponding to the plurality of PE files further includes: In the case where the file information is numerical information, using the numerical information as the feature vector.

4. The method according to claim 1, wherein The clustering each feature vector in the feature vector set to obtain a cluster set specifically includes: Taking any one feature vector in the feature vector set as a starting vector, and clustering other feature vectors included in the feature vector set that are similar to the starting vector and have a similarity less than or equal to the first threshold to form a cluster; Taking any feature vector in the feature vector set that is more similar to the starting vector than the first threshold as a new starting point, and returning to execute clustering other feature vectors included in the feature vector set that are similar to the starting point and have a similarity less than or equal to the first threshold to form a cluster, to obtain the cluster set.

5. The method according to claim 1, characterized in that The cluster set includes M clusters, where M is a positive integer. The clustering each cluster in the cluster set to obtain a target cluster set specifically includes: Taking any one feature vector in the Nth cluster as a starting vector, where N is an integer greater than 0 and less than or equal to M; Determining the similarity between the starting vector and the feature vectors included in the (N + K)th cluster, where K is an integer greater than 0 and less than or equal to M - N; In the case where the similarity is less than or equal to the second threshold, taking the Nth cluster and the (N + K)th cluster as target clusters to obtain the target cluster set.

6. A malware analysis device based on the PE file format, characterized in that The device includes: An obtaining module, configured to obtain a plurality of PE files; An analysis module, configured to analyze the plurality of PE files to obtain file information, where the file information includes the format information and general information of the plurality of PE files; The analysis module is further configured to determine, according to the file information, the feature vectors corresponding to the plurality of PE files; A processing module for scaling a plurality of feature vectors by using a scaling function to obtain a plurality of scaled feature vectors, wherein when the value at the -th position in any one of the plurality of feature vectors is greater than zero, the value at the -th position of the scaled feature vector is less than the value at the -th position of the feature vector before scaling, and the scaling function satisfies the following formula: wherein, is the value at the -th position of the feature vector, is a set of feature vectors; The processing module is further configured to cluster each feature vector in the feature vector set to obtain a cluster set, where the similarity between any feature vectors included in any cluster in the cluster set is less than or equal to a first threshold, and the feature vector set includes feature vectors corresponding to each of the multiple PE files; The processing module is further configured to cluster each cluster in the cluster set to obtain a target cluster set, where the similarity between any clusters included in any target cluster in the target cluster set is less than or equal to a second threshold, and the second threshold is greater than the first threshold; The display module is configured to display the target cluster set.

7. An electronic device, characterized in that, Comprising: A memory and a controller; The memory is configured to store program instructions; The controller is configured to call the program instructions stored in the memory and execute the method according to any one of claims 1-5 according to the obtained program.

8. A computer storage medium stores computer-executable instructions, characterized in that, The computer-executable instructions are used to execute the steps of the method according to any one of claims 1-5.

9. A computer program product, characterized in that, The computer program product includes: computer program code, which when running on a computer, causes the computer to execute the steps of the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Text clustering method and device

    CN112182206A

  • Text classification model training method and device, text classification method and device, equipment and medium

    CN114741517A