Software homology analysis method, device, computer equipment and readable storage medium

By using the target classification model and family gene resource center, the family gene gene resource center is used to automatically extract family genes of malware, solving the problem of inefficient traceability of complex malware in the existing technology, and achieving efficient homology analysis.

CN119987788BActive Publication Date: 2025-07-11CHINA TELECOM CLOUD TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510473377.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-11
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

The existing traceability analysis methods require a lot of time and labor costs to repeatedly analyze and compare complex malware samples, which is inefficient.

Method used

By obtaining the pre-trained target classification model, family classification is performed on the analysis software, family genes are obtained using interpreted analysis, and family gene databases and software samples are obtained from the family gene resource center, and the homology analysis results of the software to be analyzed and the target family are determined.

Benefits of technology

It realizes rapid traceability of analysis software, improves the efficiency of homology analysis, shortens analysis time, and reduces labor costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119987788B_ABST
    Figure CN119987788B_ABST
Patent Text Reader

Abstract

The present application relates to a software homology analysis method, apparatus, computer device, and readable storage medium. Among them, by obtaining the target family classification result of the software to be analyzed using a pre-trained target classification model, and based on the target family classification result, performing an interpretive analysis on the target classification model, the influence degree of each binary code segment in the software to be analyzed on the output result of the target classification model can be obtained, so that the family genes of the software to be analyzed can be obtained. Then, according to the target family classification result, the family gene library of the target family and the software samples in the target family are obtained from a pre-established family gene resource center, and based on the family genes of the software to be analyzed, the family gene library of the target family, and the software samples in the target family, the homology analysis result between the software to be analyzed and the target family is determined, realizing the traceability of the software to be analyzed, shortening the time for tracing the software to be analyzed, and improving the efficiency of homology analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of network security technology, and in particular, to a method, device, computer device, and computer-readable storage medium for software homology analysis. Background Art

[0002] With the development of computer technology, the spread and attack of malicious software are more frequent, and it is crucial to build a network attack defense system. In building a network attack defense system, tracing the source of attack events is a very important link. In the current field of traceability analysis, the common method is to manually compare binary codes to find and attribute and trace the samples. However, for some relatively complex samples, a large amount of time is required for repeated analysis and comparison work, which will consume a large amount of labor costs. Summary of the Invention

[0003] Based on this, in view of the above technical problems, it is necessary to provide a method, device, computer device, and computer-readable storage medium for software homology analysis that can improve the traceability efficiency.

[0004] In a first aspect, this application provides a method for software homology analysis, and the method includes:

[0005] Obtain the target family classification result of the software to be analyzed by the pre-trained target classification model;

[0006] Based on the target family classification result, perform an interpretive analysis on the target classification model to obtain the family gene of the software to be analyzed;

[0007] According to the target family classification result, obtain the family gene library of the target family and the software samples in the target family from the pre-established family gene resource center; the family classification result of the software samples in the target family is the same as the target family classification result;

[0008] Determine the homology analysis result of the software to be analyzed and the target family according to the family gene of the software to be analyzed, the family gene library of the target family, and the software samples in the target family.

[0009] In one embodiment, the target classification model includes multiple network layers; the step of, based on the target family classification result, performing an interpretive analysis on the target classification model to obtain the family gene of the software to be analyzed includes:

[0010] Determine background samples from the software sample set; the software sample set is used to train the target classification model; the background samples include at least one software sample in the software sample set;

[0011] Obtain the family classification results of each software sample in the background sample by the target classification model;

[0012] Obtain the propagation difference values of the software to be analyzed and each software sample in the background sample at each network layer of the target classification model;

[0013] According to the family classification results and propagation difference values of each software sample in the background sample, obtain the contribution degrees of each basic block in the software to be analyzed to the target family classification result;

[0014] According to the contribution degrees of each basic block in the software to be analyzed to the target family classification result, obtain the family gene of the software to be analyzed.

[0015] In one embodiment, the attribute features of the basic block include at least one of string constant, numerical constant, number of transfer instructions, number of call instructions, number of arithmetic instructions, total number of instructions, number of child nodes, and betweenness centrality; the obtaining the contribution degrees of each basic block in the software to be analyzed to the target family classification result according to the family classification results and propagation difference values of each software sample in the background sample includes:

[0016] According to the family classification results and the propagation difference values of each software sample in the background sample, calculate the marginal contribution value corresponding to each attribute feature of each basic block in the software to be analyzed respectively;

[0017] For each basic block of the software to be analyzed, obtain the corresponding contribution degree according to the marginal contribution values corresponding to all the attribute features of each basic block.

[0018] In one embodiment, the background sample includes multiple software samples; the calculating the marginal contribution value corresponding to each attribute feature of each basic block in the software to be analyzed respectively according to the family classification results and propagation difference values of each software sample in the background sample includes:

[0019] According to the family classification results and propagation difference values of each software sample in the background sample, calculate the initial marginal contribution value of each attribute feature of each basic block in the software to be analyzed relative to different software samples;

[0020] For each basic block in the software to be analyzed, determine the average value of the initial marginal contribution values of each attribute feature relative to different software samples as the marginal contribution value corresponding to each attribute feature.

[0021] In one embodiment, the determining the homology analysis result between the software to be analyzed and the target family according to the family gene of the software to be analyzed, the family gene library of the target family, and the software samples in the target family includes:

[0022] Match the family genes of the software to be analyzed with the software samples in the target family to obtain the number of successfully matched software.

[0023] Match the family genes in the family gene library of the target family with the family genes of the software to be analyzed to obtain the number of successfully matched genes.

[0024] According to the number of software and the number of genes, obtain the homology analysis result of the software to be analyzed and the target family.

[0025] In one embodiment, the method further includes:

[0026] Obtain the attribute control flow chart of each software sample in the software sample set.

[0027] Based on the attribute control flow chart, construct at least one sample classification model; the at least one sample classification model includes the target classification model.

[0028] Use each sample classification model to classify the software samples in the software sample set to obtain the family classification results of each software sample in the software sample set.

[0029] Conduct an interpretive analysis of the family classification results of each software sample in the software sample set to obtain the family genes of each software sample.

[0030] According to the family genes of each software sample, construct a family gene library for the corresponding family, and the family gene libraries of multiple families constitute the family gene resource center.

[0031] In one embodiment, the method further includes:

[0032] Obtain multiple pre-trained sample classification models.

[0033] Use each sample classification model to classify the software to be analyzed into a family, and determine the target classification model from each sample classification model.

[0034] In a second aspect, the present application also provides a software homology analysis device, and the device includes:

[0035] A classification result acquisition module, configured to obtain the target family classification result of the software to be analyzed by using a pre-trained target classification model.

[0036] A family gene acquisition module, configured to conduct an interpretive analysis of the target classification model based on the target family classification result to obtain the family genes of the software to be analyzed.

[0037] A homology analysis module is used to obtain the family gene library of the target family and the software samples in the target family from a pre-established family gene resource center according to the target family classification result; the family classification result of the software samples in the target family is the same as the target family classification result; and determine the homology analysis result of the software to be analyzed and the target family according to the family genes of the software to be analyzed, the family gene library of the target family, and the software samples in the target family.

[0038] In a third aspect, the present application also provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of the software homology analysis method provided in any of the above embodiments.

[0039] In a fourth aspect, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the software homology analysis method provided in any of the above embodiments.

[0040] In the above software homology analysis method, device, computer device, and computer-readable storage medium, by obtaining the target family classification result of the software to be analyzed by a pre-trained target classification model, and based on the target family classification result, performing an interpretive analysis on the target classification model, the influence degree of each binary code segment in the software to be analyzed on the output result of the target classification model can be obtained, so that the family genes of the software to be analyzed can be obtained. Then, according to the target family classification result, the family gene library of the target family and the software samples in the target family are obtained from a pre-established family gene resource center, and the homology analysis result of the software to be analyzed and the target family is determined according to the family genes of the software to be analyzed, the family gene library of the target family, and the software samples in the target family, realizing the traceability of the software to be analyzed, shortening the time for tracing the software to be analyzed, and improving the efficiency of homology analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required to be used in the description of the embodiments of the present application or related technologies. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other related drawings can be obtained without creative efforts based on these drawings.

[0042] Figure 1 It is a schematic flowchart of the software homology analysis method in an embodiment;

[0043] Figure 2Schematic diagram of the process of obtaining the family genes of the software to be analyzed for the target classification model based on the target family classification result in an embodiment;

[0044] Figure 3 Schematic diagram of the process of determining the homology analysis result between the software to be analyzed and the target family according to the family genes of the software to be analyzed, the family gene library of the target family, and the software samples in the target family in an embodiment;

[0045] Figure 4 Schematic diagram of the process of constructing a family gene library in an embodiment;

[0046] Figure 5 Architecture diagram of the software homology analysis method in an embodiment;

[0047] Figure 6 Schematic diagram of the process of the software homology analysis method in another embodiment;

[0048] Figure 7 Structural block diagram of the software homology analysis device in an embodiment;

[0049] Figure 8 Internal structure diagram of a computer device in an embodiment. Detailed implementation manners

[0050] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0051] In one embodiment, as Figure 1 shown, the present application provides a software homology analysis method, including step 102-step 108.

[0052] Step 102, obtaining the target family classification result of the target classification model for the software to be analyzed.

[0053] Among them, the target classification model can be a family classification model obtained by performing in-depth training using a one-dimensional convolutional neural network, such as neural network algorithms such as ResNet50 convolutional neural network, VGG19 convolutional neural network, and Xception convolutional neural network. The software to be analyzed is mainly malicious software that threatens network security or user interests, and it can be a binary code segment or a computer program product, etc. In the present application, a batch of malicious software that is very likely to be written by the same organization or person is called a "family". The target family classification result includes the belonging family of the software to be analyzed.

[0054] If the software to be analyzed contains a shell, before classifying the software to be analyzed into a family using a pre-trained target classification model, it is also necessary to perform unpacking on the software to be analyzed. Exemplarily, the Unipacker tool can be used to perform unpacking on the software to be analyzed.

[0055] Step 104: Based on the target family classification result, perform an interpretive analysis on the target classification model to obtain the family genes of the software to be analyzed.

[0056] The family genes of the software to be analyzed actually refer to the binary code segments in the software to be analyzed. The software to be analyzed can be divided into multiple consecutive binary code segments. By performing an interpretive analysis on the target classification model based on the target family classification result, the influence degree of each binary code segment in the software to be analyzed on the output result of the target classification model can be obtained. In this way, the binary code segment with the greatest influence degree can be used as the family genes of the software to be analyzed. Exemplarily, the SHAP (SHapley Additive exPlanations) framework can be used to perform an interpretive analysis on the target classification model to obtain the family genes of the software to be analyzed.

[0057] Step 106: According to the target family classification result, obtain the family gene library of the target family and the software samples in the target family from the pre-established family gene resource center; the family classification result of the software samples in the target family is the same as the target family classification result.

[0058] The family gene resource center can be a database containing family gene libraries and software samples of different families established based on the data accumulated in the early stage. The genes in the family gene library are binary code segments with similar behaviors in malicious software of the same family. The similar behaviors can be using the same IP address or the same domain name, etc.

[0059] Step 108: Determine the homology analysis result of the software to be analyzed and the target family according to the family genes of the software to be analyzed, the family gene library of the target family, and the software samples in the target family.

[0060] After obtaining the family genes of the software to be analyzed, the family gene library and software samples of the target family can be respectively matched with the family genes of the software to be analyzed to determine the homology analysis result of the software to be analyzed and the target family, so as to trace the origin of the software to be analyzed.

[0061] In the embodiments of the present application, the software to be analyzed is classified into families by the "model + gene" method. Specifically, the target family classification result of the software to be analyzed is obtained by using a pre-trained target classification model. Based on the target family classification result, the target classification model is interpreted and analyzed, and the influence degree of each binary code segment in the software to be analyzed on the output result of the target classification model can be obtained, so that the family gene of the software to be analyzed can be obtained. Then, according to the target family classification result, the family gene library of the target family and the software samples in the target family are obtained from the pre-established family gene resource center, and the homology analysis result between the software to be analyzed and the target family is determined according to the family gene of the software to be analyzed, the family gene library of the target family, and the software samples in the target family, realizing the traceability of the software to be analyzed, shortening the time for tracing the software to be analyzed, and improving the efficiency of homology analysis.

[0062] In one embodiment, as Figure 2 shown, based on the target family classification result, for the target classification model, obtaining the family gene of the software to be analyzed includes steps 202 - 210.

[0063] Step 202, determine background samples from the software sample set. The software sample set is used to train the target classification model, and the background samples include at least one software sample in the software sample set.

[0064] The software sample set includes benign software and malicious software of known families. Benign software can be installed in batches through software installation, and binary files with suffixes such as exe, sys, and dll on the disk after the installation of benign software are collected. Malicious software with family tags is obtained by using threat intelligence sharing platforms and actual combat collection, etc. If the obtained benign software and malicious software include shells, the Unipacker tool can be used to unpack the software samples respectively. The software sample set is constructed according to the obtained benign software and malicious software. The software sample set can be used to train the target classification model. When using the software sample set to train the target classification model, the software sample set can be divided into a training set and a test set. The training set is used to train the target classification model, and the test set can be used to verify the accuracy of the family classification result output by the target classification model. Exemplarily, the software sample set can be divided into a training set and a test set according to a ratio of 8:2, and after performing a shuffle operation and redistributing the data in the software sample set, the test set is used to train the target classification model. At least one malicious software can be randomly selected from the software sample set as the background sample.

[0065] Step 204, obtain the family classification results of the target classification model for each software sample in the background samples.

[0066] Denote the background samples as , each software sample in the background sample is , . The family classification results output by the target classification model for each software sample in the background sample are .

[0067] Step 206: Obtain the propagation difference values of the software to be analyzed and each software sample in the background sample in each network layer of the target classification model.

[0068] The target classification model includes multiple network layers. Exemplarily, it includes an input layer, an output layer, and at least one hidden layer. In order to attribute the output result differences of the target classification model for the software to be analyzed and the background sample to the input features, it is necessary to first obtain the propagation difference values of the software to be analyzed and each software sample in the background sample in each network layer of the target classification model.

[0069] Assume that the number of network layers of the target classification model is L. For the software x to be analyzed and the software sample , The propagation difference value of layer is:

[0070]

[0071] Among them, is the activation value of layer and can be obtained by backpropagating the propagation difference of the next layer:

[0072]

[0073] Among them, represents the weight of layer represents the bias of layer represents the activation function.

[0074] Step 208: According to the family classification results and propagation difference values of each software sample in the background sample, obtain the contribution degrees of each basic block in the software to be analyzed to the target family classification result.

[0075] When analyzing the software to be analyzed in this application, the control flow information of the code in the software to be analyzed is represented in the form of a directed graph, which is composed of a series of basic blocks. A basic block is composed of a series of code statements. When executed, a basic block has two characteristics. One is that it can only enter the basic block from the first jump statement and cannot enter the middle of the basic block in some way. The other is that the statements in the basic block must leave from the last statement when executed and cannot jump to other basic blocks halfway through. The attribute characteristics of a basic block include at least one of string constants, numeric constants, the number of transfer instructions, the number of call instructions, the number of arithmetic instructions, the total number of instructions, the number of child nodes, and betweenness centrality. Exemplarily, the Radare2 tool can be used to parse the software to be analyzed and extract the Attributed Control Flow Graph (ACFG) of the sample to be analyzed.

[0076] When obtaining the contribution degree of each basic block in the software to be analyzed to the target family classification result, the input difference between the software to be analyzed and the background samples can be decomposed into each attribute characteristic of the basic block, and the marginal contribution value of each attribute characteristic can be calculated.

[0077] Specifically, the marginal contribution value corresponding to each attribute characteristic of each basic block in the software to be analyzed can be calculated respectively according to the family classification result and the propagation difference value of each software sample in the background samples. The attribute characteristic The marginal contribution value Can be calculated by the following formula:

[0078]

[0079] Among them, Is the partial derivative of the target classification model with respect to the feature At the software sample , and the following formula is used for approximate calculation:

[0080]

[0081] If there are multiple software samples in the background samples, then based on Equation (3), the initial marginal contribution values of each attribute characteristic of each basic block in the software to be analyzed with respect to different software samples can be calculated respectively according to the family classification results and the propagation difference values of each software sample in the background samples. . For each basic block in the software to be analyzed, the average value of the initial marginal contribution values of each attribute characteristic with respect to different software samples is determined as the marginal contribution value corresponding to each attribute characteristic. Exemplarily, for the attribute characteristic i in a single basic block, the marginal contribution value can be obtained through the following formula :

[0082]

[0083] Among them, represents the number of software samples in the background sample .

[0084] After obtaining the marginal contribution value corresponding to each attribute feature of each basic block, the marginal contribution value corresponding to each attribute feature of each basic block can be stored as a numpy file in the order corresponding to the attribute feature of the basic block one by one.

[0085] Furthermore, for each basic block of the software to be analyzed, according to the marginal contribution values corresponding to all the attribute features of each basic block, the corresponding contribution degree is obtained. The numpy file can be read, and the marginal contribution values of the attribute features in each dimension of the same basic block are weighted and averaged to obtain the contribution degree corresponding to the basic block. The weights of the attribute features in each dimension of the same basic block can be set manually, and the embodiments of the present application do not limit this here.

[0086] Step 210, according to the contribution degree of each basic block in the software to be analyzed to the target family classification result, obtain the family gene of the software to be analyzed.

[0087] The contribution degree of each basic block to the target family classification result reflects the influence degree of each basic block on the target family classification result. The basic blocks in the software to be analyzed can be sorted in descending order according to the contribution degree, and the binary data corresponding to the preset number of basic blocks ranked at the top can be used as the family gene of the software to be analyzed. The preset number can be reasonably set according to empirical rules. Exemplarily, the preset number can be one, two, five or more, etc.

[0088] In this embodiment, the difference between the output results of the target classification model for the software to be analyzed and the background sample is attributed to the input features through layer-by-layer backpropagation. By obtaining the contribution degree of each basic block in the software to be analyzed to the target family classification result, the key basic blocks that affect the target family classification result in the software to be analyzed are obtained, and the binary data corresponding to the key basic blocks is used as the family gene of the software to be analyzed, realizing the automatic extraction of the family gene of the software to be analyzed.

[0089] In one embodiment, as Figure 3 shown, according to the family gene of the software to be analyzed, the family gene library of the target family, and the software samples in the target family, determine the homology analysis result between the software to be analyzed and the target family, including Step 302 - Step 306.

[0090] Step 302, match the family gene of the software to be analyzed with the software samples in the target family to obtain the number of successfully matched software.

[0091] The successful matching of the family genes of the software to be analyzed with the software samples can mean that the coincidence rate of the software samples and the family genes of the software to be analyzed exceeds a first preset value. The first preset value can be reasonably set according to experience.

[0092] The first YARA (Yet Another Recursive Acronym) rule can be defined based on the binary data after removing the location information of the family genes of the software to be analyzed. The first YARA rule is used to match in the software samples of the target family to obtain the number of successfully matched software. The definition process of the first YARA rule can refer to the definition of the conventional YARA rule, which will not be described in detail here.

[0093] Step 304: Match the family genes in the family gene library of the target family with the family genes of the software to be analyzed to obtain the number of successfully matched genes.

[0094] The successful matching of the family genes in the family gene library of the target family with the family genes of the software to be analyzed can mean that the coincidence rate of the family genes in the family gene library of the target family and the family genes of the software to be analyzed exceeds a second preset value. The second preset value can be reasonably set according to experience.

[0095] The second YARA rule can be defined based on the binary data after removing the location information of the family genes in the family gene library of the target family. The second YARA rule is used to match in the software to be analyzed to obtain the number of successfully matched genes. The definition process of the second YARA rule can refer to the definition of the conventional YARA rule, which will not be described in detail here.

[0096] Step 306: Obtain the homology analysis result of the software to be analyzed and the target family according to the number of software and the number of genes.

[0097] Specifically, when either the number of software or the number of genes is higher than the corresponding preset value, it can be considered that the correlation between the software to be analyzed and the target family is high. When both the number of software and the number of genes are lower than the corresponding preset values, it can be considered that the correlation between the software to be analyzed and the target family is low. The preset values corresponding to the number of software and the number of genes can be reasonably set based on empirical rules. Exemplarily, they can be 0 or 1.

[0098] In this embodiment, by matching the family genes of the software to be analyzed with the software samples in the target family, and by matching the family genes in the family gene library of the target family with the family genes of the software to be analyzed, the homology analysis result of the software to be analyzed and the target family is determined through a cross - verification method, ensuring the accuracy of the homology analysis result.

[0099] In one embodiment, as Figure 4As shown, the software homology analysis method of the present application further includes steps 402 to 410.

[0100] Step 402: Obtain the attribute control flow chart of each software sample in the software sample set.

[0101] The attribute control flow chart is a directed graph composed of a series of basic blocks with attribute characteristics. The attribute characteristics of the basic blocks include eight characteristics such as string constants, numerical constants, number of transfer instructions, number of call instructions, number of arithmetic instructions, total number of instructions, number of child nodes, and betweenness centrality. The software sample set has been described above and will not be elaborated here. It should be noted that in this embodiment, the software sample set can be used to train multiple sample classification models.

[0102] Step 404: Construct at least one sample classification model based on the attribute control flow chart, and at least one sample classification model includes a target classification model.

[0103] First, each basic block in multiple ACFG graphs can be arranged in the order of non-jump priority to obtain multiple ACFG feature sequences. Non-jump priority is for basic blocks with conditional jump instructions at the end. For example, if basic block A has a conditional jump instruction at the end, jumping to execute basic block C and not jumping to execute basic block B, then the position of basic block B in the ACFG feature sequence is before basic block C.

[0104] The currently commonly used one-dimensional convolutional neural network algorithm can be used to learn multiple ACFG feature sequences respectively to obtain multiple sample classification models. That is, each one-dimensional convolutional neural network algorithm corresponds to a sample classification model.

[0105] Step 406: Use each sample classification model to classify each software sample in the software sample set into families, and obtain the family classification results of each software sample in the software sample set.

[0106] Step 408: Interpret and analyze the family classification results of each software sample in the software sample set to obtain the family genes of each software sample.

[0107] Specifically, the classification effects of each sample classification model on each software sample in the software sample set can be statistically analyzed first. Exemplarily, the F1 Score can be used to measure the classification effect of each sample classification model on the same family. The F1 Score is an index in statistics used to measure the accuracy of a binary classification model, which takes into account both the precision and recall of the classification model. For the same family, the higher the F1 Score of the sample classification model, the better the classification effect. When interpreting and analyzing the family classification results of each software sample in the software sample set, only the sample classification model with the best classification effect needs to be analyzed based on the same principle as in steps 202 - 210 to obtain the family genes of each software sample.

[0108] Step 410, construct a family gene library for the corresponding family according to the family genes of each software sample, and the family gene libraries of multiple families constitute a family gene resource center.

[0109] It can be understood that the family gene library of a family can be constructed according to the family genes of the software samples belonging to the same family in each software sample. Then, a family gene resource center can be constructed according to the family gene libraries of different families.

[0110] In this embodiment, by interpreting and analyzing the family classification results of each software sample in the software sample set, the family genes of each software sample are obtained, and a family gene library for the corresponding family is constructed according to the family genes of each software sample, realizing the rapid construction of the family gene library without having to compare each sample of the malware family one by one. Moreover, when training the sample classification model, the software training set used not only includes malware but also benign software. Incorporating benign software into the consideration scope of the sample classification model can reduce the false positive rate when matching the family genes of the software to be analyzed extracted based on the sample classification model with the family gene library and software samples of the target family.

[0111] In one embodiment, the software homology analysis method of the present application further includes steps of obtaining multiple pre-trained sample classification models, using each sample classification model to perform family classification on the software to be analyzed, and determining a target classification model from each sample classification model.

[0112] It can be understood that the target classification model is the sample classification model with the best classification effect on the software to be analyzed among multiple sample classification models. Exemplarily, Model A classifies the software to be analyzed into Family X, Model B classifies the software to be analyzed into Family Y, the F1 Score of Model A for Family X is 0.75, and the F1 Score of Model B for Family Y is 0.91. Then, Model B will be selected as the target classification model for the software to be analyzed.

[0113] For better understanding, as Figure 5 andFigure 6 As shown, a more specific embodiment is enumerated to illustrate the software homology analysis method of the present application.

[0114] Step 602: Perform unpacking processing on the software sample set. The software sample set includes benign software and malware of known families.

[0115] Step 604: Extract the ACFG features of each basic block in the software sample.

[0116] Step 606: Arrange the basic blocks in each software sample in the order of non-jump priority to obtain multiple ACFG feature sequences.

[0117] Step 608: Use multiple ACFG feature sequences to train sample classification models of different convolutional neural network architectures.

[0118] Step 610: Statistically analyze the classification effects of multiple sample classification models to determine the sample classification model with the best classification effect corresponding to each family.

[0119] Step 612: For the software to be explained, analyze the key basic blocks that affect the family classification results of the sample classification model. The software to be explained is one of the software samples in the software sample set that have not been analyzed and the samples to be analyzed.

[0120] Step 614: Use the binary data corresponding to the key basic blocks as the family genes of the software to be explained.

[0121] Step 616: In the case where the software to be explained is a software sample in the software sample set that has not been analyzed, construct a corresponding family gene library according to the family genes of the software to be explained.

[0122] Step 618: In the case where the software to be explained is the software to be analyzed, convert the extracted family genes of the software to be analyzed into the first YARA rule and match them with the genes in the family gene library of the target family, and convert the genes in the family gene library of the target family into the second YARA rule and match them with the family genes of the software to be analyzed to determine the homology analysis result.

[0123] It should be understood that although the steps in the flowcharts involved in the above embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0124] Based on the same inventive concept, an embodiment of the present application further provides a software homology analysis device for implementing the software homology analysis method involved above. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the software homology analysis device provided below can refer to the limitations on the software homology analysis method in the above text, and will not be repeated here.

[0125] In an exemplary embodiment, as Figure 7 shown, a software homology analysis device is provided, including: a classification result acquisition module 702, a family gene acquisition module 704, and a homology analysis module 706, where:

[0126] The classification result acquisition module 702 is used to obtain the target family classification result of the software to be analyzed by the pre-trained target classification model;

[0127] The family gene acquisition module 704 is used to perform interpretive analysis on the target classification model based on the target family classification result to obtain the family gene of the software to be analyzed;

[0128] The homology analysis module 706 is used to obtain the family gene library of the target family and the software samples in the target family from the pre-established family gene resource center according to the target family classification result; the family classification result of the software samples in the target family is the same as the target family classification result; according to the family gene of the software to be analyzed, the family gene library of the target family, and the software samples in the target family, determine the homology analysis result of the software to be analyzed with the target family.

[0129] In one embodiment, the family gene acquisition module is further configured to determine background samples from the software sample set; the software sample set is used to train the target classification model; the background samples include at least one software sample in the software sample set; obtain the family classification results of each software sample in the background samples by the target classification model; obtain the propagation difference values of the software to be analyzed and each software sample in the background samples in each network layer of the target classification model; obtain the contribution degrees of each basic block in the software to be analyzed to the target family classification result according to the family classification results and propagation difference values of each software sample in the background samples; obtain the family genes of the software to be analyzed according to the contribution degrees of each basic block in the software to be analyzed to the target family classification result.

[0130] In one embodiment, the family gene acquisition module is further configured to calculate the marginal contribution value corresponding to each attribute feature of each basic block in the software to be analyzed respectively according to the family classification results and propagation difference values of each software sample in the background samples; for each basic block of the software to be analyzed, obtain the corresponding contribution degree according to the marginal contribution values corresponding to all attribute features of each basic block.

[0131] In one embodiment, the family gene acquisition module is further configured to calculate the initial marginal contribution value of each attribute feature of each basic block in the software to be analyzed relative to different software samples respectively according to the family classification results and propagation difference values of each software sample in the background samples; for each basic block in the software to be analyzed, determine the average value of the initial marginal contribution values of each attribute feature relative to different software samples as the marginal contribution value corresponding to each attribute feature.

[0132] In one embodiment, the homology analysis module is further configured to match the family genes of the software to be analyzed with the software samples in the target family to obtain the number of successfully matched software; match the family genes in the family gene library of the target family with the family genes of the software to be analyzed to obtain the number of successfully matched genes; obtain the homology analysis result of the software to be analyzed and the target family according to the number of software and the number of genes.

[0133] In one embodiment, the software homology analysis device further includes a model training module and a gene library construction module.

[0134] Among them, the model training module is configured to obtain the attribute control flow chart of each software sample in the software sample set; construct at least one sample classification model based on the attribute control flow chart; at least one sample classification model includes the target classification model.

[0135] The gene bank construction module is used to classify the software samples in the software sample set by family using each sample classification model, and obtain the family classification results of each software sample in the software sample set; interpret and analyze the family classification results of each software sample in the software sample set to obtain the family genes of each software sample; construct a family gene bank for the corresponding family according to the family genes of each software sample, and the family gene banks of multiple families constitute the family gene resource center.

[0136] In one embodiment, the software homology analysis device further includes a target model determination module, which is used to obtain multiple pre-trained sample classification models; classify the software to be analyzed by family using each sample classification model, and determine a target classification model from each sample classification model.

[0137] Each module in the above software homology analysis device can be implemented in whole or in part by software, hardware and their combination. The above modules can be embedded in the processor in the computer device in hardware form or independent of it, or stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0138] In an exemplary embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 8 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the extracted family genes, etc. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a software homology analysis method.

[0139] Those skilled in the art can understand that Figure 8 the structure shown in

[0140] In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the software homology analysis method provided in any of the above embodiments is implemented.

[0141] In an embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the software homology analysis method provided in any of the above embodiments is implemented.

[0142] In an embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the software homology analysis method provided in any of the above embodiments is implemented.

[0143] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., and are not limited thereto. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., and are not limited thereto.

[0144] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in the present application.

[0145] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A software homology analysis method, characterized in that The method includes: Obtaining the target family classification result of the software to be analyzed by the pre-trained target classification model; Determining background samples from the software sample set; the software sample set is used to train the target classification model; the background samples include at least one software sample in the software sample set; Obtaining the family classification results of each software sample in the background samples by the target classification model; Obtaining the propagation difference values of the software to be analyzed and each software sample in the background samples in each network layer of the target classification model; Obtaining the contribution degree of each basic block in the software to be analyzed to the target family classification result according to the family classification results and propagation difference values of each software sample in the background samples; Obtaining the family gene of the software to be analyzed according to the contribution degree of each basic block in the software to be analyzed to the target family classification result, including sorting the basic blocks in the software to be analyzed in descending order of contribution degree, and using the binary data corresponding to the preset number of basic blocks ranked at the front as the family gene of the software to be analyzed; According to the target family classification result, obtaining the family gene library of the target family and the software samples in the target family from the pre-established family gene resource center; the family classification results of the software samples in the target family are the same as the target family classification result; Determining the homology analysis result between the software to be analyzed and the target family according to the family gene of the software to be analyzed, the family gene library of the target family, and the software samples in the target family.

2. The method according to claim 1, characterized in that The attribute features of the basic block include at least one of string constant, numerical constant, number of transfer instructions, number of call instructions, number of arithmetic instructions, total number of instructions, number of child nodes, and betweenness centrality; the obtaining the contribution degree of each basic block in the software to be analyzed to the target family classification result according to the family classification results and propagation difference values of each software sample in the background samples includes: Calculating the marginal contribution value corresponding to each attribute feature of each basic block in the software to be analyzed respectively according to the family classification results and the propagation difference values of each software sample in the background samples; For each basic block of the software to be analyzed, obtaining the corresponding contribution degree according to the marginal contribution values corresponding to all attribute features of each basic block.

3. The method according to claim 2, wherein The background samples include multiple software samples; the calculating the marginal contribution value corresponding to each attribute feature of each basic block in the software to be analyzed respectively according to the family classification results and propagation difference values of each software sample in the background samples includes: Calculating the initial marginal contribution value of each attribute feature of each basic block in the software to be analyzed relative to different software samples according to the family classification results and propagation difference values of each software sample in the background samples; For each basic block in the software to be analyzed, determining the average value of the initial marginal contribution values of each attribute feature relative to different software samples as the marginal contribution value corresponding to each attribute feature.

4. The method according to claim 1, wherein Determining the homology analysis result between the software to be analyzed and the target family according to the family gene of the software to be analyzed, the family gene library of the target family, and the software samples in the target family, includes: Matching the family gene of the software to be analyzed with the software samples in the target family to obtain the number of successfully matched software; Matching the family genes in the family gene library of the target family with the family gene of the software to be analyzed to obtain the number of successfully matched genes; Obtaining the homology analysis result between the software to be analyzed and the target family according to the number of software and the number of genes.

5. The method according to any one of claims 1 to 4, characterized in that The method further includes: Obtaining the attribute control flow chart of each software sample in the software sample set; Constructing at least one sample classification model based on the attribute control flow chart; the at least one sample classification model includes the target classification model; Using each of the sample classification models to perform family classification on each software sample in the software sample set to obtain the family classification results of each software sample in the software sample set; Performing interpretive analysis on the family classification results of each software sample in the software sample set to obtain the family genes of each software sample; Constructing a family gene library for the corresponding family according to the family genes of each software sample, and the family gene libraries of multiple families constitute the family gene resource center.

6. The method according to any one of claims 1-4, characterized in that, The method further includes: Obtaining multiple pre-trained sample classification models; Using each of the sample classification models to perform family classification on the software to be analyzed and determining the target classification model from each of the sample classification models.

7. A software homology analysis device, characterized in that, The device includes: A classification result acquisition module, configured to obtain the target family classification result of the software to be analyzed by using the pre-trained target classification model; A family gene acquisition module, configured to determine background samples from the software sample set; the software sample set is used to train the target classification model; the background samples include at least one software sample in the software sample set; obtaining the family classification results of each software sample in the background samples by the target classification model; obtaining the propagation difference values of the software to be analyzed and each software sample in the background samples at each network layer of the target classification model; obtaining the contribution degree of each basic block in the software to be analyzed to the target family classification result according to the family classification results and propagation difference values of each software sample in the background samples; obtaining the family gene of the software to be analyzed according to the contribution degree of each basic block in the software to be analyzed to the target family classification result, including sorting the basic blocks in the software to be analyzed in descending order of contribution degree, and using the binary data corresponding to the preset number of basic blocks with the highest ranking as the family gene of the software to be analyzed; A homology analysis module, configured to obtain a family gene library of a target family and software samples in the target family from a pre-established family gene resource center according to the target family classification result; the family classification result of the software samples in the target family is the same as the target family classification result; and determine a homology analysis result of the software to be analyzed and the target family according to the family genes of the software to be analyzed, the family gene library of the target family, and the software samples in the target family.

8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Malware detection method based on software genetic technology

    CN108932430A

  • Malicious code classification method and device, electronic equipment and storage medium

    CN118070276A