Software homology analysis method and device, computer equipment and readable storage medium

By using pre-trained target classification model and family gene resource center, software homology analysis is automatically performed, which solves the problem of low traceability analysis efficiency in the existing technology, and achieves fast and efficient homology analysis.

CN119987788AActive Publication Date: 2025-05-13CHINA TELECOM CLOUD TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510473377.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-05-13
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

The prior art requires a lot of labor time and cost in traceability analysis, especially when processing complex samples, which are less efficient.

Method used

By obtaining the target family classification results of the pre-trained target classification model for the analysis software, the family genes of the software to be analyzed are obtained, and the family gene database and software samples of the target family are obtained from the pre-established family gene resource center are determined to determine the homology analysis results of the software to be analyzed and the target family.

Benefits of technology

It realizes rapid traceability of analysis software, shortens analysis time, improves the efficiency of homology analysis, and reduces labor costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119987788A_ABST
    Figure CN119987788A_ABST
Patent Text Reader

Abstract

The invention relates to a software homology analysis method and device, computer equipment and a readable storage medium. The method comprises the following steps: acquiring a target family classification result of to-be-analyzed software by a pre-trained target classification model, and interpreting and analyzing the target classification model based on the target family classification result, so as to acquire the influence degree of each binary code snippet in the to-be-analyzed software on an output result of the target classification model; then, according to a target family classification result, a family gene pool of a target family and a software sample in the target family are obtained from a pre-established family gene resource center, and according to the family gene of the to-be-analyzed software, the family gene pool of the target family and the software sample in the target family, the software sample in the to-be-analyzed software can be analyzed according to the family gene of the to-be-analyzed software, the family gene pool of the target family and the software sample in the target family. According to the method, the homology analysis result of the to-be-analyzed software and the target family is determined, the to-be-analyzed software is traced, the time for tracing the to-be-analyzed software is shortened, and the homology analysis efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of network security technology, and in particular to a software homology analysis method, apparatus, computer equipment, and computer-readable storage medium. Background Art

[0002] With the development of computer technology, the spread and attacks of malware have become more frequent, and it is crucial to build a network attack defense system. In building a network attack defense system, tracing the source of attack events is a very important link. In the current field of source tracing analysis, the common method is to manually compare and search binary codes to attribute and trace the source of samples. However, when facing some more complex samples, it takes a lot of time to perform repeated analysis and comparison, which will consume a lot of labor costs. Summary of the invention

[0003] Based on this, it is necessary to provide a software homology analysis method, device, computer equipment and computer-readable storage medium that can improve traceability efficiency in response to the above technical problems.

[0004] In a first aspect, the present application provides a method for software homology analysis, the method comprising:

[0005] Obtain the classification results of the target family of the software to be analyzed by the pre-trained target classification model;

[0006] Based on the target family classification result, interpret and analyze the target classification model to obtain the family gene of the software to be analyzed;

[0007] According to the target family classification result, a family gene library of the target family and software samples in the target family are obtained from a pre-established family gene resource center; the family classification result of the software samples in the target family is the same as the target family classification result;

[0008] According to the family genes of the software to be analyzed, the family gene library of the target family and the software samples in the target family, the homology analysis results of the software to be analyzed and the target family are determined.

[0009] In one embodiment, the target classification model includes multiple network layers; the step of obtaining the family gene of the software to be analyzed from the target classification model based on the target family classification result includes:

[0010] Determine a background sample from a software sample set; the software sample set is used to train the target classification model; the background sample includes at least one software sample in the software sample set;

[0011] Obtaining the family classification result of the target classification model on each software sample in the background sample;

[0012] Obtaining propagation difference values ​​of the software to be analyzed and each software sample in the background samples at each network layer of the target classification model;

[0013] According to the family classification results and propagation difference values ​​of each software sample in the background sample, obtaining the contribution of each basic block in the software to be analyzed to the target family classification result;

[0014] According to the contribution of each basic block in the software to be analyzed to the target family classification result, the family gene of the software to be analyzed is obtained.

[0015] In one embodiment, the attribute characteristics of the basic block include at least one of a string constant, a numerical constant, a transfer instruction number, a call instruction number, an arithmetic instruction number, a total instruction number, a child node number, and a betweenness centrality; and obtaining the contribution of each basic block in the software to be analyzed to the target family classification result according to the family classification result and the propagation difference value of each software sample in the background sample includes:

[0016] Calculate the marginal contribution value corresponding to each attribute feature of each basic block in the software to be analyzed according to the family classification result of each software sample in the background sample and the propagation difference value;

[0017] For each basic block of the software to be analyzed, the corresponding contribution degree is obtained according to the marginal contribution values ​​corresponding to all attribute features of each basic block.

[0018] In one embodiment, the background sample includes a plurality of software samples; and the step of calculating the marginal contribution value corresponding to each attribute feature of each basic block in the software to be analyzed according to the family classification result and the propagation difference value of each software sample in the background sample includes:

[0019] According to the family classification results and propagation difference values ​​of each software sample in the background sample, respectively calculating the initial marginal contribution value of each attribute feature of each basic block in the software to be analyzed relative to different software samples;

[0020] For each basic block in the software to be analyzed, an average value of the initial marginal contribution values ​​of each attribute feature relative to different software samples is determined as a marginal contribution value corresponding to each attribute feature.

[0021] In one embodiment, determining the homology analysis result between the software to be analyzed and the target family according to the family gene of the software to be analyzed, the family gene library of the target family and the software samples in the target family includes:

[0022] Matching the family gene of the software to be analyzed with the software samples in the target family to obtain the number of successfully matched software;

[0023] Matching the family genes in the family gene library of the target family with the family genes of the software to be analyzed to obtain the number of successfully matched genes;

[0024] According to the number of software and the number of genes, homology analysis results of the software to be analyzed and the target family are obtained.

[0025] In one embodiment, the method further comprises:

[0026] Obtaining a property control flow chart of each software sample in a software sample set;

[0027] Constructing at least one sample classification model based on the attribute control flow chart; the at least one sample classification model includes the target classification model;

[0028] Using each of the sample classification models to perform family classification on each of the software samples in the software sample set, and obtaining a family classification result for each of the software samples in the software sample set;

[0029] Explain and analyze the family classification results of each software sample in the software sample set to obtain the family gene of each software sample;

[0030] According to the family genes of each of the software samples, a family gene library of the corresponding family is constructed, and the family gene libraries of multiple families constitute the family gene resource center.

[0031] In one embodiment, the method further comprises:

[0032] Get multiple pre-trained sample classification models;

[0033] The software to be analyzed is classified into families using each of the sample classification models, and the target classification model is determined from each of the sample classification models.

[0034] In a second aspect, the present application also provides a software homology analysis device, the device comprising:

[0035] A classification result acquisition module, used to obtain the classification results of the target family of the software to be analyzed by the pre-trained target classification model;

[0036] A family gene acquisition module is used to interpret and analyze the target classification model based on the target family classification result to obtain the family gene of the software to be analyzed;

[0037] A homology analysis module is used to obtain a family gene library of a target family and software samples in the target family from a pre-established family gene resource center according to the target family classification result; the family classification result of the software samples in the target family is the same as the target family classification result; and determine the homology analysis results between the software to be analyzed and the target family according to the family genes of the software to be analyzed, the family gene library of the target family and the software samples in the target family.

[0038] In a third aspect, the present application further provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of the software homology analysis method provided in any of the above embodiments are implemented.

[0039] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the software homology analysis method provided in any of the above embodiments are implemented.

[0040] In the above-mentioned software homology analysis method, device, computer equipment and computer-readable storage medium, by obtaining the target family classification results of the software to be analyzed by a pre-trained target classification model, and based on the target family classification results, the target classification model is interpreted and analyzed, so that the degree of influence of each binary code fragment in the software to be analyzed on the output results of the target classification model can be obtained, so that the family genes of the software to be analyzed can be obtained, and then, according to the target family classification results, the family gene library of the target family and the software samples in the target family are obtained from a pre-established family gene resource center, and according to the family genes of the software to be analyzed, the family gene library of the target family and the software samples in the target family, the homology analysis results of the software to be analyzed and the target family are determined, so as to achieve the traceability of the software to be analyzed, shorten the time for tracing the software to be analyzed, and improve the efficiency of homology analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the drawings required for use in the embodiments of the present application or related technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0042] Figure 1 A schematic diagram of a software homology analysis method in one embodiment;

[0043] Figure 2A schematic diagram of a process of obtaining family genes of software to be analyzed for a target classification model based on a target family classification result in an embodiment;

[0044] Figure 3 It is a flowchart of determining the homology analysis result between the software to be analyzed and the target family according to the family gene of the software to be analyzed, the family gene library of the target family and the software samples in the target family in one embodiment;

[0045] Figure 4 A schematic diagram of a process for constructing a family gene library in one embodiment;

[0046] Figure 5 is an architecture diagram of a software homology analysis method in one embodiment;

[0047] Figure 6 is a schematic diagram of a process of software homology analysis method in another embodiment;

[0048] Figure 7 is a structural block diagram of a software homology analysis device in one embodiment;

[0049] Figure 8 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0050] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0051] In one embodiment, Figure 1 As shown, the present application provides a software homology analysis method, including steps 102 to 108.

[0052] Step 102, obtaining the classification result of the target family of the software to be analyzed by the pre-trained target classification model.

[0053] Among them, the target classification model can be a family classification model obtained by deep training using a one-dimensional convolutional neural network, such as a ResNet50 convolutional neural network, a VGG19 convolutional neural network, an Xception convolutional neural network and other neural network algorithms. The software to be analyzed is mainly malware that threatens network security or user interests, which can be a binary code fragment or a computer program product. In this application, malware is classified, and a batch of malware that is likely to be written by the same organization or person is called a "family". The target family classification result includes the family to which the software to be analyzed belongs.

[0054] If the software to be analyzed contains a shell, before using the pre-trained target classification model to classify the software to be analyzed into families, the software to be analyzed needs to be unpacked first. For example, the Unipacker tool can be used to unpack the software to be analyzed.

[0055] Step 104 , based on the target family classification result, interpret and analyze the target classification model to obtain the family gene of the software to be analyzed.

[0056] The family gene of the software to be analyzed actually refers to the binary code fragments in the software to be analyzed. The software to be analyzed can be divided into multiple continuous binary code fragments, and the target classification model can be interpreted and analyzed based on the target family classification results. The degree of influence of each binary code fragment in the software to be analyzed on the output result of the target classification model can be obtained. In this way, the binary code fragment with the greatest influence can be used as the family gene of the software to be analyzed. Exemplarily, the SHAP (SHapley Additive exPlanations) framework can be used to interpret and analyze the target classification model to obtain the family gene of the software to be analyzed.

[0057] Step 106, according to the target family classification result, obtain the family gene library of the target family and the software samples in the target family from a pre-established family gene resource center; the family classification result of the software samples in the target family is the same as the target family classification result.

[0058] The family gene resource center can be a database containing family gene libraries and software samples of different families established based on the data accumulated in the early stage. The genes in the family gene library are binary code fragments with similar behaviors in the malware of the same family. Similar behaviors can be using the same IP address or the same domain name, etc.

[0059] Step 108 , determining the homology analysis result between the software to be analyzed and the target family according to the family genes of the software to be analyzed, the family gene library of the target family and the software samples in the target family.

[0060] After obtaining the family genes of the software to be analyzed, the family gene library and software samples of the target family can be matched with the family genes of the software to be analyzed respectively to determine the homology analysis results between the software to be analyzed and the target family, thereby tracing the source of the software to be analyzed.

[0061] In an embodiment of the present application, the software to be analyzed is classified into families by a "model + gene" approach. Specifically, a target family classification result of the software to be analyzed is obtained by obtaining a pre-trained target classification model, and based on the target family classification result, the target classification model is interpreted and analyzed, so that the degree of influence of each binary code fragment in the software to be analyzed on the output result of the target classification model can be obtained, thereby obtaining the family genes of the software to be analyzed. Then, according to the target family classification result, the family gene library of the target family and the software samples in the target family are obtained from a pre-established family gene resource center, and according to the family genes of the software to be analyzed, the family gene library of the target family and the software samples in the target family, the homology analysis results of the software to be analyzed and the target family are determined, so as to achieve traceability of the software to be analyzed, shorten the time for tracing the software to be analyzed, and improve the efficiency of homology analysis.

[0062] In one embodiment, Figure 2 As shown, based on the target family classification result, the target classification model is used to obtain the family gene of the software to be analyzed, including steps 202 to 210.

[0063] Step 202: determine a background sample from a software sample set, where the software sample set is used to train a target classification model, and the background sample includes at least one software sample in the software sample set.

[0064] The software sample set includes benign software and malware of known families. Benign software obtained through official channels can be installed in batches through software installation, and binary files with suffixes such as exe, sys, and dll on the disk after the benign software is installed can be collected. Malware with family labels can be obtained by using threat intelligence sharing platforms and actual combat collection. If the obtained benign software and malware include shells, the software samples can be unpacked using the Unipacker tool respectively. A software sample set is constructed based on the obtained benign software and malware. The software sample set can be used to train a target classification model. When using the software sample set to train the target classification model, the software sample set can be divided into a training set and a test set. The training set is used to train the target classification model, and the test set can be used to verify the accuracy of the family classification results output by the target classification model. For example, the software sample set can be divided into a training set and a test set in a ratio of 8:2, and a shuffle operation is performed. After the data in the software sample set is redistributed, the target classification model is trained using the test set. At least one malware can be randomly selected from the software sample set as a background sample.

[0065] Step 204: Obtain the family classification results of each software sample in the background sample by the target classification model.

[0066] The background sample is recorded as , the software samples in the background sample are , The family classification results output by the target classification model for each software sample in the background sample are: .

[0067] Step 206, obtaining propagation difference values ​​of the software to be analyzed and each software sample in the background samples at each network layer of the target classification model.

[0068] The target classification model includes multiple network layers, illustratively including an input layer, an output layer and at least one hidden layer. In order to attribute the difference in the output results of the target classification model for the software to be analyzed and the background samples to the input features, it is necessary to first obtain the propagation difference values ​​of each software sample in the software to be analyzed and the background samples in each network layer of the target classification model.

[0069] Assume that the number of network layers of the target classification model is L, for the software to be analyzed x and the software sample , The propagation difference value of the layer for:

[0070]

[0071] in, for The activation value of the layer, It can be obtained by back-transferring the propagation difference of the next layer:

[0072]

[0073] in, express The weight of the layer, express The bias of the layer, Represents the activation function.

[0074] Step 208 , obtaining the contribution of each basic block in the software to be analyzed to the target family classification result according to the family classification result and the propagation difference value of each software sample in the background sample.

[0075] In this application, when analyzing the software to be analyzed, the control flow information of the code in the software to be analyzed is represented in the form of a directed graph, and the directed graph consists of a series of basic blocks. The basic block consists of a series of code statements. The basic block has two characteristics when it is executed. One is that the basic block can only be executed from the first jump statement, and it cannot jump into the middle of the basic block in some way. The second is that the statements in the basic block must leave from the last statement when executing, and cannot jump to other basic blocks halfway through execution. The attribute characteristics of the basic block include at least one of string constants, numerical constants, the number of transfer instructions, the number of call instructions, the number of arithmetic instructions, the total number of instructions, the number of child nodes, and the betweenness centrality. Exemplarily, the Radare2 tool can be used to parse the software to be analyzed and extract the attributed control flow graph (ACFG) of the sample to be analyzed.

[0076] When obtaining the contribution of each basic block in the software to be analyzed to the classification result of the target family, the input difference between the software to be analyzed and the background sample can be decomposed into each attribute feature of the basic block, and the marginal contribution value of each attribute feature can be calculated.

[0077] Specifically, the marginal contribution value corresponding to each attribute feature of each basic block in the software to be analyzed can be calculated based on the family classification results and propagation difference values ​​of each software sample in the background sample. The marginal contribution value It can be calculated by the following formula:

[0078]

[0079] in, Is the target classification model in the software sample Pair characteristics The partial derivative of is approximated by the following formula:

[0080]

[0081] If the background sample includes multiple software samples, then based on formula (3), according to the family classification results and propagation difference values ​​of each software sample in the background sample, the initial marginal contribution value of each attribute feature of each basic block in the software to be analyzed relative to different software samples can be calculated respectively. For each basic block in the software to be analyzed, the average of the initial marginal contribution values ​​of each attribute feature relative to different software samples is determined as the marginal contribution value corresponding to each attribute feature. Exemplarily, for an attribute feature i in a single basic block, the marginal contribution value can be obtained by the following formula: :

[0082]

[0083] in, Represents software samples in background samples The number of

[0084] After obtaining the marginal contribution value corresponding to each attribute feature of each basic block, the marginal contribution value corresponding to each attribute feature of each basic block can be stored as a numpy file in a one-to-one correspondence order between the marginal contribution value and the attribute feature of the basic block.

[0085] Furthermore, for each basic block of the software to be analyzed, the corresponding contribution is obtained according to the marginal contribution values ​​corresponding to all the attribute features of each basic block. The numpy file can be read, and the marginal contribution values ​​of the attribute features of each dimension in the same basic block are weighted and averaged to obtain the contribution corresponding to the basic block. The weights of the attribute features of each dimension in the same basic block can be set manually, and the embodiment of the present application does not limit this.

[0086] Step 210 , obtaining the family gene of the software to be analyzed according to the contribution of each basic block in the software to be analyzed to the target family classification result.

[0087] The contribution of each basic block to the target family classification result reflects the degree of influence of each basic block on the target family classification result. The basic blocks in the software to be analyzed can be sorted in descending order according to the contribution, and the binary data corresponding to the preset number of basic blocks ranked at the top are used as the family genes of the software to be analyzed. The preset number can be reasonably set according to empirical rules. For example, the preset number can be one, two, five or more.

[0088] In this embodiment, the difference in the output results of the target classification model for the software to be analyzed and the background samples is attributed to the input features through layer-by-layer reverse transmission. By obtaining the contribution of each basic block in the software to be analyzed to the target family classification result, the key basic blocks in the software to be analyzed that affect the target family classification result are obtained. The binary data corresponding to the key basic blocks are used as the family genes of the software to be analyzed, thereby realizing the automatic extraction of the family genes of the software to be analyzed.

[0089] In one embodiment, Figure 3 As shown, according to the family genes of the software to be analyzed, the family gene library of the target family and the software samples in the target family, the homology analysis results of the software to be analyzed and the target family are determined, including steps 302 to 306.

[0090] Step 302: Match the family gene of the software to be analyzed with the software samples in the target family to obtain the number of successfully matched software.

[0091] Successful matching of the family gene of the software to be analyzed with the software sample may mean that the overlap rate between the software sample and the family gene of the software to be analyzed exceeds a first preset value. The first preset value may be reasonably set based on experience.

[0092] According to the binary data after removing the position information of the family gene of the software to be analyzed, the first YARA (YetAnother Recursive Acronym) rule can be defined, and the first YARA rule is used to match the software samples of the target family to obtain the number of successfully matched software. The definition process of the first YARA rule can refer to the definition of conventional YARA rules, which will not be described in detail here.

[0093] Step 304, matching the family genes in the family gene library of the target family with the family genes of the software to be analyzed, and obtaining the number of successfully matched genes.

[0094] The successful matching of the family genes in the family gene library of the target family with the family genes of the software to be analyzed may mean that the overlap rate between the family genes in the family gene library of the target family and the family genes of the software to be analyzed exceeds a second preset value, and the second preset value may be reasonably set based on experience.

[0095] The second YARA rule can be defined based on the binary data after the position information of the family genes in the family gene library of the target family is removed, and the second YARA rule is used to perform matching in the software to be analyzed to obtain the number of genes that are successfully matched. The definition process of the second YARA rule can refer to the definition of conventional YARA rules, and will not be described in detail here.

[0096] Step 306, obtaining homology analysis results of the software to be analyzed and the target family according to the number of software and the number of genes.

[0097] Specifically, when any one of the software quantity and the gene quantity is higher than the corresponding preset value, it can be considered that the correlation between the software to be analyzed and the target family is high, and when both the software quantity and the gene quantity are lower than the corresponding preset value, it can be considered that the correlation between the software to be analyzed and the target family is low. The preset values ​​corresponding to the software quantity and the gene quantity can be reasonably set based on empirical rules, and can be 0 or 1, for example.

[0098] In this embodiment, the family genes of the software to be analyzed are matched with the software samples in the target family, and the family genes in the family gene library of the target family are matched with the family genes of the software to be analyzed. The homology analysis results of the software to be analyzed and the target family are determined by cross-validation, thereby ensuring the accuracy of the homology analysis results.

[0099] In one embodiment, Figure 4As shown, the software homology analysis method of the present application also includes steps 402 to 410.

[0100] Step 402: Obtain the attribute control flow chart of each software sample in the software sample set.

[0101] The attribute control flow graph is a directed graph composed of a series of basic blocks with attribute features. The attribute features of the basic blocks include eight features, namely, string constants, numerical constants, number of transfer instructions, number of call instructions, number of arithmetic instructions, total number of instructions, number of child nodes, and betweenness centrality. The software sample set has been described above and will not be repeated here. It should be noted that, in this embodiment, the software sample set can be used to train multiple sample classification models.

[0102] Step 404: construct at least one sample classification model based on the attribute control flow chart, wherein the at least one sample classification model includes a target classification model.

[0103] The basic blocks in multiple ACFG graphs can be arranged in a non-jump priority manner to obtain multiple ACFG feature sequences. The non-jump priority is for basic blocks with conditional jump instructions at the end. For example, if there is a conditional jump instruction at the end of basic block A, it jumps to basic block C and does not jump to basic block B. Then the position of basic block B in the ACFG feature sequence is before basic block C.

[0104] The currently commonly used one-dimensional convolutional neural network algorithm can be used to learn multiple ACFG feature sequences respectively to obtain multiple sample classification models. That is, each one-dimensional convolutional neural network algorithm corresponds to a sample classification model.

[0105] Step 406 , using each sample classification model to perform family classification on each software sample in the software sample set, and obtaining a family classification result of each software sample in the software sample set.

[0106] Step 408: interpret and analyze the family classification results of each software sample in the software sample set to obtain the family gene of each software sample.

[0107] Specifically, the classification effect of each sample classification model on the family classification of each software sample in the software sample set can be counted first. Exemplarily, the F1 score (F1 Score) can be used to measure the classification effect of each sample classification model on the same family. The F1 score is an indicator used in statistics to measure the accuracy of a binary classification model. It takes into account both the precision and recall of the classification model. For the same family, the higher the F1 score of the sample classification model, the better the classification effect. When interpreting and analyzing the family classification results of each software sample in the software sample set, it is only necessary to interpret and analyze the sample classification model with the best classification effect based on the same principles as steps 202-210 to obtain the family genes of each software sample.

[0108] Step 410: construct a family gene library of the corresponding family according to the family gene of each software sample. The family gene libraries of multiple families constitute a family gene resource center.

[0109] It can be understood that the family gene library of the family can be constructed based on the family genes of the software samples belonging to the same family in each software sample. Then, the family gene resource center can be constructed based on the family gene libraries of different families.

[0110] In this embodiment, by interpreting and analyzing the family classification results of each software sample in the software sample set, the family gene of each software sample is obtained, and the family gene library of the corresponding family is constructed according to the family gene of each software sample, so that the family gene library is quickly constructed without the need to compare samples of the malware family one by one. In addition, when training the sample classification model, the software training set used not only includes malware but also benign software. Including benign software in the consideration scope of the sample classification model can reduce the false alarm rate when matching the family gene of the software to be analyzed extracted based on the sample classification model with the family gene library and software samples of the target family.

[0111] In one embodiment, the software homology analysis method of the present application also includes the steps of obtaining multiple pre-trained sample classification models, using each sample classification model to classify the software to be analyzed into families, and determining a target classification model from each sample classification model.

[0112] It can be understood that the target classification model is the sample classification model with the best classification effect on the software to be analyzed among multiple sample classification models. For example, model A classifies the software to be analyzed into family X, and model B classifies the software to be analyzed into family Y. The classification F1 score of model A for family X is 0.75, and the classification F1 score of model B for family Y is 0.91, then model B will be selected as the target classification model for the software to be analyzed.

[0113] For a better understanding, Figure 5 and Figure 6 As shown, a more specific embodiment is listed to illustrate the software homology analysis method of the present application.

[0114] Step 602: Unpacking the software sample set. The software sample set includes benign software and malware of known families.

[0115] Step 604: extract the ACFG features of each basic block in the software sample.

[0116] Step 606, arrange the basic blocks in each software sample in a non-jump priority manner to obtain multiple ACFG feature sequences.

[0117] Step 608: Use multiple ACFG feature sequences to train sample classification models of different convolutional neural network architectures.

[0118] Step 610 , statistically analyze the classification effects of multiple sample classification models to determine the sample classification model with the best classification effect for each family.

[0119] Step 612: for the software to be explained, explain and analyze the key basic blocks that affect the family classification result of the sample classification model. The software to be explained is one of the software samples that have not been explained and analyzed in the software sample set and the sample to be analyzed.

[0120] Step 614: Use the binary data corresponding to the key basic block as the family gene of the software to be interpreted.

[0121] Step 616: when the software to be interpreted is a software sample in the software sample set that has not been interpreted and analyzed, a corresponding family gene library is constructed according to the family genes of the software to be interpreted.

[0122] Step 618, when the software to be interpreted is the software to be analyzed, the extracted family genes of the software to be analyzed are converted into the first YARA rules and matched with the genes in the family gene library of the target family, and the genes in the family gene library of the target family are converted into the second YARA rules and matched with the family genes of the software to be analyzed to determine the homology analysis results.

[0123] It should be understood that, although the steps in the flowcharts involved in the above embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps is not strictly limited in order, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.

[0124] Based on the same inventive concept, the embodiment of the present application also provides a software homology analysis device for implementing the software homology analysis method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more software homology analysis device embodiments provided below can refer to the limitations of the software homology analysis method above, and will not be repeated here.

[0125] In an exemplary embodiment, Figure 7 As shown, a software homology analysis device is provided, including: a classification result acquisition module 702, a family gene acquisition module 704 and a homology analysis module 706, wherein:

[0126] The classification result acquisition module 702 is used to obtain the classification result of the target family of the software to be analyzed by the pre-trained target classification model;

[0127] The family gene acquisition module 704 is used to interpret and analyze the target classification model based on the target family classification result, and obtain the family gene of the software to be analyzed;

[0128] The homology analysis module 706 is used to obtain the family gene library of the target family and the software samples in the target family from a pre-established family gene resource center according to the target family classification results; the family classification results of the software samples in the target family are the same as the target family classification results; based on the family genes of the software to be analyzed, the family gene library of the target family and the software samples in the target family, determine the homology analysis results between the software to be analyzed and the target family.

[0129] In one embodiment, the family gene acquisition module is also used to determine background samples from the software sample set; the software sample set is used to train the target classification model; the background sample includes at least one software sample in the software sample set; the family classification results of the target classification model for each software sample in the background sample are obtained; the propagation difference values ​​of the software to be analyzed and each software sample in the background sample in each network layer of the target classification model are obtained; based on the family classification results and propagation difference values ​​of each software sample in the background sample, the contribution of each basic block in the software to be analyzed to the target family classification result is obtained; based on the contribution of each basic block in the software to be analyzed to the target family classification result, the family gene of the software to be analyzed is obtained.

[0130] In one embodiment, the family gene acquisition module is also used to calculate the marginal contribution value corresponding to each attribute feature of each basic block in the software to be analyzed based on the family classification results and propagation difference values ​​of each software sample in the background sample; for each basic block of the software to be analyzed, the corresponding contribution degree is obtained based on the marginal contribution values ​​corresponding to all attribute features of each basic block.

[0131] In one embodiment, the family gene acquisition module is also used to calculate the initial marginal contribution value of each attribute feature of each basic block in the software to be analyzed relative to different software samples based on the family classification results and propagation difference values ​​of each software sample in the background sample; for each basic block in the software to be analyzed, the average value of the initial marginal contribution value of each attribute feature relative to different software samples is determined as the marginal contribution value corresponding to each attribute feature.

[0132] In one embodiment, the homology analysis module is also used to match the family genes of the software to be analyzed with the software samples in the target family to obtain the number of successfully matched software; match the family genes in the family gene library of the target family with the family genes of the software to be analyzed to obtain the number of successfully matched genes; and obtain the homology analysis results of the software to be analyzed and the target family based on the number of software and the number of genes.

[0133] In one embodiment, the software homology analysis device further includes a model training module and a gene library construction module.

[0134] Among them, the model training module is used to obtain the attribute control flow chart of each software sample in the software sample set; build at least one sample classification model based on the attribute control flow chart; at least one sample classification model includes a target classification model.

[0135] The gene library construction module is used to use each sample classification model to perform family classification on each software sample in the software sample set, and obtain the family classification results of each software sample in the software sample set; interpret and analyze the family classification results of each software sample in the software sample set to obtain the family genes of each software sample; according to the family genes of each software sample, construct the family gene library of the corresponding family, and the family gene libraries of multiple families constitute the family gene resource center.

[0136] In one embodiment, the software homology analysis device further includes a target model determination module for obtaining a plurality of pre-trained sample classification models; performing family classification on the software to be analyzed using each sample classification model, and determining a target classification model from each sample classification model.

[0137] Each module in the above-mentioned software homology analysis device can be implemented in whole or in part by software, hardware and a combination thereof. Each of the above-mentioned modules can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.

[0138] In an exemplary embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in FIG. Figure 8 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. Among them, the processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store extracted family genes, etc. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a software homology analysis method is implemented.

[0139] Those skilled in the art will understand that Figure 8 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0140] In an exemplary embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the software homology analysis method provided in any of the above embodiments is implemented.

[0141] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the software homology analysis method provided in any of the above embodiments is implemented.

[0142] In one embodiment, a computer program product is provided, including a computer program, which, when executed by a processor, implements the software homology analysis method provided in any of the above embodiments.

[0143] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., but are not limited to this.

[0144] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0145] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A software homology analysis method, characterized in that: The method comprises: Obtain the classification results of the target family of the software to be analyzed by the pre-trained target classification model; Based on the target family classification result, interpret and analyze the target classification model to obtain the family gene of the software to be analyzed; According to the target family classification result, a family gene library of the target family and software samples in the target family are obtained from a pre-established family gene resource center; the family classification result of the software samples in the target family is the same as the target family classification result; According to the family genes of the software to be analyzed, the family gene library of the target family and the software samples in the target family, the homology analysis results of the software to be analyzed and the target family are determined.

2. The method according to claim 1, characterized in that The target classification model includes multiple network layers; based on the target family classification result, obtaining the family gene of the software to be analyzed from the target classification model includes: Determine a background sample from a software sample set; the software sample set is used to train the target classification model; the background sample includes at least one software sample in the software sample set; Obtaining the family classification result of the target classification model on each software sample in the background sample; Obtaining propagation difference values ​​of the software to be analyzed and each software sample in the background samples at each network layer of the target classification model; According to the family classification results and propagation difference values ​​of each software sample in the background sample, obtaining the contribution of each basic block in the software to be analyzed to the target family classification result; According to the contribution of each basic block in the software to be analyzed to the target family classification result, the family gene of the software to be analyzed is obtained.

3. The method according to claim 2, characterized in that The attribute features of the basic block include at least one of a string constant, a numerical constant, a transfer instruction number, a call instruction number, an arithmetic instruction number, a total instruction number, a child node number, and a betweenness centrality; the step of obtaining the contribution of each basic block in the software to be analyzed to the target family classification result based on the family classification result and the propagation difference value of each software sample in the background sample includes: Calculate the marginal contribution value corresponding to each attribute feature of each basic block in the software to be analyzed according to the family classification result of each software sample in the background sample and the propagation difference value; For each basic block of the software to be analyzed, the corresponding contribution degree is obtained according to the marginal contribution values ​​corresponding to all attribute features of each basic block.

4. The method according to claim 3, characterized in that The background sample includes a plurality of software samples; and the marginal contribution value corresponding to each attribute feature of each basic block in the software to be analyzed is calculated according to the family classification result and the propagation difference value of each software sample in the background sample, including: According to the family classification results and propagation difference values ​​of each software sample in the background sample, respectively calculating the initial marginal contribution value of each attribute feature of each basic block in the software to be analyzed relative to different software samples; For each basic block in the software to be analyzed, an average value of the initial marginal contribution values ​​of each attribute feature relative to different software samples is determined as a marginal contribution value corresponding to each attribute feature.

5. The method according to claim 1, characterized in that: The step of determining the homology analysis result between the software to be analyzed and the target family according to the family gene of the software to be analyzed, the family gene library of the target family and the software samples in the target family includes: Matching the family gene of the software to be analyzed with the software samples in the target family to obtain the number of successfully matched software; Matching the family genes in the family gene library of the target family with the family genes of the software to be analyzed to obtain the number of successfully matched genes; According to the number of software and the number of genes, homology analysis results of the software to be analyzed and the target family are obtained.

6. The method according to any one of claims 1 to 5, characterized in that: The method further comprises: Obtaining a property control flow chart of each software sample in a software sample set; Constructing at least one sample classification model based on the attribute control flow chart; the at least one sample classification model includes the target classification model; Using each of the sample classification models to perform family classification on each of the software samples in the software sample set, and obtaining a family classification result for each of the software samples in the software sample set; Explain and analyze the family classification results of each software sample in the software sample set to obtain the family gene of each software sample; According to the family genes of each of the software samples, a family gene library of the corresponding family is constructed, and the family gene libraries of multiple families constitute the family gene resource center.

7. The method according to any one of claims 1 to 5, characterized in that: The method further comprises: Get multiple pre-trained sample classification models; The software to be analyzed is classified into families using each of the sample classification models, and the target classification model is determined from each of the sample classification models.

8. A software homology analysis device, characterized in that: The device comprises: A classification result acquisition module, used to obtain the classification results of the target family of the software to be analyzed by the pre-trained target classification model; A family gene acquisition module is used to interpret and analyze the target classification model based on the target family classification result to obtain the family gene of the software to be analyzed; A homology analysis module is used to obtain a family gene library of a target family and software samples in the target family from a pre-established family gene resource center according to the target family classification result; the family classification result of the software samples in the target family is the same as the target family classification result; and determine the homology analysis results between the software to be analyzed and the target family according to the family genes of the software to be analyzed, the family gene library of the target family and the software samples in the target family.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Malware detection method based on software genetic technology

    CN108932430A

  • Malicious code family identification method and device

    CN110414234A

  • Malicious software family classification avoidance method based on deep reinforcement learning

    CN111552971A

  • Malicious software analysis method and device, storage medium and equipment

    CN117171738A

  • Malicious code classification method and device, electronic equipment and storage medium

    CN118070276A