Multi-level feature fusion binary code similarity detection method and system

Through the multi-level feature fusion method, combined with the characteristics of control flow diagrams and pseudo-code, the problem of insufficient analysis of different levels of binary code in the prior art is solved, and the accuracy and robustness of binary code similarity detection are improved.

CN120163140AActive Publication Date: 2025-06-17CHONGQING UNIV OF POSTS & TELECOMM

Patent Information

Application Number
CN202510232580.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-17
Estimated Expiration
2045-02-28

AI Technical Summary

Technical Problem

The existing technology lacks comprehensive analysis of different levels of binary code in cross-architecture and cross-compiler scenarios, and deep learning networks have shortcomings in capturing the global structural features of the code.

Method used

A binary code similarity detection method with multi-level feature fusion is used to obtain the code representation of the binary file through disassembly, combining the control flow diagram and pseudo-code, and using the feature fusion module to integrate the features of assembly instructions and pseudo-code to calculate the final similarity score.

Benefits of technology

It improves the robustness and detection accuracy of binary code similarity detection, fully captures the global structural characteristics of the code, and enhances the detection capabilities in cross-architecture and cross-compiler scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163140A_ABST
    Figure CN120163140A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of computer security, and particularly relates to a multi-level feature fusion binary code similarity detection method and system.The method comprises the steps that two binary files to be detected are obtained, the two binary files to be detected are disassembled respectively, and code representations of the two binary files are obtained; code representations of the two binary files are preprocessed and then input into the trained binary code similarity detection model, and a similarity score is obtained; the binary code similarity detection model comprises a CFG branch, a pseudo code branch, a structural feature extraction module and a feature fusion module; according to the method, characteristics of assembly instructions and pseudo codes are respectively extracted through CFG and pseudo code branches to capture bottom-layer and high-layer information, the similarity between the assembly instructions and the similarity between the pseudo codes are calculated according to the characteristics, and different levels of information are fused by fusing the similarity between the assembly instructions and the similarity between the pseudo codes. And the robustness and the detection precision of binary code similarity detection are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer security, and particularly relates to a binary code similarity detection method and system with multi-level feature fusion. Background Art

[0002] Binary code similarity detection refers to analyzing and comparing the similarities between different binary files to help identify potential code duplication, malicious tampering, or unauthorized modifications. Due to the accelerating speed of software development, developers usually prioritize implementing the core functions of products and bringing them to market as soon as possible, while paying relatively less attention to software security.

[0003] Traditional binary code similarity detection methods mainly rely on manual features. For example, the opcode sequence in assembly instructions is used to analyze the low-level structure of the code, or the homology of functions is judged by analyzing the function call graph (CG) and the control flow relationship between basic blocks. Deep learning-based methods automatically extract code features by learning higher-level representations, especially performing well when dealing with complex code representations. For example, some methods use graph neural networks (GNNs) to model the control flow graph (CFG) and the advanced control flow graph (ACFG) of the code, such as methods like Geinus, Gemini, and Vulseeker.

[0004] Although the learning-based BCSD methods have better performance than traditional methods. However, there are still the following problems in scenarios such as cross-architecture and cross-compiler:

[0005] (1) Most of the existing technologies rely on a single representation form of binary code and lack comprehensive analysis of different levels of code features. Among them, assembly instructions mainly reflect the underlying structural features of the code, while pseudocode is more inclined to high-level abstract semantics. Therefore, how to effectively fuse multi-level features is a problem worthy of in-depth study.

[0006] (2) Although deep learning networks can automatically extract high-level semantic features of the code, the deep features learned by them may not be able to fully capture the global structural features of the code in some cases. Summary of the Invention

[0007] To solve the above-mentioned existing technologies, on the one hand, the present invention adopts a binary code similarity detection method with multi-level feature fusion, including: obtaining two binary files to be detected, respectively disassembling the two binary files to be detected to obtain the code representations of the two binary files, preprocessing the code representations of the two binary files and then inputting them into a trained binary code similarity detection model to obtain a similarity score; the binary code similarity detection model includes: a CFG branch, a pseudocode branch, a structural feature extraction module, and a feature fusion module; the training process of the binary code similarity detection model includes:

[0008] S1: Obtain two binary files, respectively disassemble the two binary files to obtain the code representations of the two binary files; wherein, the code representation includes a control flow graph CFG n and pseudocode, the nodes of the control flow graph represent basic blocks, the edges represent the relationships between basic blocks, and each basic block contains multiple assembly instructions;

[0009] S2: Respectively preprocess the control flow graphs CFG n of the two binary files to obtain the preprocessed control flow graphs CFG″ n of the two binary files; wherein, n is the index of the binary file;

[0010] S3: Input the preprocessed control flow graphs CFG″ n of the two binary files into the CFG branch to obtain the similarity X between the assembly instructions of the two binary files;

[0011] S4: Input the pseudocodes of the two binary files into the pseudocode branch to obtain the similarity Y between the pseudocodes of the two binary files;

[0012] S5: Input the preprocessed control flow graphs CFG″ n and the pseudocode representations of the two binary files into the structural feature extraction module to obtain the structural features F n of the preprocessed control flow graphs CFG″ CFG,n and the pseudocode representations of the two binary files, P,n F;

[0013] S6: Input the structural features F n of the preprocessed control flow graphs CFG″ CFG,n and the pseudocode representations of the two binary files, F P,n and the similarity X and Y into the feature fusion module to obtain the final similarity score;

[0014] S7: Calculate the loss function value according to the final similarity score, update the model parameters according to the loss function value, and when the loss function value is the smallest, obtain the trained binary code similarity detection model.

[0015] On the other hand, the present invention adopts a binary code similarity detection system with multi-level feature fusion, which is used to execute the above-mentioned binary code similarity detection method with multi-level feature fusion, including:

[0016] A data input module, which is used to obtain two binary files to be detected, and use a disassembly tool to disassemble the two binary files to be detected, so as to obtain the code representation of each binary file; the code representation includes a control flow graph and pseudocode, the nodes of the control flow graph represent basic blocks, the edges represent the relationships between basic blocks, and each basic block contains multiple assembly instructions;

[0017] A data preprocessing module, which is used to preprocess the code representations of the two binary files to obtain a preprocessed control flow graph CFG″ n ;

[0018] An assembly similarity calculation module, which is used to calculate the similarity X between the assembly instructions of the two binary files according to the preprocessed control flow graph CFG″ n ;

[0019] A pseudocode similarity calculation module, which is used to calculate the similarity Y between the pseudocodes of the two binary files;

[0020] A structural feature extraction module, which is used to extract the structural features of the preprocessed control flow graph and pseudocode of the two binary files;

[0021] A feature fusion module, which is used to calculate the final similarity according to the structural features of the control flow graph and pseudocode and the similarities X and Y.

[0022] Beneficial effects:

[0023] 1. The present invention captures low-level information and high-level information by extracting the features of assembly instructions and pseudocode, calculates the similarity X between assembly instructions and the similarity Y between pseudocodes respectively according to the features of assembly instructions and pseudocode, and fully fuses information at different levels by deeply fusing the similarity X between assembly instructions and the similarity Y between pseudocodes, improving the robustness and detection accuracy of binary code similarity detection; 2. On the basis of obtaining the similarities in the CFG branch and the pseudocode branch, the present invention incorporates structural features, fully capturing the global structural features of the code, thereby improving the detection accuracy. Description of the drawings

[0024] Figure 1 is a flowchart of a binary code similarity detection method with multi-level feature fusion provided by an embodiment of the present invention;

[0025] Figure 2 is a preprocessing flowchart provided by an embodiment of the present invention;

[0026] Figure 3 The structure diagram of the binary code similarity detection system provided by the embodiment of the present invention;

[0027] Figure 4 The schematic diagram of the data input module of the binary code similarity detection system according to the embodiment of the present invention;

[0028] Figure 5 The schematic diagram of the data preprocessing module of the binary code similarity detection system according to the embodiment of the present invention;

[0029] Figure 6 The schematic diagram of the assembly similarity calculation module of the binary code similarity detection system according to the embodiment of the present invention;

[0030] Figure 7 The schematic diagram of the pseudocode similarity calculation module of the binary code similarity detection system according to the embodiment of the present invention;

[0031] Figure 8 The schematic diagram of the structure feature extraction module of the binary code similarity detection system according to the embodiment of the present invention;

[0032] Figure 9 The schematic diagram of the feature fusion module of the binary code similarity detection system according to the embodiment of the present invention. Specific implementation manners

[0033] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0034] On the one hand, the present invention adopts a multi-level feature fusion binary code similarity detection method, as Figure 1 shown, including: obtaining two binary files to be detected, disassembling the two binary files to be detected respectively by using a disassembler to obtain the code representations of each binary file, preprocessing the code representations of the two binary files and inputting them into a trained binary code similarity detection model to obtain a similarity score; the binary code similarity detection model includes: a CFG branch, a pseudocode branch, a structure feature extraction module, and a feature fusion module; the training process of the binary code similarity detection model includes:

[0035] S1: Obtain two binary files, disassemble the two binary files respectively to obtain the code representations of the two binary files; wherein, the code representation includes a control flow graph CFG nAnd the pseudo-code, the nodes of the control flow graph represent basic blocks, and the edges represent the relationships between basic blocks. Each basic block contains multiple assembly instructions;

[0036] Disassembling the two binary files respectively includes:

[0037] Converting the binary file into a sequence of assembly instructions through a disassembling tool;

[0038] Generating a control flow graph by analyzing the jump instructions and call instructions in the sequence of assembly instructions; among them, each node of the control flow graph represents a basic block, and the edges represent the control flow relationships between basic blocks;

[0039] Converting the assembly instructions into pseudo-code through a decompiling tool. The pseudo-code retains the logical structure of the source code, facilitating subsequent extraction of high-level semantic features.

[0040] S2: For the control flow graphs CFG of the two binary files respectively n Perform preprocessing to obtain the preprocessed control flow graphs CFG″ of the two binary files n ; where n is the index of the binary file;

[0041] As Figure 2 shown, performing preprocessing on the control flow graph CFG n includes:

[0042] S21. Normalize the assembly instructions of each basic block of the control flow graph CFG n to obtain the normalized control flow graph CFG′ n ;

[0043] S22. Set a threshold A, and perform filling or truncating processing on the control flow graph according to the threshold A to obtain a control flow graph containing A basic blocks

[0044] Specifically, performing filling or truncating processing on the control flow graph includes: Judging whether the number of basic blocks of the control flow graph CFG′ n exceeds the threshold A. If so, truncate the first A basic blocks in the control flow graph CFG′ n to obtain a control flow graph containing A basic blocks Otherwise, perform filling and complementing on the control flow graph CFG′ n to obtain a control flow graph containing A basic blocks

[0045] S23. Perform filling or truncating processing on each basic block of the control flow graph according to the threshold A to obtain the preprocessed control flow graph CFG″ n; Among them, the control flow graph CFG″ n Each basic block of contains A assembly instructions.

[0046] Specifically, the process of padding or truncating each basic block of the control flow graph includes: judging whether the number of assembly instructions of the basic block of the control flow graph exceeds the threshold A. If so, truncate the first A assembly instructions of the basic block to obtain a basic block containing A assembly instructions; otherwise, pad and complete the assembly instructions of the basic block to obtain a basic block containing A assembly instructions.

[0047] Preferably, the threshold A is 20.

[0048] Normalizing the assembly instructions of each basic block includes:

[0049] All immediate numbers will be uniformly replaced by "const". For example, the instruction "sub rsp,07h" will be converted to "sub rsp,const"; for operators such as addition, subtraction, and multiplication in the operands, they remain unchanged. For example, "mov [rbp - 08h],rax" will become "mov [rbp - const],rax"; for the operand after the "call" instruction, it is replaced by a unified placeholder "foo". For example, "call sub_401E85" will be represented as "call foo"; the offsets in the operands, such as "offset", will be retained and used as normal, and for those cases where the operands contain "var" or "arg", which are usually related to pointers, such symbols will be replaced by "ptr". For example, "test [ebp + arg_0],1" will be converted to "test [ebp + ptr],const"; in addition, other symbols will be uniformly represented by "tag". These rules help to eliminate architecture and implementation differences and provide standardized input for subsequent similarity detection.

[0050] The control flow graph CFG″ n includes the adjacency matrix between basic blocks; when there is an edge between basic blocks, the corresponding position in the adjacency matrix is 1, otherwise it is 0.

[0051] S3: Input the pre - processed control flow graph CFG″ of the two binary files n into the CFG branch to obtain the similarity X between the assembly instructions of the two binary files;

[0052] The CFG branch includes a semantic embedding module and a function - level embedding module; the CFG branch processes the control flow graph CFG″ of the two binary files n including: respectively inputting the control flow graph CFG″ of the two binary files nInput semantic embedding module to obtain the control flow graph CFG″ of two binary files n The high-dimensional embedding matrix H of each basic block n,i ; respectively input the high-dimensional embedding matrix H of each basic block of the control flow graph CFG″ of the two binary files n and the adjacency matrix into the function-level embedding module to obtain the feature vectors of the assembly instructions of the two binary files, and calculate the similarity X between the assembly instructions of the two binary files according to the feature vectors of the assembly instructions of the two binary files; where i is the index of the basic block. n,i The process of processing by the semantic embedding module of the CFG branch includes:

[0053] Each assembly instruction carries low-level semantic information, which directly reflects the execution logic of the code. It is crucial for the model to convert the assembly instructions into high-dimensional embedding vectors to understand the semantics of the basic blocks.

[0054] Specifically, input each basic block of the control flow graph CFG″ of each binary file into the Word2Vec model to obtain the high-dimensional embedding vectors of each basic block of the control flow graph CFG″ of each binary file

[0055] ; n The formula for embedding each instruction can be expressed as: n The high-dimensional embedding vectors of each basic block

[0056] Each instruction embedding formula can be expressed as:

[0057] h n,i,j = W emb (instr n,i,j )

[0058] where h n,i,j represents the high-dimensional embedding vector of the j-th assembly instruction in the i-th basic block of the n-th binary file, W emb is the pre-trained word vector matrix, and instr n,i,j is the j-th instruction in the basic block i of the n-th binary file.

[0059] Optionally, the Word2Vec model is a pre-trained Continuous Bag-of-Words (CBOW) model.

[0060] Combine the high-dimensional embedding vectors h n,i,j of all instructions in the basic block i to obtain the high-dimensional embedding matrix where L i is the number of instructions in the basic block, and d is the dimension of the word vector. This vectorized representation lays the foundation for the subsequent intra-block and inter-block learning of the model.

[0061] The goal of function-level embedding is to aggregate the basic block features within each function into an overall function feature representation, providing an input for similarity detection. Specifically, the function-level embedding module includes: Bi-GRU, GCN, and the Bi-GRU with an attention module. The function-level embedding module in the CFG branch processes the high-dimensional embedding matrix H n of each basic block in the control flow graph CFG″ n,i and the adjacency matrix, including:

[0062] Respectively input the high-dimensional embedding matrix H of each basic block n,i into the Bi-GRU network to learn the global dependencies within each basic block, obtaining the intra-block global dependency features H' of each basic block n,i ; Input the intra-block global dependency features H' of each basic block n,i and the adjacency matrix of the control flow graph CFG″ n into the GCN to learn the inter-relationships between the basic blocks of the control flow graph CFG″ n and obtain the inter-block features H″ of each basic block n,i ; Input the inter-block features H″ of each basic block n,i into the Bi-GRU with an attention module to obtain the feature vectors of the assembly instructions of the binary file. Among them, Bi-GRU is a bidirectional gated recurrent unit, and GCN is a graph convolutional network.

[0063] Through the combination of Bi-GRU and the attention mechanism, the model can not only capture the dependencies between basic blocks within a function but also automatically identify the most representative basic blocks in the function. This mechanism enables the model to more effectively process cross-architecture function representations, improving the accuracy and robustness of similarity calculation.

[0064] Calculating the similarity X between the assembly instructions of two binary files based on the feature vectors of the assembly instructions of the two binary files includes: concatenating the two feature vectors, then passing through a 2-layer fully connected network to achieve feature dimensionality reduction and enhanced classification, and finally outputting the similarity score through the sigmoid activation function. The output similarity score ranges from 0 to 1, where the closer to 1, the more similar, and the closer to 0, the less similar.

[0065] S4: Input the pseudocode of the two binary files into the pseudocode branch to obtain the similarity Y between the pseudocodes of the two binary files;

[0066] The pseudocode branch processes the pseudocode of each binary file, including:

[0067] The pseudocodes of two binary files are respectively input into the Word2Vec model to obtain the embedding vectors of the pseudocodes of the two binary files. Specifically, after each pseudocode undergoes word embedding, it is converted into a vector representation of dimension d, and the specific function can be expressed as:

[0068] e = W word2vec ·p

[0069] where e represents the embedding vector of the pseudocode, p represents the pseudocode, and W word2vec is the pre-trained word vector matrix.

[0070] The embedding vectors of the pseudocodes of two binary files are respectively input into the R-CNN model to obtain the feature vectors of the pseudocodes of the two binary files. Among them, the full name of R-CNN is Region-CNN (Region Convolutional Neural Network).

[0071] R-CNN can not only capture the local context dependencies in the pseudocode but also capture the long-distance dependencies in the sequence through its recursive structure. The formula of R-CNN is as follows:

[0072] h t = ReLU(W conv ·[e t-m ,…,e t ,…,e t+m +b conv )

[0073] where h t is the hidden state at time step t, e t represents the embedding vector of the pseudocode at time step t, e t-m represents the embedding vector of the pseudocode at time step t - m, W conv is the convolutional kernel, and b conv is the bias term.

[0074] Calculate the similarity Y between the pseudocodes of two binary files according to the feature vectors of the pseudocodes of two binary files. Here, the steps of similarity calculation are consistent with the CFG branches.

[0075] S5: Input the preprocessed control flow graphs CFG″ n and the pseudocode representations of two binary files into the structural feature extraction module to obtain the structural features F n of the control flow graphs CFG″ CFG,n and the pseudocode representations of two binary files, F P,n ;

[0076] The structural feature extraction module extracts the structural features F n of the control flow graphs CFG″ CFG,n and the pseudocode, F P,nIncluding:

[0077] Analyze CFG″ n and the structural features in the pseudocode, and extract a series of statistical information. Specifically, 11 structural features are extracted from CFG″ n and 8 structural features are extracted from the pseudocode. As shown in Table 1, these features help to capture the execution logic and structural similarity of the code.

[0078] Table 1 Structural Feature Table

[0079]

[0080]

[0081] S6: Input the control flow graphs CFG n ″ of the two binary files and the structural features F CFG,n and F P,n as well as the similarities X and Y into the feature fusion module to obtain the final similarity score;

[0082] Concatenate the structural features F CFG,n and F P,n as well as the similarities X and Y to obtain the concatenated features; input the concatenated features into a multi-layer perceptron (MLP) to calculate the final similarity score s. The MLP consists of multiple hidden layers with non-linear activation functions, which enables it to learn the complex relationship between the input features and the final similarity.

[0083] The final similarity score s represents the overall similarity between the two binary codes being compared, which takes into account the features and similarities from the assembly instruction branches and the pseudocode branches.

[0084] S7: Calculate the loss function value according to the final similarity score, update the model parameters according to the loss function value, and when the loss function value is the smallest, obtain the trained binary code similarity detection model.

[0085] The loss function L is:

[0086] L = -(ylog(p))+(1 - y)log(1 - p)

[0087] where y is the true similarity label (0 or 1) and p is the predicted value of the model.

[0088] On the other hand, as Figure 3 shown, the present invention adopts a binary code similarity detection system based on multi-level feature fusion, which is used to execute the above-mentioned binary code similarity detection method based on multi-level feature fusion, including:

[0089] A data input module, which is used to obtain two binary files to be detected, and use a disassembly tool to disassemble the two binary files to be detected, so as to obtain the code representation of each binary file; the code representation includes a control flow graph and pseudocode, the nodes of the control flow graph represent basic blocks, the edges represent the relationships between basic blocks, and each basic block contains multiple assembly instructions;

[0090] As Figure 4 shown, the data input module includes a disassembly unit, a control flow graph generation unit, and a pseudocode generation unit;

[0091] The disassembly unit converts the binary code into an assembly instruction sequence through a disassembly tool, providing basic data for subsequent control flow graph generation and pseudocode generation.

[0092] The control flow graph generation unit generates a control flow graph by analyzing jump instructions and call instructions in the assembly instruction sequence; among them, each node of the control flow graph represents a basic block, and the edges represent the control flow relationships between basic blocks.

[0093] The pseudocode generation unit converts the assembly instructions into pseudocode through a decompilation tool. The pseudocode retains the logical structure of the source code, facilitating subsequent high-level semantic feature extraction.

[0094] A data preprocessing module, which is used to preprocess the code representations of the two binary files for subsequent semantic feature extraction;

[0095] As Figure 5 shown, the data preprocessing module includes a control flow graph preprocessing unit and a result output unit;

[0096] The control flow graph preprocessing unit is used to preprocess the control flow graph CFG n to obtain the preprocessed control flow graph CFG″ n ; where n is the index of the binary file. The preprocessing process of the control flow graph preprocessing unit is step S2 of the binary code similarity detection method.

[0097] The result output unit is used to output the preprocessed control flow graph CFG″ n for use by subsequent modules.

[0098] An assembly similarity calculation module, which is used to calculate the similarity X between the assembly instructions of the two binary files according to the preprocessed control flow graph CFG″ n ;

[0099] As Figure 6 shown, the assembly similarity calculation module includes a word embedding unit, a function-level embedding unit, and a similarity calculation unit:

[0100] The word embedding unit embeds the control flow graphs CFG″ of two binary files into high-dimensional vector representations through a pre-trained Word2Vec model, obtaining the high-dimensional embedding matrices H of each basic block of the control flow graphs CFG″ of the two binary files. n of the two binary files. n for each basic block. n,i .

[0101] The function-level embedding unit performs function-level embedding on the high-dimensional embedding matrices H of each basic block of the control flow graph CFG″ of the binary file through Bi-GRU, GCN, and Bi-GRU with an attention module, obtaining the feature vectors of the assembly instructions of the two binary files. n for each basic block of the control flow graph CFG″ of the binary file. n,i of the two binary files.

[0102] The similarity calculation unit calculates the similarity X between the assembly instructions of the two binary files according to the feature vectors of the assembly instructions of the two binary files.

[0103] A pseudocode similarity calculation module for calculating the similarity Y between the pseudocodes of two binary files;

[0104] As Figure 7 shown, the pseudocode similarity calculation module includes a word embedding unit, a function embedding unit, and a similarity calculation unit:

[0105] The word embedding unit embeds the pseudocode into a high-dimensional vector representation through a pre-trained Word2Vec model, obtaining the embedding vectors of the pseudocode, providing input for subsequent function embedding.

[0106] The function embedding unit embeds the pseudocode into a function-level vector representation through a recursive convolutional neural network (R-CNN), obtaining the feature vectors of the pseudocode.

[0107] The similarity calculation unit calculates the similarity Y between the pseudocodes of the two binary files according to the feature vectors of the pseudocodes of the two binary files.

[0108] A structural feature extraction module for extracting the structural features of the control flow graph and pseudocode of the binary file;

[0109] As Figure 8 shown, this module includes a control flow graph structural feature extraction unit, a pseudocode structural feature extraction unit, and a result output unit:

[0110] The control flow graph structural feature extraction unit is used to extract structural features from the control flow graph, such as the number of loops, the number of strongly connected components, the average in-degree and out-degree of nodes, etc. These features can reflect the structure of the control flow graph and provide more abundant information for similarity detection.

[0111] The pseudo-code structure feature extraction unit is used to extract structure features from the pseudo-code, such as the number of loops, the number of conditional statements, the number of function calls, etc. These features can reflect the logical structure of the pseudo-code and provide additional information for similarity detection.

[0112] The result output unit is used to output the extracted control flow graph and the structure features of the pseudo-code for use by the subsequent feature fusion module.

[0113] The feature fusion module is used to calculate the final similarity based on the control flow graph, the structure features of the pseudo-code, and the similarities X and Y.

[0114] As Figure 9 shown, this module includes a feature splicing unit, a similarity calculation unit, and a result output unit:

[0115] The feature splicing unit splices the control flow graph, the structure features of the pseudo-code, and the similarities X and Y to obtain the spliced features.

[0116] The similarity calculation unit calculates the final similarity score based on the spliced features using a multi-layer perceptron (MLP).

[0117] The result output unit is used to output the final similarity score for further analysis and decision-making by the user.

[0118] The above-described embodiments further elaborate on the purpose, technical solution, and advantages of the present invention. It should be understood that the above-described embodiments are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made to the present invention within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A binary code similarity detection method based on multi-level feature fusion, characterized in that: include: Obtain two binary files to be detected, disassemble the two binary files to be detected respectively, obtain code representations of the two binary files, pre-process the code representations of the two binary files and input them into a trained binary code similarity detection model to obtain a similarity score; The binary code similarity detection model includes: CFG branch, pseudo code branch, structural feature extraction module and feature fusion module; the training process of the binary code similarity detection model includes: S1: Obtain two binary files, disassemble the two binary files respectively, and obtain the code representation of the two binary files; wherein the code representation includes the control flow graph CFG n and pseudocode, the nodes of the control flow graph represent basic blocks, the edges represent the relationships between basic blocks, and each basic block contains multiple assembly instructions; S2: Control flow graph CFG of two binary files respectively n Perform preprocessing to obtain the control flow graph CFG of the two binary files after preprocessing n ; Where n is the index of the binary file; S3: Control flow graph CFG" after preprocessing the two binary files n Input the CFG branch and get the similarity X between the assembly instructions of the two binary files; S4: input the pseudocodes of the two binary files into the pseudocode branch to obtain the similarity Y between the pseudocodes of the two binary files; S5: Control flow graph CFG" after preprocessing the two binary files n The pseudo code represents the input structure feature extraction module, and the control flow graph CFG of the two binary files is obtained. n And the structural features of pseudocode F CFG,n 、F P,n ; S6: Convert the control flow graph CFG of the two binary files to n And the structural features of pseudocode F CFG,n and F P,n And the similarity X and Y are input into the feature fusion module to obtain the final similarity score; S7: Calculate the loss function value according to the final similarity score, update the model parameters according to the loss function value, and obtain the trained binary code similarity detection model when the loss function value is the smallest.

2. According to the multi-level feature fusion binary code similarity detection method of claim 1, it is characterized in that: Control Flow Graph CFG n Preprocessing includes: S21. Control Flow Graph CFG n The assembly instructions of each basic block are normalized to obtain the normalized control flow graph CFG′ n ; S22, set a threshold A, and control the control flow graph CFG′ according to the threshold A n Fill or cut to get a control flow graph containing A basic blocks S23, control flow graph according to threshold A Each basic block is filled or intercepted to obtain the preprocessed control flow graph CFG" n ; Among them, the control flow graph CFG″ n Each basic block of consists of A assembly instructions.

3. According to the multi-level feature fusion binary code similarity detection method of claim 1, it is characterized in that: The CFG branch includes a semantic embedding module and a function-level embedding module; CFG branch control flow graph CFG" for two binary files n The processing includes: respectively converting the control flow graph CFG" of the two binary files n Input the semantic embedding module to obtain the control flow graph CFG of the two binary files. n The high-dimensional embedding matrix H of each basic block n,i ;Respectively convert the control flow graph CFG of the two binary files into n The high-dimensional embedding matrix H of each basic block n,i The adjacency matrix is ​​input into the function-level embedding module to obtain the feature vectors of the assembly instructions of the two binary files, and the similarity X between the assembly instructions of the two binary files is calculated according to the feature vectors of the assembly instructions of the two binary files; wherein i is the index of the basic block.

4. According to the multi-level feature fusion binary code similarity detection method of claim 3, it is characterized in that: The semantic embedding module is the Word2Vec model, where Word2Vec is a word vector model.

5. According to the multi-level feature fusion binary code similarity detection method of claim 3, it is characterized in that: Function-level embedding modules include: Bi-GRU, GCN, and Bi-GRU and attention modules; the function-level embedding module controls the control flow graph CFG of the binary file n The high-dimensional embedding matrix H of each basic block n,i The processing of the adjacency matrix includes: embedding the high-dimensional matrix H of each basic block separately n,i Input the Bi-GRU network to obtain the global dependency feature H′ within each basic block n,i ; The global dependency feature H′ within each basic block n,i and control flow graph CFG″ n The adjacency matrix of is input into GCN to obtain the inter-block feature H″ of each basic block n,i , the inter-block feature H″ of each basic block n,i Input Bi-GRU and attention module to obtain the feature vector of assembly instruction; Bi-GRU is a bidirectional gated recurrent unit and GCN is a graph convolutional network.

6. The binary code similarity detection method of multi-level feature fusion according to claim 1 is characterized in that: The pseudocode branch includes: a Word2Vec model and an R-CNN model; the pseudocode branch processes the pseudocodes of the two binary files by: respectively inputting the pseudocodes of the two binary files into the Word2Vec model to obtain embedding vectors of the pseudocodes of the two binary files, respectively inputting the embedding vectors of the pseudocodes of the two binary files into the R-CNN model to obtain feature vectors of the pseudocodes of the two binary files, and calculating the similarity Y between the pseudocodes of the two binary files according to the feature vectors of the pseudocodes of the two binary files; wherein, Word2Vec is a word vector model, and R-CNN is a region-based convolutional neural network.

7. The binary code similarity detection method of multi-level feature fusion according to claim 1 is characterized in that: The feature fusion module uses the structural feature F CFG,n and F P,n And similarity X and Y processing includes: structural feature F CFG,n and F P,n The similarities X and Y are concatenated to obtain concatenated features. The concatenated features are processed using a multi-layer perceptron to obtain the final similarity score.

8. The binary code similarity detection method of multi-level feature fusion according to claim 1 is characterized in that: The structural characteristics of the control flow graph include: the number of loops, the number of strongly connected components, the number of back edges, CFG″ n Maximum width, CFG″ n The maximum depth, the number of nodes in each loop, the number of nodes in each strongly connected component, the average number of nodes in the loop, the average number of nodes in the strongly connected component, the average in-degree of the nodes, and the average out-degree of the nodes; the structural features of the pseudocode include: the number of loops, the number of conditional statements, the number of function calls, the number of jump statements, the number of variable declarations, the number of pointer operations, the number of arithmetic operations, and the number of memory allocations.

9. A binary code similarity detection system with multi-level feature fusion, the system is used to execute a binary code similarity detection method with multi-level feature fusion as claimed in any one of claims 1 to 8, characterized in that: include: A data input module is used to obtain two binary files to be tested, and use a disassembly tool to disassemble the two binary files to be tested to obtain a code representation of each binary file; the code representation includes a control flow graph and a pseudo code, the nodes of the control flow graph represent basic blocks, the edges represent the relationship between the basic blocks, and each basic block contains multiple assembly instructions; The data preprocessing module is used to preprocess the code representation of the two binary files to obtain the preprocessed control flow graph CFG" n ; Assembly similarity calculation module, used to calculate the similarity of the control flow graph CFG according to the preprocessed n Calculate the similarity X between the assembly instructions of two binary files; A pseudocode similarity calculation module is used to calculate the similarity Y between the pseudocodes of two binary files; The structural feature extraction module is used to extract the structural features of the control flow graph and pseudo code after preprocessing of the two binary files; The feature fusion module is used to calculate the final similarity based on the structural features of the control flow graph and the pseudocode and the similarities X and Y.

10. A binary code similarity detection system with multi-level feature fusion according to claim 9, characterized in that: The feature fusion module includes: a feature splicing unit, a similarity calculation unit and a result output unit; the feature splicing unit is used to splice the structural features of the control flow graph and the pseudocode as well as the similarities X and Y to obtain the splicing features; the similarity calculation unit is used to calculate the final similarity according to the splicing features; and the result output unit is used to output the final similarity.

Citation Information

Patent Citations

  • Binary program malicious code detection method, terminal equipment and storage medium

    CN112948828A

  • Method and device for detecting code similarity, storage medium and electronic equipment

    CN116956065A

  • Binary file similarity detection method and device, equipment, medium and program product

    CN118798156A

  • Cross-architecture binary code similarity detection method, system, equipment and medium

    CN118885827A

  • Display apparatus

    KR1020230003713A

Cited By

  • Binary program similarity analysis method and system

    CN121187640A