Binary code similarity detection method and system based on multi-modal feature fusion

Through the binary code similarity detection method of multimodal feature fusion, disassembled code and control flow graph features, combined with deep learning models for feature fusion, the problems of low detection accuracy and weak generalization ability in the existing technology are solved, and more efficient binary code similarity detection is achieved.

CN120277432AInactive Publication Date: 2025-07-08NAT UNIV OF DEFENSE TECH

Patent Information

Application Number
CN202510756489.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-07-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing binary code similarity detection methods are not fully used at the feature level, resulting in low detection accuracy and weak generalization ability of detection scenarios.

Method used

The multimodal feature fusion method is used to extract disassembled code from binary files as text sequence features and control flow graphs as topology structure features. The CLAP and GMN models are used for embedded characterization, and the feature fusion is performed through the multimodal fusion characterization model, and similarity is detected using Cosine vector distance calculation.

Benefits of technology

It realizes a more comprehensive representation of binary code, improves detection accuracy and generalization capabilities of detection scenarios, and can more accurately identify vulnerabilities, copyrights and malicious code.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277432A_ABST
    Figure CN120277432A_ABST
Patent Text Reader

Abstract

The invention provides a binary code similarity detection method and system based on multi-modal feature fusion, and belongs to the technical field of network security. According to the method, a program analysis method is utilized, disassembling codes are extracted from a binary file to serve as text sequence features, and a control flow graph is extracted to serve as topological graph structure features; aiming at two different modes of text sequence features and topological graph structure features, respectively using different representation models to carry out embedding representation; carrying out fusion processing on the embedded representation vectors of different modals by using a multi-modal fusion representation model, and generating a fused embedded representation vector; and based on the fused embedded representation vector, detecting the similarity degree in combination with a vector distance calculation formula, and completing similarity detection according to a preset threshold. According to the invention, the problems of low detection accuracy and weak detection scene generalization ability caused by incomplete use of the feature level of the binary code similarity detection method are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of network security, and in particular relates to a binary code similarity detection method and system based on multi-modal feature fusion. Background Art

[0002] In the contemporary software development field, code reuse has become a key practice for improving development efficiency and promoting innovation. Developers frequently draw on existing code libraries, open-source projects, and even third-party components to accelerate product iteration and function implementation. However, while this trend has brought significant productivity improvements, it has also inevitably been accompanied by a series of hidden security challenges. The quality of the reused code varies, and once the hidden unknown vulnerabilities and security defects are triggered, they may lead to catastrophic consequences. In the face of this situation, the industry uses code similarity detection technology in an attempt to identify and eliminate those code segments carrying security risks at the source. However, current technologies mostly focus on the analysis of single-modal features, such as text sequence features based on code or topological graph structure features based on programs, and do not fully consider the fused representation of multi-modal features. Although single-modal analysis methods can analyze code similarity to a certain extent, when faced with complex codes, there are problems of insufficient representation, resulting in low detection accuracy.

[0003] One of the existing technologies is based on the Transformer and BERT structures to build a deep learning model. The masked language model (MLM) is used to perform representation learning on the jump relationships between instructions in binary code. Through the pre-trained model, a relatively high success rate of jump relationship prediction can be finally obtained. After fine-tuning the model on the similarity detection task, the generated representation vectors can achieve a higher accuracy of similarity comparison and CVE vulnerability detection effect, but it still only considers the features of the text sequence modality.

[0004] Another existing technology fuses the semantics of control flow graphs, data flow graphs, and call graphs, uses BERT to embed instructions and operands at the same time, and uses GGNN to fuse the graph structure for representation, in order to use one model to jointly solve binary analysis tasks at the program level and function level; it redefines the data flow graph, where the nodes are each instruction representing the data flow direction, and the edges represent the data dependency relationships between instructions. It embeds and combines two types of features, text sequences and graph structures, but at the representation level, it still focuses on the topological graph structure features.

[0005] The third prior art is based on IDA Pro to process binary files to obtain microcode and control flow graphs. First, the RoBERTa model is used to process the instruction text features, and the obtained text representation vectors are used as the node attributes of the control flow graph. Second, the GCN model is used to process the function graph structure features, and the function semantic representation vectors are also generated with the graph structure as the main modality. After that, information such as basic blocks, string constants, and imported functions is added and combined through concatenation at the feature level, still without considering the interaction relationship between different modality features for fusion.

[0006] Therefore, the existing binary code similarity detection methods are not fully utilized at the feature level, resulting in low detection accuracy and weak generalization ability for detection scenarios. Summary of the Invention

[0007] For this reason, the present invention proposes a binary code similarity detection scheme based on multi-modal feature fusion, aiming to solve the technical problems of incomplete utilization at the feature level in the existing binary code similarity detection methods, resulting in low detection accuracy and weak generalization ability for detection scenarios.

[0008] The first aspect of the present invention proposes a binary code similarity detection method based on multi-modal feature fusion, and the method includes: Step S1, Binary code feature extraction: Using a program analysis method, disassembled code is extracted from the binary file as text sequence features, and a control flow graph is extracted as topological graph structure features; Step S2, Multi-modal feature independent embedding representation: For two different modalities of text sequence features and topological graph structure features, different representation models are respectively used for embedding representation to obtain embedding representation vectors of different modalities; Step S3, Multi-modal feature fusion embedding representation: Using a multi-modal fusion representation model, the embedding representation vectors of different modalities are fused and processed to generate fused embedding representation vectors; Step S4, Similarity detection: Based on the fused embedding representation vectors, the Cosine vector distance calculation formula is combined to detect the similarity degree, and the similarity detection is completed according to a preset threshold.

[0009] According to the method of the first aspect of the present invention, in step S1, disassembled code is extracted from the binary file as text sequence features, and a control flow graph is extracted as topological graph structure features; specifically including: using the decompilation tool Radare2 to convert the binary code into assembly code, and the generated assembly code is used as text sequence features; according to the assembly code, a program analysis tool is used to extract the control flow graph as topological graph structure features; the text sequence features and topological graph structure features are respectively output in the form of text sequences and adjacency matrices.

[0010] According to the method of the first aspect of the present invention, in step S2, for two different modalities, namely text sequence features and topological graph structure features, different representation models are respectively used for embedding representation to obtain embedding representation vectors of different modalities; specifically including: for text sequence features and topological graph structure features, different deep learning models are selected for processing and embedding representation; for text sequence features, the CLAP model is used as the text representation model for embedding representation; for topological graph structure features, the GMN model is used as the graph structure representation model for embedding representation.

[0011] According to the method of the first aspect of the present invention, in step S3, a multi-modal fusion representation model is used to fuse the embedding representation vectors of different modalities and generate a fused embedding representation vector; specifically including: using two fully connected layers to map the embedding vectors of text sequence features and topological graph structure features to the same dimension, using a mutual attention mechanism layer to preliminarily fuse text sequence features and topological graph structure features, using a Transformer encoder layer for further fusion, and outputting a fused embedding representation vector.

[0012] According to the method of the first aspect of the present invention, in step S3, the training process of the multi-modal fusion representation model specifically includes: Given the embedding vectors of text sequence features and topological graph structure features, determine whether they come from the same binary code; Given the embedding vectors of text sequence features and topological graph structure features, perform a masking operation on the embedding vector of the topological graph structure feature with a certain ratio, and use the graph structure representation model to recover the values at the corresponding positions; Given the embedding vectors of text sequence features and topological graph structure features from different binary codes, determine whether they come from the same source code according to the similarity between the generated fused embedding representation vectors.

[0013] The second aspect of the present invention proposes a binary code similarity detection system based on multi-modal feature fusion. The system includes a processing unit, and the processing unit is configured to execute: Binary code feature extraction: Using a program analysis method, disassembled code is extracted from a binary file as text sequence features, and a control flow graph is extracted as topological graph structure features; Multi-modal feature independent embedding representation: For two different modalities, namely text sequence features and topological graph structure features, different representation models are respectively used for embedding representation to obtain embedding representation vectors of different modalities; Multi-modal feature fusion embedding representation: Using a multi-modal fusion representation model to fuse the embedding representation vectors of different modalities and generate a fused embedding representation vector; Similarity detection: Based on the fused embedding representation vectors, the cosine vector distance calculation formula is combined to detect the similarity degree, and the similarity detection is completed according to the preset threshold.

[0014] According to the system of the second aspect of the present invention, the processing unit is configured to execute: extracting the disassembly code from the binary file as the text sequence feature, and extracting the control flow graph as the topological graph structure feature; specifically including: using the decompilation tool Radare2 to convert the binary code into assembly code, and the generated assembly code is used as the text sequence feature; according to the assembly code, using the program analysis tool to extract the control flow graph as the topological graph structure feature; outputting the text sequence feature and the topological graph structure feature in the form of a text sequence and an adjacency matrix respectively.

[0015] According to the system of the second aspect of the present invention, the processing unit is configured to execute: for two different modalities of the text sequence feature and the topological graph structure feature, different representation models are respectively used for embedding representation to obtain embedding representation vectors of different modalities; specifically including: for the text sequence feature and the topological graph structure feature, different deep learning models are selected for processing and embedding representation; for the text sequence feature, the CLAP model is used as the text representation model for embedding representation; for the topological graph structure feature, the GMN model is used as the graph structure representation model for embedding representation.

[0016] According to the system of the second aspect of the present invention, the processing unit is configured to execute: using a multi-modal fusion representation model to perform fusion processing on the embedding representation vectors of different modalities and generate a fused embedding representation vector; specifically including: using two fully connected layers to map the embedding vectors of the text sequence feature and the embedding vectors of the topological graph structure feature to the same dimension, using the mutual attention mechanism layer to perform preliminary fusion on the text sequence feature and the topological graph structure feature, using the Transformer encoder layer for further fusion, and outputting the fused embedding representation vector.

[0017] According to the system of the second aspect of the present invention, the training process of the multi-modal fusion representation model specifically includes: given the embedding vector of the text sequence feature and the embedding vector of the topological graph structure feature, determining whether they come from the same binary code; given the embedding vector of the text sequence feature and the embedding vector of the topological graph structure feature, performing a certain proportion of masking operation on the embedding vector of the topological graph structure feature, and using the graph structure representation model to recover the values at the corresponding positions; given the embedding vector of the text sequence feature and the embedding vector of the topological graph structure feature from different binary codes, determining whether they come from the same source code according to the similarity between the generated fused embedding representation vectors.

[0018] A third aspect of the present invention discloses an electronic device. The electronic device includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, a method for detecting binary code similarity based on multi-modal feature fusion according to the first aspect of the present disclosure is implemented.

[0019] A fourth aspect of the present invention discloses a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, a method for detecting binary code similarity based on multi-modal feature fusion according to the first aspect of the present disclosure is implemented.

[0020] The binary code similarity detection scheme based on multi-modal feature fusion proposed by the present invention is directed to service scenarios such as vulnerability mining, software copyright infringement detection, and malicious code detection. Based on binary code features of multiple modalities, a deep learning model structure is designed and combined with a multi-modal feature fusion task to optimize the parameters of the model. Finally, a more comprehensive and detailed representation of the binary code is achieved, so as to more accurately analyze and detect the similarity of the binary code. This scheme has the following technical effects: (1) In a variety of complex software, it matches whether there is a situation of reusing disclosed vulnerable code, so as to implement the vulnerability detection task; (2) Characterize the software function from multiple aspects to discover whether there are copyright issues in commercial software; (3) More comprehensively characterize the features of malicious software code, so as to achieve more accurate malicious code detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0022] Figure 1 It is a schematic diagram of the binary code similarity detection process based on multi-modal feature fusion.

[0023] Figure 2 It is a schematic diagram of the binary code feature extraction process.

[0024] FIG. 3(a) is a schematic diagram of the independent embedding representation process of the topological graph modal feature.

[0025] FIG. 3(b) is a schematic diagram of the independent embedding representation process of the text modal feature.

[0026] Figure 4 It is a schematic diagram of the multi-modal feature fusion embedding representation process.

[0027] Figure 5(a) is a schematic diagram of Training Task 1 during the multi-modal feature fusion training process.

[0028] Figure 5(b) is a schematic diagram of Training Task 2 and Task 3 during the multi-modal feature fusion training process. Detailed implementation manners

[0029] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only a part rather than all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0030] A method for detecting binary code similarity based on multi-modal feature fusion is proposed in the first aspect of the present invention. The method includes: Step S1, binary code feature extraction: Using a program analysis method, disassembled code is extracted from a binary file as text sequence features, and a control flow graph is extracted as topological graph structure features; Step S2, multi-modal feature independent embedding representation: For two different modalities of text sequence features and topological graph structure features, different representation models are respectively used for embedding representation to obtain embedding representation vectors of different modalities; Step S3, multi-modal feature fusion embedding representation: Using a multi-modal fusion representation model, the embedding representation vectors of different modalities are fused and processed to generate a fused embedding representation vector; Step S4, similarity detection: Based on the fused embedding representation vector, the similarity degree is detected by combining with the Cosine vector distance calculation formula, and the similarity detection is completed according to a preset threshold.

[0031] According to the method of the first aspect of the present invention, in Step S1, disassembled code is extracted from a binary file as text sequence features, and a control flow graph is extracted as topological graph structure features; specifically including: using the decompilation tool Radare2 to convert the binary code into assembly code, and the generated assembly code is used as text sequence features; according to the assembly code, using a program analysis tool to extract the control flow graph as topological graph structure features; the text sequence features and topological graph structure features are respectively output in the form of a text sequence and an adjacency matrix.

[0032] According to the method of the first aspect of the present invention, in step S2, for two different modalities of text sequence features and topological graph structure features, different representation models are respectively used for embedding representation to obtain embedding representation vectors of different modalities; specifically including: for text sequence features and topological graph structure features, different deep learning models are selected for processing and embedding representation; for text sequence features, the CLAP model is used as the text representation model for embedding representation; for topological graph structure features, the GMN model is used as the graph structure representation model for embedding representation.

[0033] According to the method of the first aspect of the present invention, in step S3, a multi-modal fusion representation model is used to fuse the embedding representation vectors of different modalities and generate a fused embedding representation vector; specifically including: using two fully connected layers to map the embedding vectors of text sequence features and topological graph structure features to the same dimension, using a mutual attention mechanism layer to preliminarily fuse text sequence features and topological graph structure features, using a Transformer encoder layer for further fusion, and outputting a fused embedding representation vector.

[0034] According to the method of the first aspect of the present invention, in step S3, the training process of the multi-modal fusion representation model specifically includes: Given the embedding vectors of text sequence features and topological graph structure features, determine whether they come from the same binary code; Given the embedding vectors of text sequence features and topological graph structure features, perform a masking operation on the embedding vectors of topological graph structure features with a certain ratio, and use the graph structure representation model to recover the values at the corresponding positions; Given the embedding vectors of text sequence features and topological graph structure features from different binary codes, determine whether they come from the same source code according to the similarity between the generated fused embedding representation vectors.

[0035] The second aspect of the present invention proposes a binary code similarity detection system based on multi-modal feature fusion. The system includes a processing unit, and the processing unit is configured to execute: Binary code feature extraction: Using a program analysis method, extract the disassembly code from the binary file as text sequence features, and extract the control flow graph as topological graph structure features; Multi-modal feature independent embedding representation: For two different modalities of text sequence features and topological graph structure features, different representation models are respectively used for embedding representation to obtain embedding representation vectors of different modalities; Multi-modal feature fusion embedding representation: Using a multi-modal fusion representation model, fuse the embedding representation vectors of different modalities and generate a fused embedding representation vector; Similarity detection: Based on the fused embedded representation vectors, the Cosine vector distance calculation formula is combined to detect the similarity degree, and the similarity detection is completed according to a preset threshold.

[0036] According to the system of the second aspect of the present invention, the processing unit is configured to execute: extracting disassembly code from a binary file as text sequence features, and extracting a control flow graph as topological graph structure features; specifically including: using the decompilation tool Radare2 to convert the binary code into assembly code, and the generated assembly code is used as text sequence features; according to the assembly code, using a program analysis tool to extract a control flow graph as topological graph structure features; and respectively outputting the text sequence features and topological graph structure features in the form of a text sequence and an adjacency matrix.

[0037] According to the system of the second aspect of the present invention, the processing unit is configured to execute: for two different modalities of text sequence features and topological graph structure features, different representation models are respectively used for embedded representation to obtain embedded representation vectors of different modalities; specifically including: for text sequence features and topological graph structure features, different deep learning models are selected for processing and embedded representation; for text sequence features, the CLAP model is used as a text representation model for embedded representation; for topological graph structure features, the GMN model is used as a graph structure representation model for embedded representation.

[0038] According to the system of the second aspect of the present invention, the processing unit is configured to execute: using a multi-modal fusion representation model to fuse the embedded representation vectors of different modalities and generate fused embedded representation vectors; specifically including: using two fully connected layers to map the embedded vectors of text sequence features and the embedded vectors of topological graph structure features to the same dimension, using a mutual attention mechanism layer to preliminarily fuse the text sequence features and topological graph structure features, using a Transformer encoder layer for further fusion, and outputting the fused embedded representation vectors.

[0039] According to the system of the second aspect of the present invention, the training process of the multi-modal fusion representation model specifically includes: given the embedded vectors of text sequence features and the embedded vectors of topological graph structure features, determining whether they come from the same binary code; given the embedded vectors of text sequence features and the embedded vectors of topological graph structure features, performing a masking operation on the embedded vectors of topological graph structure features at a certain ratio, and using the graph structure representation model to recover the values at the corresponding positions; given the embedded vectors of text sequence features and the embedded vectors of topological graph structure features from different binary codes, determining whether they come from the same source code according to the similarity between the generated fused embedded representation vectors.

[0040] The third aspect of the present invention discloses an electronic device. The electronic device includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, a binary code similarity detection method based on multi-modal feature fusion according to the first aspect of the present disclosure is implemented.

[0041] The fourth aspect of the present invention discloses a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, a binary code similarity detection method based on multi-modal feature fusion according to the first aspect of the present disclosure is implemented.

[0042] The binary code similarity detection scheme based on multi-modal feature fusion proposed by the present invention is directed to business scenarios such as vulnerability mining, software copyright plagiarism detection, and malicious code detection. Based on binary code features of multiple modalities, a deep learning model structure is designed and integrated with a multi-modal feature fusion task to optimize the parameters of the model. Finally, a more comprehensive and detailed representation of the binary code is achieved, so as to more accurately analyze and detect the similarity of the binary code. The scheme has the following technical effects: (1) In a variety of complex software, it matches whether there is a situation of reusing disclosed vulnerable code, so as to implement the vulnerability detection task; (2) Characterize the software functions from multiple aspects to discover whether there are copyright issues in commercial software; (3) More comprehensively characterize the features of malicious software code, so as to achieve more accurate malicious code detection.

[0043] The first embodiment

[0044] The binary code similarity detection process based on multi-modal feature fusion is as Figure 1 shown, and mainly includes four parts (modules): binary code feature extraction, multi-modal feature independent embedding representation, multi-modal feature fusion embedding representation, and similarity detection. First, the binary code feature extraction module uses a program analysis method to extract disassembly code from a binary file as text sequence features and extract a control flow graph as topological graph structure features; then the multi-modal feature independent embedding representation module uses different representation models for the two modalities of features respectively to perform embedding representation and obtain embedding representation vectors of different modalities; secondly, the multi-modal feature fusion embedding representation module uses a multi-modal fusion representation model to perform fusion processing on the representation vectors from different modalities and generate fused embedding representation vectors; finally, the similarity detection module, based on the fused embedding representation vectors, combines the Cosine vector distance calculation formula to detect the similarity degree and determines whether they are similar according to a preset threshold.

[0045] 1. Binary code feature extraction The binary code feature extraction module proposed by this method uses program analysis methods to extract disassembly code from binary files as text sequence features and extract control flow graphs as topological graph structure features. As Figure 2 shown, first use a decompilation tool (such as Radare2, etc.) to convert the binary code into assembly code, and the generated assembly code can be used as text sequence class features; then according to the assembly code, use a program analysis tool to extract the control flow graph, which can be used as graph structure class features. Finally, the module outputs the features in the form of text sequences and adjacency matrices for the two types of features respectively.

[0046] 2. Independent Embedding Representation of Multimodal Features The independent embedding representation module of multimodal features proposed by this method uses different representation models for the features of two modalities respectively for embedding representation. As shown in Figures 3(a) and 3(b), first, for the two types of modal features extracted by the previous module: text class features and graph structure features, select different deep learning models for processing; then use the corresponding deep learning models to perform embedding representation on the features. For text class features, use the CLAP model for embedding representation; for graph structure features, use the GMN model for embedding representation.

[0047] 3. Fusion Embedding Representation of Multimodal Features The fusion embedding representation module of multimodal features proposed by this method uses a multimodal fusion representation model to perform fusion processing on the representation vectors from different modalities and generate an embedded representation vector. As Figure 4 shown in the model structure, first use two fully connected layers to map the embedded vectors of text features and graph structure features to the same dimension, and then use a mutual attention mechanism layer to perform preliminary fusion on the two types of features. Finally, use a Transformer encoder layer for further fusion and output the feature fusion embedded vector.

[0048] As shown in Figures 5(a) and 5(b), the model training process designs three training tasks, namely: First, given the embedded vectors of text features and graph structure features, determine whether they come from the same binary code; Second, given the embedded vectors of text features and graph structure features, perform a 15% masking operation on the embedded vector of the graph structure feature, and use the model to recover the corresponding position values; Third, given the embedded vectors of text features and graph structure features from different binary codes, determine whether they come from the same source code according to the similarity between the generated feature fusion representation vectors. Among them, Task 2 and Task 3 can be carried out simultaneously.

[0049] 4. Similarity Detection The similarity detection module proposed by this method is based on the embedded representation vectors generated from different binary codes. Combining with the Cosine vector distance calculation formula, it detects the similarity degree of different binary codes and determines whether they come from the same source code according to a preset threshold.

[0050] The second embodiment

[0051] Taking the similarity detection of the version_etc function code in the md5sum binary program in the coreutils suite as an example. Coreutils is part of the GNU project and provides a set of basic file, shell, and text operation tools. These tools are an integral part of Unix and Unix-like systems (such as Linux) and provide many common command-line utilities. Coreutils contains a large number of commands covering multiple aspects such as file management, text processing, and system information.

[0052] 1. Binary code feature extraction First, select the md5sum binary program in the 6.5 version coreutils suite with the x86_32 architecture compiled by the 4.0 version clang compiler under the O0 option, and the md5sum binary program in the 6.5 version coreutils suite with the mips_32 architecture compiled by the 7.3.0 version gcc compiler under the O0 option as the comparison targets. Feature extraction is performed on the version_etc function in each of the two binary programs to extract text sequence features and topological graph structure features.

[0053] 2. Multi-modal feature independent embedded representation For features of different modalities, different embedded representation models are used to generate embedded vectors after embedding the features. For text sequence features, the CLAP model is used for embedded representation; for topological graph structure features, the GMN model is used for embedded representation, and embedded vectors are obtained after embedded representation.

[0054] 3. Multi-modal feature fusion embedded representation For two different coreutils-md5sum binary programs, the text sequence features and topological graph structure features of the version_etc function are respectively input into the multi-modal feature fusion module to generate vectorized features of the binary code.

[0055] 4. Binary code similarity detection Read the embedded vectors after multi-modal feature fusion corresponding to the version_etc functions in the two programs respectively, and calculate the distance using the cosine similarity distance. It is found that the cosine similarity between the vectors is 1, so it is determined to be similar. Therefore, this method can accurately represent the features of the same function code in programs under different compilation conditions and can be well used for binary code similarity detection.

[0056] In summary, the goal of the binary code similarity detection scheme proposed by the present invention is: given different binary codes, it is possible to generate an embedded representation vector by fusing features of multiple modalities according to the code semantics, and calculate the similarity based on the vector distance, so as to determine whether they come from the same source code.

[0057] The main technical improvements involved in the present invention include: 1. Design of the fusion representation model structure: First, use a fully connected layer to unify the dimensions of the embedded vectors of different modalities. Secondly, use a mutual attention mechanism layer for preliminary modality fusion. Finally, use a Transformer encoder layer for semantic understanding-level fusion and output the representation vector after feature fusion.

[0058] 2. Training of the fusion representation model parameters: Training the model parameters is a key step for the model to play its role. By designing three training tasks, the model is enabled to have semantic understanding ability and can output a more semantic-rich binary code representation vector.

[0059] Among them, the three training tasks include: (1) Given the embedded vectors of text features and the embedded vectors of graph structure features, determine whether they come from the same binary code; (2) Given the embedded vectors of text features and the embedded vectors of graph structure features, perform a 15% masking operation on the embedded vectors of graph structure features, and use the model to recover the corresponding position values; (3) Given the embedded vectors of text features and the embedded vectors of graph structure features from different binary codes, determine whether they come from the same source code according to the similarity between the generated feature fusion representation vectors.

[0060] Please note that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, all possible combinations of the technical features in the above embodiments are not described. However, as long as the combinations of these technical features do not conflict, they should be considered as within the scope described in this specification. The above embodiments only represent several implementation manners of the present application, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be pointed out that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

Claims

1. A binary code similarity detection method based on multi-modal feature fusion, characterized in that The method includes: Step S1, binary code feature extraction: Using a program analysis method, disassembled code is extracted from a binary file as text sequence features, and a control flow graph is extracted as topological graph structure features; Step S2, multi-modal feature independent embedding representation: For two different modalities of text sequence features and topological graph structure features, different representation models are used for embedding representation respectively to obtain embedding representation vectors of different modalities; Step S3, multi-modal feature fusion embedding representation: Using a multi-modal fusion representation model, the embedding representation vectors of different modalities are fused and processed to generate fused embedding representation vectors; Step S4, similarity detection: Based on the fused embedding representation vectors, the Cosine vector distance calculation formula is combined to detect the similarity degree, and the similarity detection is completed according to a preset threshold.

2. The binary code similarity detection method based on multi-modal feature fusion according to claim 1, wherein, In step S1, disassembled code is extracted from a binary file as text sequence features, and a control flow graph is extracted as topological graph structure features; specifically including: using the decompilation tool Radare2 to convert the binary code into assembly code, and the generated assembly code is used as text sequence features; according to the assembly code, a program analysis tool is used to extract the control flow graph as topological graph structure features; the text sequence features and topological graph structure features are respectively output in the form of a text sequence and an adjacency matrix.

3. The binary code similarity detection method based on multi-modal feature fusion according to claim 2, wherein In step S2, for two different modalities of text sequence features and topological graph structure features, different representation models are used for embedding representation respectively to obtain embedding representation vectors of different modalities; specifically including: for text sequence features and topological graph structure features, different deep learning models are selected for processing and embedding representation; the CLAP model is used as the text representation model for embedding representation of text sequence features; the GMN model is used as the graph structure representation model for embedding representation of topological graph structure features.

4. The binary code similarity detection method based on multi-modal feature fusion according to claim 3, wherein, In step S3, using a multi-modal fusion representation model, the embedding representation vectors of different modalities are fused and processed to generate fused embedding representation vectors; specifically including: using two fully connected layers to map the embedding vectors of text sequence features and the embedding vectors of topological graph structure features to the same dimension, using a mutual attention mechanism layer to preliminarily fuse the text sequence features and topological graph structure features, using a Transformer encoder layer for further fusion, and outputting the fused embedding representation vectors.

5. A binary code similarity detection method based on multi-modal feature fusion according to claim 4, characterized in that In step S3, the training process of the multi-modal fusion representation model specifically includes: Given the embedding vectors of text sequence features and the embedding vectors of topological graph structure features, it is judged whether they come from the same binary code; Given the embedding vectors of text sequence features and the embedding vectors of topological graph structure features, a masking operation with a certain ratio is performed on the embedding vectors of topological graph structure features, and the graph structure representation model is used to recover the values at the corresponding positions; Given the embedding vectors of text sequence features and the embedding vectors of topological graph structure features from different binary codes, it is judged whether they come from the same source code according to the similarity between the generated fused embedding representation vectors.

6. A binary code similarity detection system based on multi-modal feature fusion, characterized in that The system includes a processing unit configured to perform: Binary code feature extraction: Using program analysis methods, extract the disassembly code from the binary file as text sequence features, and extract the control flow graph as topological graph structure features; Multi-modal feature independent embedding representation: For the two different modalities of text sequence features and topological graph structure features, use different representation models for embedding representation respectively to obtain embedding representation vectors of different modalities; Multi-modal feature fusion embedding representation: Use a multi-modal fusion representation model to perform fusion processing on the embedding representation vectors of different modalities and generate fused embedding representation vectors; Similarity detection: Based on the fused embedding representation vectors, combine the Cosine vector distance calculation formula to detect the similarity degree, and complete the similarity detection according to the preset threshold.

7. The binary code similarity detection system based on multi-modal feature fusion according to claim 6, wherein The processing unit is configured to perform: Extract the disassembly code from the binary file as text sequence features, and extract the control flow graph as topological graph structure features; specifically include: Use the decompilation tool Radare2 to convert the binary code into assembly code, and the generated assembly code is used as text sequence features; According to the assembly code, use a program analysis tool to extract the control flow graph as topological graph structure features; Output the features of the text sequence features and topological graph structure features in the form of text sequences and adjacency matrices respectively; For the two different modalities of text sequence features and topological graph structure features, use different representation models for embedding representation respectively to obtain embedding representation vectors of different modalities; specifically include: For text sequence features and topological graph structure features, select different deep learning models for processing and embedding representation; Use the CLAP model as the text representation model for embedding representation of text sequence features; Use the GMN model as the graph structure representation model for embedding representation of topological graph structure features.

8. A binary code similarity detection system based on multi-modal feature fusion according to claim 7, characterized in that, The processing unit is configured to perform: Use a multi-modal fusion representation model to perform fusion processing on the embedding representation vectors of different modalities and generate fused embedding representation vectors; specifically include: Use two fully connected layers to map the embedding vectors of text sequence features and the embedding vectors of topological graph structure features to the same dimension, use the mutual attention mechanism layer to perform preliminary fusion on text sequence features and topological graph structure features, use the Transformer encoder layer for further fusion, and output the fused embedding representation vectors; The training process of the multi-modal fusion representation model specifically includes: Given the embedding vectors of text sequence features and the embedding vectors of topological graph structure features, determine whether they come from the same binary code; Given the embedding vectors of text sequence features and the embedding vectors of topological graph structure features, perform a certain proportion of masking operation on the embedding vectors of topological graph structure features, and use the graph structure representation model to recover the values at the corresponding positions; Given the embedding vectors of text sequence features and the embedding vectors of topological graph structure features from different binary codes, determine whether they come from the same source code according to the similarity between the generated fused embedding representation vectors.

9. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, a method for detecting binary code similarity based on multi-modal feature fusion according to any one of claims 1-5 is implemented.

10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, a method for detecting binary code similarity based on multi-modal feature fusion according to any one of claims 1-5 is implemented.

Citation Information

Patent Citations

  • Comparable learning object generation method and system based on graph matching network

    CN113095361A

  • Vulnerability detection method and device, electronic equipment, storage medium and program product

    CN116663008A

  • Multi-mode and adversarial learning-based cross-site subject identity linking method, system and equipment

    CN118821052A

  • Cross-architecture binary code similarity detection method, system, equipment and medium

    CN118885827A

  • Multi-modal semantic alignment method oriented to classroom teaching guidance

    CN119669710A

Cited By

  • Binary program similarity analysis method and system

    CN121187640A