Fine-grained binary code similarity detection method based on code retrieval and subgraph matching

By matching the code retrieval with subgraphs, the existing binary code similarity detection has low accuracy and weak generalization ability in complex scenarios, and the fine-grained similarity detection of binary code is realized, which can identify the reuse of function inline and code modification, improving the accuracy and applicability of detection.

CN120256980AActive Publication Date: 2025-07-04NAT UNIV OF DEFENSE TECH

Patent Information

Application Number
CN202510756501.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-07-04
Estimated Expiration
2045-06-09

AI Technical Summary

Technical Problem

The existing binary code similarity detection methods can only meet the detection requirements of one-to-one matching, and cannot cope with complex detection scenarios such as function inline and code modification and reuse. The detection accuracy is low and the detection scenario generalization ability is weak.

Method used

Using a method based on code search matching with subgraphs, the binary code to be tested is vectorized embedded at the basic block level and searched in the binary code database, and the similarity calculation and graph matching are calculated using the IVF-PQ algorithm and the SeedGNN model to achieve fine-grained binary code similarity detection.

Benefits of technology

It realizes more accurate and finer-grained similarity detection of binary code in complex detection scenarios, can identify the reuse of function inline and code modification, and improves the accuracy and generalization ability of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256980A_ABST
    Figure CN120256980A_ABST
Patent Text Reader

Abstract

The invention provides a fine-grained binary code similarity detection method based on code retrieval and subgraph matching, and belongs to the technical field of network security. The method comprises the following steps of: performing vectorization embedding on a binary code to be tested at a basic block level on the basis of code snippet-level fine-grained retrieval; and retrieving the embedded vector at the basic block level in an existing binary code database, thereby performing binary code matching at the basic block level. According to the method, each matched basic block is expanded to an adjacent basic block to form a sub-graph, and the matched basic blocks are used as seed nodes; and performing seed graph matching on the plurality of discontinuous sub-graphs, and judging the similarity between the binary code to be detected and the code in the binary code database, thereby completing the similarity detection of the fine-grained binary code. According to the method, the problems of low accuracy and weak generalization ability of a detection scene of traditional binary code similarity detection in complex detection scenes such as function inline and multiplexing after code modification are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of network security, and particularly relates to a fine-grained binary code similarity detection method based on code retrieval and subgraph matching. Background Art

[0002] In the contemporary software development field, code reuse has become a key practice for improving development efficiency and promoting innovation. Developers frequently draw on existing code libraries, open-source projects, and even third-party components to accelerate product iteration and function implementation. However, while this trend has brought significant productivity improvements, it has also inevitably been accompanied by a series of hidden security challenges. The quality of the reused code varies widely, and once the hidden unknown vulnerabilities and security defects are triggered, they may lead to catastrophic consequences. In the face of this situation, the industry uses code similarity detection technology in an attempt to identify and eliminate those code fragments carrying security risks at the source. However, current technologies mostly focus on the analysis of code at the function level and fail to fully consider the more fine-grained binary code similarity detection within functions required for downstream applications such as copyright detection and vulnerability detection. The core code of a software project, or the function code containing vulnerabilities, often has a large amount of project-specific content added by developers during the reuse process, resulting in a situation where, at the function level, the original code and the reused code cannot be matched one by one, and there is essentially an inclusion relationship between the two. In response to this phenomenon, similarity detection needs to be carried out at a code granularity finer than the function granularity.

[0003] At the binary code level, different from the source code level where string matching can be used for matching, due to the source code compilation step, even if the binary code is derived from the same source code, differences in the compilation method will result in significant changes in the binary code. Therefore, both the one-to-one matching at the function level and the fine-grained inclusion relationship detection are technically difficult to implement.

[0004] One of the existing technologies detects template class code in C++ that has been inlined and optimized by extracting the fingerprints of the binary code and matching the fingerprints. In the fingerprint extraction part, it mainly captures the syntax and semantic features in the assembly code, as well as the control flow graph structure. In addition to the extracted fingerprints, subgraph isomorphism detection technology is also used to assist in identifying the calls of inline functions in the target binary code. In the fingerprint library generation part, based on the known template class source code, inlined calls are made and compiled into binary files, and then the fingerprints are extracted and stored in the library. This method is limited by the collected fingerprint library, and if the detection scope exceeds the content in the library, the detection accuracy is limited.

[0005] The second prior art is to detect inline functions based on topological graph theory. It proposes an instruction topological graph (ITG) to represent the data flow dependencies of instructions in basic blocks, and with the help of ITG, the problem of distinguishing inline instructions from caller instructions is transformed into a graph connectivity problem, which is finally solved by calculating the minimum vertex partition set. If there is a minimum connected graph, the corresponding position code can be detected as the inlined function code according to the corresponding threshold setting. This method is somewhat inspiring for fine-grained detection, but it cannot be directly migrated to the field of similarity detection.

[0006] The third prior art first built a cross-inline binary code retrieval database, compiled 51 projects using 9 compilers, 4 optimization levels, 6 architectures, and 2 inline flags, and generated two data sets with 216 combinations each. Through the analysis of the inline data set, three function inline modes were found, and different detection models were selected and trained for different modes to achieve detection for all inline modes. However, the effectiveness of the model detection of this method is limited by the data used. Once the data distribution changes in the application scenario, the detection accuracy will be reduced.

[0007] It can be seen that the existing binary code similarity detection method can only meet the detection requirements of one-to-one matching, but cannot cope with detection in complex detection scenarios such as function inlining and code reuse after modification. That is, the detection accuracy in the corresponding scenarios is low and the generalization ability of the detection scenarios is weak. Summary of the invention

[0008] To this end, the present invention proposes a fine-grained binary code similarity detection scheme based on code retrieval and subgraph matching, aiming to solve the technical problem that the existing binary code similarity detection method can only meet the detection requirements of one-to-one matching, but cannot cope with detection in complex detection scenarios such as function inlining and code reuse after modification, that is, the detection accuracy in the corresponding scenarios is low and the generalization ability of detection scenarios is weak.

[0009] The first aspect of the present invention provides a fine-grained binary code similarity detection method based on code retrieval and subgraph matching, the method comprising:

[0010] Step S1, based on fine-grained retrieval at the code snippet level, vectorize and embed the binary code to be tested at the basic block level; retrieve the embedding vector at the basic block level in the existing binary code database, so as to perform binary code matching at the basic block level;

[0011] Step S2: Each matched basic block is extended to its adjacent basic blocks to form a subgraph, and the matched basic block is used as a seed node; by performing seed graph matching on multiple discontinuous subgraphs, the similarity between the binary code to be tested and the code in the binary code database is determined, thereby completing fine-grained binary code similarity detection.

[0012] According to the method of the first aspect of the present invention, in step S1, the binary code to be tested is vectorized and embedded at the basic block level; specifically, it includes:

[0013] The binary code to be tested at the binary function level is segmented into code fragments based on the control flow graph, and divided into two levels of code units, namely:

[0014] A code unit at the continuous basic block level composed of two adjacent basic blocks; and

[0015] A code unit at the single basic block level composed of all individual basic blocks;

[0016] By segmenting the binary code to be tested, the detection granularity is converted. Using the sem2vec tool, the continuous binary assembly code in the basic block is input, and an embedded representation in vector form is output.

[0017] According to the method of the first aspect of the present invention, in step S1, the embedded vectors at the basic block level are retrieved in the existing binary code database; specifically, it includes:

[0018] The IVF-PQ algorithm is used to perform code retrieval based on vector similarity. IVF-PQ consists of two parts, namely product quantization PQ and inverted file IVF;

[0019] Product quantization PQ is based on vector compression technology, which is used to reduce the dimension of the embedded vector and accelerate the calculation speed of vector similarity; there are N vectors with dimension D in the vector library, and each floating-point value is represented by d bits. PQ divides the high-dimensional vector into M groups of sub-vectors with dimension D / M, performs K-means clustering on the M groups of sub-vectors, so that each group generates K clustering centers and incorporates them into the codebook of the sub-vectors in this group;

[0020] The inverted file IVF is applied to the retrieval process. The inverted file index uses the fragmented content of the file as the key of the index and the serial numbers of all files containing the corresponding content as the values, so as to retrieve specific file content; the points in the space are divided into n list units through K-means clustering. During the query, the target vector is compared with the centers of all units, and n probe nearest units are selected, and all vectors in the selected units are compared to obtain the final result.

[0021] According to the method of the first aspect of the present invention, in step S1:

[0022] The IVF-PQ algorithm reduces the vector range for retrieval through IVF, uses PQ for similarity calculation and retrieval of similar vectors, and obtains one or a cluster of vectors that are most similar to the vector to be measured; thus, it realizes the similarity calculation of basic block-level embedded vectors in the binary code database.

[0023] After calculating the similarities between all basic block units and the basic block units in the binary code to be measured, perform similarity sorting, and select the top K B code segments with basic block unit similarities higher than the threshold T top and a relatively large number of matching units as matching candidates for similarity detection.

[0024] According to the method of the first aspect of the present invention, in step S2, seed graph matching refers to detecting whether the target code exists in the binary code to be measured or detecting whether the target code contains the binary code to be measured based on the matched basic blocks and basic block sequences.

[0025] According to the method of the first aspect of the present invention, in step S2, among the K top code segments, add adjacent nodes to multiple matching basic block units with similarities higher than the threshold, thereby expanding them into multiple subgraphs, and perform similarity matching through the graph matching algorithm; where:

[0026] When the number of matched similar subgraphs reaches the threshold T G then it is determined that the binary code similarity matching is successful; and based on the already matched nodes, use them as seed nodes to further perform graph structure matching.

[0027] According to the method of the first aspect of the present invention, in step S2, use the supervised seed graph matching model SeedGNN, take the matched basic block units as seed nodes, and perform seed graph matching; where:

[0028] SeedGNN consists of multiple layers of neural networks, uses convolutional operations for node attribute propagation, uses a filtering module to deal with the situation where seed nodes in the nodes are interfered by noise nodes, and uses the Hungarian algorithm to pre-match the nodes with the maximum similarity between the two graphs, thereby distinguishing the real seed nodes from the noise nodes. After masking the noise nodes, use the real seed nodes for the operation of the graph neural network;

[0029] As the number of convolutional layers increases, new seed nodes are continuously added to the operation, and the message passing in the graph neural network deepens step by step. Finally, the similarity values between each pair of nodes in the two target graph structures are obtained; determine the matching result of the graph by setting a threshold for screening, thereby determining the matching ratio in the subgraphs of the target code segment and the binary code segment to be measured, and further determining whether the two code segments are similar.

[0030] In a second aspect of the present invention, a fine-grained binary code similarity detection system based on code retrieval and subgraph matching is proposed. The system includes a processing unit configured to execute:

[0031] Based on fine-grained retrieval at the code snippet level, vectorize and embed the binary code to be tested at the basic block level; retrieve the embedded vectors at the basic block level in the existing binary code database, so as to perform binary code matching at the basic block level;

[0032] Expand each matched basic block to its adjacent basic blocks to form a subgraph, and use the matched basic block as a seed node; judge the similarity degree between the binary code to be tested and the code in the binary code database by performing seed graph matching on multiple discontinuous subgraphs, so as to complete the fine-grained binary code similarity detection.

[0033] In a third aspect of the present invention, an electronic device is disclosed. The electronic device includes a memory and a processor. When the processor executes the computer program stored in the memory, a fine-grained binary code similarity detection method based on code retrieval and subgraph matching according to the first aspect of the present disclosure is implemented.

[0034] In a fourth aspect of the present invention, a computer-readable storage medium is disclosed. A computer program is stored on the computer-readable storage medium. When the computer program is executed by the processor, a fine-grained binary code similarity detection method based on code retrieval and subgraph matching according to the first aspect of the present disclosure is implemented.

[0035] The fine-grained binary code similarity detection scheme based on code retrieval and subgraph matching proposed by the present invention is aimed at business scenarios such as vulnerability mining, software copyright infringement detection, and malicious code detection. Based on code retrieval and subgraph matching, it realizes more fine-grained and accurate similarity detection of binary codes. It has the following technical effects: (1) In various complex software, match whether there is a situation of reusing disclosed vulnerable codes, so as to realize the vulnerability detection task; (2) Characterize the software functions from multiple aspects, so as to discover whether there are copyright issues in commercial software; (3) More comprehensively characterize the features of malicious software codes, so as to realize more accurate malicious code detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0037] Figure 1 Schematic diagram of fine-grained binary code similarity detection based on code retrieval and subgraph matching.

[0038] Figure 2 Schematic diagram of the principle of the product quantization PQ algorithm.

[0039] Figure 3 Schematic diagram of the SeedGNN network structure.

[0040] Figure 4 Schematic diagram of the SeedGNN seed graph matching processing process. Specific implementation manners

[0041] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0042] The first embodiment

[0043] The present invention proposes a fine-grained binary code similarity detection scheme based on code retrieval and subgraph matching. As Figure 1 shown, it mainly includes two parts (modules): code text retrieval and seed graph matching. First, through fine-grained retrieval at the basic block level of the binary code, the binary code to be tested is vectorized and embedded at the basic block level. Then, in the existing binary code database, these embedded vectors at the basic block level are efficiently retrieved to achieve accurate matching of the binary code at the basic block level. Next, each matched basic block is extended to its adjacent basic blocks to form a subgraph, and these matched basic blocks are used as seed nodes. By performing seed graph matching on multiple discontinuous subgraphs, the similarity degree between the binary code to be tested and the code in the database is judged, and finally, fine-grained binary code similarity detection is realized.

[0044] 1. Code text retrieval

[0045] The code text retrieval proposed by this method performs vectorized representation and indexing on the binary code at the basic block level, and uses an efficient vector retrieval algorithm for similarity matching to finally obtain the matching result of the binary code at the basic block level.

[0046] The binary code vectorization representation part first divides the binary function-level code based on the control flow graph into code units at two levels, namely, code units at the continuous basic block level composed of two adjacent basic blocks and code units at the single basic block level composed of all individual basic blocks. After the detection granularity is converted by splitting the binary code to be detected, sem2vec is used to output a vector-form representation by inputting the continuous binary assembly code in the basic block.

[0047] The efficient vector retrieval part, since the dimension of the generated representation vector is usually large and dimensionality reduction will lead to a decline in representation performance, adopts the IVF-PQ approximate nearest neighbor (ANN) algorithm to achieve efficient code retrieval based on vector similarity. IVF-PQ mainly consists of two parts, namely, product quantization (PQ) and inverted file (IVF).

[0048] PQ is a vector compression technology, that is, a quantization technology, used to reduce the dimension of the embedded vector and speed up the calculation speed of vector similarity. Suppose there are N vectors with dimension D in the vector library, and each floating-point value is represented by d bits. PQ divides the high-dimensional vector into M sub-vectors with dimension D / M, and performs clustering such as K-means on M groups of sub-vectors, so that K clustering centers are generated in each group and incorporated into the codebook of this group of sub-vectors, as Figure 2 shown.

[0049] IVF is commonly used in retrieval systems. Different from the traditional method of finding the file content according to the file number (the number is the key and the file content is the value), the inverted file index uses the fragmented content of the file as the key of the index and the file numbers of all files containing the corresponding content as the value. In this way, the specific file content can be retrieved from which file or files in a shorter time. It divides the points in the space into n list units through clustering methods such as K-means. During the query, the target vector is first compared with the centers of all units to select n probe nearest units, and then all vectors in these selected units are compared to obtain the final result.

[0050] The IVF-PQ algorithm first reduces the vector range for retrieval through IVF, and then uses PQ for efficient similarity calculation and retrieval of similar vectors, ultimately obtaining one or a cluster of vectors in the library that are most similar to the vector to be measured. In this way, fast similarity calculation of basic block-level embedded vectors can be achieved in a large-scale binary code database. After calculating the similarity between all basic block units in the database and the basic block units in the code to be measured, a similarity ranking is performed, and the first K B segments of code with basic block unit similarity higher than the threshold T top , and a relatively large number of matched units are used as matching candidates for the next similarity detection.

[0051] 2. Seed graph matching

[0052] This method involves the seed graph matching process. Based on the matched basic blocks and basic block sequences, it detects whether the target code truly exists in the code to be measured, or whether the target code contains the code to be measured.

[0053] Among the K top segments of candidate code obtained in the previous module, multiple matching basic block units with similarity higher than the threshold are added with adjacent nodes to expand into multiple subgraphs, and similarity matching is performed through the graph matching algorithm. When the number of similar subgraphs matched reaches the threshold T G , it can be determined that the two binary codes are similar matches. Since there are already some matched basic blocks in the subgraph, the seed graph matching algorithm can be used, that is, based on the already matched nodes, taking them as seed nodes to further perform graph structure matching.

[0054] The project plans to use the supervised seed graph matching model SeedGNN, taking the matched basic block units as seed nodes for seed graph matching. SeedGNN consists of multiple layers of neural networks, and the structure of each layer is as Figure 3 shown. First, convolutional operations are used for node attribute propagation; then, a filtering module is used to address the problem that seed nodes in the nodes are affected by noise nodes (with high similarity but not corresponding nodes). The Hungarian algorithm is used to pre-match the nodes with the maximum similarity between the two graphs first, so as to distinguish the true seed nodes from the noise nodes, and then after masking the noise nodes, the true seed nodes are continued to be used for graph neural network operations.

[0055] As the number of convolutional layers increases, new seed nodes are continuously added to the operation, and the message passing in the graph neural network also deepens level by level, as Figure 4As shown, the similarity values between each pair of nodes in the two target graph structures can be finally obtained. By setting a threshold for screening, the matching result of the graph can be judged, so that the matching ratio of the subgraphs of the target code snippet and the code snippet to be tested can be determined, and finally whether the two code snippets are similar can be determined.

[0056] As Figure 4 shown, where: S l is the input seed node of the l-th layer, that is, the output seed node of the (l - 1)-th layer; S l+1 is the output seed node of the l-th layer, that is, the input seed node of the (l + 1)-th layer; H l = A1S l A2; A1 and A2 are respectively the input value S l and two parameter matrices in the l-th convolutional layer, and the matrix multiplication operation is performed, and the operation result is H l . Softmax is used to convert each element in a vector into a probability value, and the range of these probability values is (0, 1), and all values add up to 1.

[0057] Second Embodiment

[0058] Taking the similarity detection of the version_etc function code in the md5sum binary program in the coreutils suite as an example.

[0059] coreutils is part of the GNU project and provides a set of basic file, shell, and text manipulation tools. These tools are an integral part of Unix and Unix-like systems (such as Linux) and provide many common command-line utilities. coreutils contains a large number of commands covering multiple aspects such as file management, text processing, and system information.

[0060] 1. Binary Code Feature Extraction

[0061] First, select the md5sum binary program in the 6.5 version coreutils suite with the x86_32 architecture compiled by the 4.0 version clang compiler under the O0 option, and the md5sum binary program in the 6.5 version coreutils suite with the mips_32 architecture compiled by the 7.3.0 version gcc compiler under the O0 option as the comparison targets. Feature extraction is performed on the version_etc function in the two binary programs respectively, and text sequence features and topological graph structure features are extracted.

[0062] 2. Multi-modal Feature Independent Embedding Representation

[0063] For features of different modalities, different embedding representation models are used to generate embedding vectors after embedding the features. For text sequence features, the CLAP model is used for embedding representation; for topological graph structure features, the GMN model is used for embedding representation, and embedding vectors are obtained after embedding representation.

[0064] 3. Embedding Representation of Multimodal Feature Fusion

[0065] For two different coreutils-md5sum binary programs, the text sequence features and topological graph structure features of the version_etc function are respectively input into the multimodal feature fusion module to generate vectorized features of the binary code.

[0066] 4. Binary Code Similarity Detection

[0067] The embedding vectors after multimodal feature fusion corresponding to the version_etc function in the two programs are read respectively, and the cosine similarity distance is used for distance calculation. It is found that the cosine similarity between the vectors is 1, so they are determined to be similar. Therefore, this method can accurately represent the features of the same function code in programs under different compilation conditions and can be well used for binary code similarity detection.

[0068] It can be seen that the goal of the binary code similarity detection method described in the present invention is: given a binary code to be tested and a known binary code database, it is possible to achieve fine-grained binary code similarity detection at the basic block level within the function by successively performing two steps of code text retrieval and seed graph matching, calculating the similarity according to the vector distance, and thus determining whether they come from the same source code.

[0069] The main technical improvements of the present invention include: (1) Design of the code text retrieval method: First, it is necessary to select a semantic representation model for accurately representing the continuous binary assembly code within the basic block to perform accurate representation at the semantic level; second, it is necessary to select a vector retrieval method with high operation efficiency to save time consumption in massive calculation scenarios; finally, a manually determined threshold is used to screen the basic blocks with a similarity matching degree higher than the threshold. (2) Design of the seed graph matching method: First, the selected matching basic blocks are extended to adjacent basic block nodes to form subgraphs; second, for each subgraph, it is necessary to select a suitable seed graph matching algorithm to match the graph matching results between the binary code to be tested and the binary code in the library; finally, a manually determined threshold is used to judge the similarity degree between the binary codes according to the number of nodes matched in the subgraph and the number of successfully matched subgraphs.

[0070] Please note that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as within the scope described in this specification. The above embodiments only express several implementation manners of the present application, and the description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. A fine-grained binary code similarity detection method based on code retrieval and subgraph matching, characterized in that The method includes: Step S1: Based on fine-grained retrieval at the code snippet level, vectorize and embed the binary code to be tested at the basic block level; retrieve the embedded vectors at the basic block level in the existing binary code database, so as to perform binary code matching at the basic block level; Step S2: Expand each matched basic block to its adjacent basic blocks to form a subgraph, and use the matched basic block as a seed node; judge the similarity between the binary code to be tested and the code in the binary code database by performing seed graph matching on multiple discontinuous subgraphs, so as to complete the fine-grained binary code similarity detection.

2. The fine-grained binary code similarity detection method based on code retrieval and subgraph matching according to claim 1, wherein In step S1, vectorize and embed the binary code to be tested at the basic block level; specifically including: Segment the binary code to be tested at the binary function level based on the control flow graph into two levels of code units, namely: Code units at the continuous basic block level composed of two adjacent basic blocks; and Code units at the single basic block level composed of all individual basic blocks; Perform detection granularity conversion by segmenting the binary code to be detected, use the sem2vec tool, input the continuous binary assembly code in the basic block, and output the embedded representation in vector form.

3. The fine-grained binary code similarity detection method based on code retrieval and subgraph matching according to claim 2, wherein, In step S1, retrieve the embedded vectors at the basic block level in the existing binary code database; specifically including: Adopt the IVF-PQ algorithm to perform code retrieval based on vector similarity. IVF-PQ consists of two parts, namely product quantization PQ and inverted file IVF; Product quantization PQ is based on vector compression technology, which is used to reduce the dimension of the embedded vector and speed up the calculation speed of vector similarity; there are N vectors with dimension D in the vector library, and each floating-point value is represented by d bits. PQ divides the high-dimensional vector into M groups of sub-vectors with dimension D / M, performs K-means clustering on the M groups of sub-vectors, so that each group generates K clustering centers and incorporates them into the codebook of the sub-vectors of this group; The inverted file IVF is applied to the retrieval process. The inverted file index uses the fragmented content of the file as the key of the index and the serial numbers of all files containing the corresponding content as the value, so as to retrieve specific file content; the points in the space are divided into n units through K-means clustering. During the query, the target vector is compared with the centers of all units, and n nearest units are selected. Then, all vectors in the selected units are compared to obtain the final result. list units. During the query, the target vector is compared with the centers of all units, and n probe nearest units are selected. Then, all vectors in the selected units are compared to obtain the final result.

4. A fine-grained binary code similarity detection method based on code retrieval and subgraph matching according to claim 3, characterized in that In step S1: The IVF-PQ algorithm reduces the vector range used for retrieval through IVF, uses PQ for similarity calculation and retrieval of similar vectors, and obtains one or a cluster of vectors that are most similar to the vector to be tested; thus realizing the similarity calculation of the embedded vectors at the basic block level in the binary code database. After calculating the similarity between all basic block units and the basic block units in the binary code to be tested, perform similarity sorting, and select the first K B segments of code with a basic block unit similarity higher than the threshold T top and a relatively large number of matched units as matching candidates for similarity detection.

5. A fine-grained binary code similarity detection method based on code retrieval and subgraph matching according to claim 4, characterized in that In step S2, seed graph matching refers to detecting whether the target code exists in the binary code to be tested or detecting whether the target code contains the binary code to be tested based on the matched basic block and basic block sequence.

6. The fine-grained binary code similarity detection method based on code retrieval and subgraph matching according to claim 5, characterized in that In step S2, at K top Among the top pieces of code, adjacent nodes are added to multiple matching basic block units with a similarity higher than the threshold, so as to be expanded into multiple subgraphs, and similarity matching is performed through a graph matching algorithm; where: When the number of matching similar subgraphs reaches the threshold T G it is determined that the binary code similarity matching is successful; and based on the already matched nodes, they are used as seed nodes to further perform graph structure matching.

7. A fine-grained binary code similarity detection method based on code retrieval and subgraph matching according to claim 6, characterized in that In step S2, use the supervised seed graph matching model SeedGNN, use the matched basic block unit as a seed node to perform seed graph matching; where: SeedGNN consists of multiple layers of neural networks, uses convolutional operations for node attribute propagation, and uses a filtering module to pre-match nodes with maximized similarity between two graphs for the case where seed nodes in a node are interfered by noise nodes by using the Hungarian algorithm, so as to distinguish true seed nodes from noise nodes. After masking the noise nodes, the operations of the graph neural network are performed using the true seed nodes; As the number of convolutional layers increases, new seed nodes are continuously added to the operation, and the message passing in the graph neural network deepens step by step, and finally the similarity values between each pair of nodes in the two target graph structures are obtained; the matching result of the graph is judged by setting a threshold for screening, so as to determine the matching ratio in the subgraphs of the target code snippet and the binary code snippet to be tested, and further determine whether the two code snippets are similar.

8. A fine-grained binary code similarity detection system based on code retrieval and subgraph matching, characterized in that The system includes a processing unit, and the processing unit is configured to execute: Based on fine-grained retrieval at the code snippet level, vectorize and embed the binary code to be tested at the basic block level; retrieve the embedded vectors at the basic block level in the existing binary code database, so as to perform binary code matching at the basic block level; Expand each matched basic block to its adjacent basic blocks to form a subgraph, and use the matched basic blocks as seed nodes; judge the similarity between the binary code to be tested and the code in the binary code database by performing seed graph matching on multiple discontinuous subgraphs, so as to complete fine-grained binary code similarity detection.

9. An electronic device, characterized in that, The electronic device includes a memory and a processor. When the processor executes the computer program stored in the memory, it implements a fine-grained binary code similarity detection method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, it implements a fine-grained binary code similarity detection method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Binary code similarity analysis method for vulnerability detection

    CN112733137A

  • Binary component detection method and system based on neural network and function call graph

    CN117454382A

  • Binary code similarity detection method and system, electronic equipment and medium

    CN117909753A

  • Malicious code author identification and code infringement detection method based on multi-feature fusion

    CN118709181A

  • Method and device for analyzing homology of binary code file and computer equipment

    CN119556978A

Cited By

  • Software security detection method and system based on code lines

    CN121188778A