Binary code similarity detection method for eliminating false alarm problem

The control flow graph feature vector of binary code is constructed through graph alignment method, which solves the problems of high false alarm rate and low detection capability in the prior art, and achieves higher search accuracy and detection efficiency.

CN119939261AActive Publication Date: 2025-05-06HANGZHOU DIANZI UNIV +1

Patent Information

Application Number
CN202411713910.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-27
Publication Date
2025-05-06
Estimated Expiration
2044-11-27

AI Technical Summary

Technical Problem

The existing binary code similarity detection methods have high false positive rates and low detection capabilities due to "embedding conflicts", the influence of compilers and optimizers, and the loss of semantic information.

Method used

The graph alignment method is used to obtain the control flow graph of the binary code through reverse analysis, construct structure and attribute feature vectors, establish a similarity matrix, and obtain the node vector representation through low-rank approximation and SVD matrix decomposition, and calculate the similarity score between the binary codes.

Benefits of technology

Effectively eliminate false positives, improve search accuracy, reduce detection time cost, and improve detection efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939261A_ABST
    Figure CN119939261A_ABST
Patent Text Reader

Abstract

The invention discloses a binary code similarity detection method for eliminating a false alarm problem, and belongs to the field of static software vulnerability detection and analysis. The method comprises the following steps: extracting control flow graph information of binary codes by utilizing a reverse analysis tool, constructing structural feature vectors and attribute feature vectors of all nodes, and generating a first similarity matrix based on the structural feature vectors and the attribute feature vectors; approximately decomposing the first similarity matrix through a low-rank matrix to obtain vector representation of each node; and quickly aligning graph nodes, converting distance measurement between vector representations of the nodes into a second similarity score, and converting the second similarity score into a final similarity score of the two binary codes. According to the method, the false alarm problem caused by'embedded collision 'in binary code similarity detection is eliminated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of static software vulnerability detection and analysis, and in particular to a binary code similarity detection method for eliminating false positive problems. Background Art

[0002] Binary code similarity detection has been used as the core foundation for various security and software engineering applications, including malware clustering, code clone detection, and vulnerability review. Current mainstream binary code similarity detection methods mainly rely on neural networks (NNs) to extract structural level information, such as control flow and data flow graphs. However, neural network-based techniques capture lexical level, control structure level, or data flow level information of binary code for representation learning, which is usually too coarse-grained and largely loses semantic information, resulting in a large number of false positives in its retrieved top-k candidates. In addition, it also exhibits low detection ability and stability for various challenging settings, such as cross-optimization levels and obfuscation.

[0003] There are many challenges in binary code similarity detection:

[0004] (1) "Embedding conflict" problem: Since the READOUT function in most binary code similarity detection methods based on neural networks is a summation function, this will cause some logically dissimilar binary code fragments to have similar sums of different node information embeddings. As a result, mainstream binary code similarity detection methods cannot be applied to eliminate such false positives.

[0005] (2) The compiler, compilation options, and optimizer will have a significant impact on the final generated binary code, resulting in different binary codes generated in different environments even if the source code is the same, thus affecting similarity detection.

[0006] (3) The neural network-based model loses a lot of semantic information when extracting the semantics of binary codes.

[0007] Graph alignment can be used to eliminate false positives caused by "embedding collisions". However, there is currently a lack of methods for equivalence testing using graph alignment in the context of binary codes. It is important and necessary to develop relevant analysis methods and analysis systems. Summary of the invention

[0008] Aiming at the problems in the existing methods, the present invention provides a low-cost and comprehensive graph alignment method to eliminate false positives in binary code similarity detection.

[0009] To achieve the above purpose, the technical solution adopted by the present invention is as follows:

[0010] A binary code similarity detection method for eliminating false positive problems comprises the following steps:

[0011] S1: Perform reverse analysis on the binary code to be detected to obtain the control flow graph of the binary code;

[0012] S2: Based on the structure of the control flow graph, construct a structural feature vector for all nodes in the control flow graph;

[0013] S3: Extract the semantic information of the control flow graph and construct attribute feature vectors for all nodes in the control flow graph based on the semantic information;

[0014] S4: construct a similarity matrix between control flow graphs based on the structural feature vectors and attribute feature vectors;

[0015] S5: Calculate the similarity scores between binary codes based on the similarity matrix, and perform binary code similarity detection based on the similarity scores.

[0016] Furthermore, the control flow graph includes basic blocks in the binary code and execution order information thereof.

[0017] Furthermore, the step S2 is specifically as follows:

[0018] S2.1: construct an undirected graph based on the structure of the control flow graph, wherein the weights of the edges in the undirected graph are all 1;

[0019] S2.2: For each node in an undirected graph, take all other nodes in the same unedged graph and assign discount factors to the taken nodes: Where k is the distance between the node and the current node;

[0020] S2.3: For each node in the undirected graph, calculate the number of neighbor nodes at different distances, and use the discount factor as the weight to construct the structural feature vector of the node:

[0021] d s [k] = len(neighbors) × δ k

[0022] Among them, d s [k] is the kth dimension of the structural feature vector, and len(neighbors) is the number of nodes that are at distance k from the current node.

[0023] Furthermore, the step S4 is specifically as follows:

[0024] S4.1: p nodes are selected from the control flow graphs of the two binary codes to be detected as marker nodes, and the first node similarity between each node and the p marker nodes is calculated based on the structural feature vector and the attribute feature vector to form an n×p first similarity matrix C, where n is the sum of the number of nodes in the control flow graphs of the two binary codes to be detected;

[0025] S4.2: According to the first similarity matrix C, a vector representation of each node is obtained by low-rank approximation and SVD matrix decomposition;

[0026] S4.3: Calculate the cosine similarity value between the vector representations of the nodes of each control flow graph as the second similarity between the nodes, and construct an n×n second similarity matrix.

[0027] Furthermore, the similarity score between binary codes is solved based on the similarity matrix in step S5, specifically: for each row of the second similarity matrix, the maximum value in a row of elements is calculated as the row maximum value, and the average of all row maximum values ​​is taken; for each column of the second similarity matrix, the maximum value in a column of elements is calculated as the column maximum value, and the average of all column maximum values ​​is taken; the average of the row maximum values ​​and the average of the column maximum values ​​are then averaged, and the final average value is the similarity score between the binary codes.

[0028] Compared with the prior art, the present invention has the following advantages:

[0029] 1. The present invention innovatively introduces graph alignment technology into binary code similarity detection, thereby overcoming the false positive problem caused by "embedding collision" in the existing binary code similarity detection technology, thereby achieving a higher search accuracy.

[0030] 2. The present invention adopts The method performs low-rank approximation and SVD matrix decomposition to represent the node information of the binary code control flow graph in the form of a low-dimensional vector, so that the similarity scores between binary codes can be obtained more efficiently, thereby overcoming the problem of high time cost of similarity detection in the prior art and achieving higher detection efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 The overall structure diagram of a low-cost and comprehensive graph alignment method for eliminating false positives in binary code similarity detection;

[0032] Figure 2 A schematic diagram for obtaining structural and attribute feature vectors from a control flow graph;

[0033] Figure 3 To construct a similarity matrix and use Methods: Matrix low-rank approximation and decomposition are performed to obtain a node vector representation schematic diagram;

[0034] Figure 4 For fast node alignment, the similarity scores of the two graphs are calculated and judged according to the set threshold. DETAILED DESCRIPTION

[0035] The present invention is further described and illustrated below in conjunction with specific embodiments. The embodiments are merely exemplary of the present disclosure and do not define the scope of limitation. The technical features of each embodiment of the present invention may be combined accordingly without conflicting with each other.

[0036] like Figure 1 As shown, the method proposed in the present invention is mainly divided into three parts: structure and attribute embedding vector extraction, similarity matrix construction and node alignment.

[0037] The entire workflow of structural feature and attribute feature vector extraction is as follows: Figure 2 As shown, the following steps are included:

[0038] S1: Perform reverse analysis on the binary code to be detected to obtain the control flow graph of the binary code.

[0039] S1.1: Collect control flow graph (CFG) information using reverse analysis tools: For binary codes (functions or program blocks) that need to be tested for similarity, use reverse analysis tools (such as IDA Pro, Ghidra, or Radare2) to disassemble the target binary file, distinguish each segment of binary code into multiple basic blocks according to jump instructions, and obtain the jump relationship between each basic block. Use these basic blocks as nodes to establish a control flow graph (CFG) for each segment of binary code to represent the execution path of instructions in the program. Preprocess the CFG to generate a unique identifier for the basic block to facilitate subsequent graph alignment and similarity calculation.

[0040] S1.2: Use NetworkX to construct an undirected graph to represent the control flow graph information. The connections between nodes in the undirected graph are the same as those in the control flow graph, but the direction of the edges is removed and the weights of all edges are assigned to 1.

[0041] S2: Based on the structure of the control flow graph, construct a structural feature vector for all nodes in the control flow graph, such as Figure 2 shown.

[0042] S2.1: Use a recursive method to construct a k-hop neighborhood for each node. The k-hop neighborhood is: all nodes that are k steps away from the target node. Specifically, for each node in an undirected graph, first find its direct neighbor node as a 1-hop neighborhood. The direct neighbor node is a node that is directly adjacent to the target node. Use a dynamic data structure to continuously track the neighbor nodes that have been processed, and recursively calculate each hop neighborhood of the current node until the farthest node is reached. In each hop, a discount factor is assigned to each neighbor node, and the discount factor is calculated based on the distance. Neighbor nodes at different distances from the current node are assigned discount factors of different sizes δ1, δ2, δ3,…, δ k ,…,in Represents the discount factor of the neighboring nodes that are k steps away from the current node.

[0043] S2.2: Construct degree distribution sequence: In the k-hop neighborhood of each node, calculate the degree distribution of all neighbor nodes, and weight the neighbor degrees according to the discount factor to construct the structural feature vector d of the node s :

[0044] d s [k] = len(neighbors) × δ k

[0045] Among them, d s [j] is the kth dimension of the structural feature vector. The number of dimensions of the structural feature vector is not less than the longest path length of the undirected graph. len(neighbors) is the number of neighbor nodes in the k-hop neighborhood of the node.

[0046] The structural feature vectors of each node are superimposed to finally generate a degree distribution matrix, in which each row represents a node and contains the neighborhood information of the node and its corresponding degree distribution characteristics.

[0047] S3: Construct attribute feature vectors for all nodes in the control flow graph based on the semantic information of the control flow graph.

[0048] In this embodiment, the NLP model is used to extract the semantic information of the binary code. The code content of the functions or basic blocks in the control flow graph and the connection relationship between them are used as the input text of the model.

[0049] S3.1: Pre-training stage: Use the Bert model to pre-train the semantic information of the binary code. The Bert model is an NLP model based on the Transformer architecture. In the masked language model task (MLM), the tags are masked to simulate semantic filling in the real context, and all adjacent nodes are identified for pre-training of the adjacent node prediction task (ANP). Two graph-level tasks, the intra-block graph task and the graph classification task, are added to deeply mine the semantic information and optimize the block representation learning. Intra-block graph task (BIG): By judging whether the sampled blocks belong to the same graph, the intra-graph correlation learning of different blocks is realized. Graph classification task (GC): Based on the different platforms / optimization levels to which the blocks belong, the graphs are classified to realize the attribution learning of the platform / optimization level.

[0050] S3.2: Attribute feature extraction: Use the control flow graph (CFG) parsing tool to obtain the information of each control flow block. Use the trained NLP model to process the block embedding vector, and use the block embedding vector as the attribute feature vector of each graph node.

[0051] S4: Constructing a similarity matrix between control flow graphs based on the structural feature vector and the attribute feature vector, including the following steps:

[0052] S4.1: Construct the first similarity matrix from all nodes in the two graphs to the "marked" node.

[0053] (1) The Euclidean distance of the structural feature vector and the Euclidean distance of the attribute feature vector are weighted and added to obtain the total distance. The total distance is converted into the first similarity score between 0 and 1 using an exponential function. The specific formula here is:

[0054] sim(u,v)=exp[-γ s ·dist(d u , d v )-γ a ·dist(f u , f v )]

[0055] Among them, γ s , γ a is the weight, d u d v is the structural feature vector of node u and node v, f u 、f v is the attribute feature vector of node u and node v, and dist(·) represents the calculation of Euclidean distance.

[0056] (2) Select a small number of landmark nodes. Here, the landmark nodes should be selected from the two graphs in proportion. The similarity of each node is only compared with a small number of p landmark nodes, where p is much smaller than the total number of nodes n in the two control flow graphs. Calculate the first similarity between the p landmark nodes and all n nodes to form an n×p first similarity matrix C.

[0057] S4.2: If Figure 3 As shown in , the decomposition representation of the low-rank approximation matrix is ​​constructed, and the similarity of each node to these landmark nodes is used to construct the representation of the node. The specific approach is: The method performs low-rank approximation and decomposes the high-dimensional kernel function to obtain the approximate low-rank part, that is, in is the pseudo-inverse matrix of the p×p similarity matrix W. The similarity matrix W is composed of the first node similarities between p marker nodes. Perform SVD decomposition, the specific formula is Where U and V are orthogonal matrices, Σ is a diagonal matrix, and then by the formula Get an n×p matrix Y, where each row of the matrix Y is the vector representation of each node;

[0058] S4.3: If Figure 4 Perform fast graph node alignment as shown.

[0059] The cosine similarity value of the vector representation between two nodes is calculated as the second similarity. A similarity matrix is ​​constructed based on the second similarity between two nodes. The element in the i-th row and j-th column of the matrix is ​​the second similarity score between the i-th node in one control flow graph and the j-th node in another control flow graph, thereby forming a second similarity matrix.

[0060] S5: Extract the second similarity matrix information and calculate the final similarity scores of the two images.

[0061] For each row of the matrix, calculate the maximum value of its elements. This aims to capture the strongest alignment signal in each row, reflecting the maximum degree of matching in the row direction. Similarly, for each column of the matrix, calculate the maximum value of its elements. This operation reflects the maximum degree of matching in the column direction. Calculate the average value of the maximum value set of the above rows and the average value of the maximum value set of the columns. Then take the average of the two and finally get the similarity score of the two images.

[0062] S6: Determine whether the two binary codes are similar based on the final score.

[0063] The judgment in this step can be implemented based on a preset score threshold or ranking threshold. If based on a preset score threshold, the final similarity score greater than the set threshold is considered similar, and the score below the threshold is considered dissimilar. For binary code retrieval scenarios, similarity can be determined based on a preset ranking threshold. Calculate the final similarity score of the query code snippet and all code snippets in the database according to the above method, sort all code snippets involved in the retrieval from large to small according to the final similarity score, the first 10 code snippets are considered similar to the query code snippet, and the rest are considered dissimilar.

[0064] In a specific implementation of the present invention, it may also include: using part of the graph nodes to perform faster graph alignment, for example, the first 70% of the nodes may be selected for alignment to achieve a system speed improvement.

[0065] The embodiments described above provide a detailed description of the technical solutions and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, supplements and equivalent substitutions made within the scope of the principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A binary code similarity detection method for eliminating false positive problems, characterized in that: The following steps are involved: S1: Perform reverse analysis on the binary code to be detected to obtain the control flow graph of the binary code; S2: Based on the structure of the control flow graph, construct a structural feature vector for all nodes in the control flow graph; S3: Extract the semantic information of the control flow graph and construct attribute feature vectors for all nodes in the control flow graph based on the semantic information; S4: construct a similarity matrix between control flow graphs based on the structural feature vectors and attribute feature vectors; S5: Calculate the similarity scores between binary codes based on the similarity matrix, and perform binary code similarity detection based on the similarity scores.

2. The binary code similarity detection method for eliminating false positives according to claim 1 is characterized in that: The control flow graph includes basic blocks in binary code and execution order information thereof.

3. The binary code similarity detection method for eliminating false positives according to claim 1, characterized in that: The step S2 is specifically as follows: S2.1: construct an undirected graph based on the structure of the control flow graph, wherein the weights of the edges in the undirected graph are all 1; S2.2: For each node in an undirected graph, take all other nodes in the same unedged graph and assign discount factors to the taken nodes: Where k is the distance between the node and the current node; S2.3: For each node in the undirected graph, calculate the number of neighbor nodes at different distances, and use the discount factor as the weight to construct the structural feature vector of the node: d s [k]=len(neighbors)×δ k Among them, d s [k] is the kth dimension of the structural feature vector, and len(neighbors) is the number of nodes that are at distance k from the current node.

4. The binary code similarity detection method for eliminating false positives according to claim 2, characterized in that: The step S3 specifically includes: inputting the control flow graph into the NLP model, extracting multiple block embedding vectors through the NLP model, and using the block embedding vectors as attribute feature vectors of corresponding nodes.

5. The binary code similarity detection method for eliminating false positives according to claim 1, characterized in that: The step S4 is specifically as follows: S4.1: p nodes are selected from the control flow graphs of the two binary codes to be detected as marker nodes, and the first node similarity between each node and the p marker nodes is calculated based on the structural feature vector and the attribute feature vector to form an n×p first similarity matrix C, where n is the sum of the number of nodes in the control flow graphs of the two binary codes to be detected; S4.2: According to the first similarity matrix C, a vector representation of each node is obtained by low-rank approximation and SVD matrix decomposition; S4.3: Calculate the cosine similarity value between the vector representations of the nodes of each control flow graph as the second similarity between the nodes, and construct an n×n second similarity matrix.

6. The binary code similarity detection method for eliminating false positives according to claim 5, characterized in that: The step S4.2 is specifically as follows: based on the first node similarities between the p marker nodes, construct a p×p similarity matrix W, and calculate the pseudo-inverse matrix of the similarity matrix W: Then the pseudo-inverse matrix By SVD decomposition, we can get Where U and V are orthogonal matrices, Σ is a diagonal matrix, and then the p×n vector representation matrix is ​​obtained based on the decomposed orthogonal matrix and diagonal matrix. Each row of the vector representation matrix Y is the vector representation of each node.

7. The binary code similarity detection method for eliminating false positives according to claim 5, characterized in that: The calculation method of the first node similarity is: sim(u,v)=exp[-γ s ·dist(d u ,d v )-γ a ·dist(f u ,f v )] Among them, γ s , γ a is the weight, d u d v is the structural feature vector of node u and node v, f u 、f v is the attribute feature vector of node u and node v, and dist(·) represents the calculation of Euclidean distance.

8. The binary code similarity detection method for eliminating false positives according to claim 5, characterized in that: The similarity scores between binary codes are solved based on the similarity matrix in step S5, specifically: For each row of the second similarity matrix, the maximum value in a row of elements is calculated as the row maximum value, and the average of all row maximum values ​​is taken; for each column of the second similarity matrix, the maximum value in a column of elements is calculated as the column maximum value, and the average of all column maximum values ​​is taken; the average of the row maximum values ​​and the average of the column maximum values ​​are then averaged, and the final average value is the similarity score between the binary codes.

9. The binary code similarity detection method for eliminating false positives according to claim 1, characterized in that: The binary code similarity detection based on the similarity score is specifically as follows: Binary code similarity detection is performed based on a preset score threshold or based on a ranking threshold.

Citation Information

Patent Citations

  • Binary file code search detection method and system based on tensor operation

    CN110688150A

  • Deep-learning-based binary code similarity detection method

    CN113554101A

  • Binary code similarity detection method and device and electronic equipment

    CN117951543A

  • Cross-architecture binary code similarity detection method, system, equipment and medium

    CN118885827A

Cited By

  • FPGA code similarity detection method and device based on weighted directed graph

    CN120670323A

  • FPGA Code Similarity Detection Method and Apparatus Based on Weighted Directed Graph

    CN120670323B