GNN-based cross-architecture binary program similarity detection method, apparatus and device

By disassembling the binary program into LLVM IR and building a program graph, the graph neural network enhanced by global attention is used to process program graphs and instruction vectors to generate fixed-dimensional graph embedding vectors, and calculate the similarity of graph embedding vectors to evaluate similarity. This solves the problems of unified representation across architectures, insufficient semantic modeling capabilities and low efficiency of large-scale library analysis in the existing technology, and achieves efficient, robust and accurate cross-architecture binary program similarity analysis.

CN120010909AActive Publication Date: 2025-05-16XIAMEN UNIV OF TECH

Patent Information

Application Number
CN202510486738.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-05-16
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

In the analysis of binary program similarity, it is difficult to achieve cross-architecture unified program representation, insufficient high-level semantic modeling capabilities, and low large-scale program library analysis and processing efficiency.

Method used

By disassembling the binary program into a low-level virtual machine intermediate representation LLVM IR, and building a program graph based on LLVM IR, using the global attention-enhanced graph neural network GNN handles the program graph and instruction vectors, a fixed-dimensional graph embedding vector is generated, and the similarity of the graph embedding vector is calculated to evaluate the similarity.

Benefits of technology

It realizes unified representation of cross-architecture binary programs, significantly improves the ability to understand the high-level semantic features and complex logical dependencies of the program, improves the efficiency of large-scale program library analysis, and solves the problem of insufficient robustness in stripped symbolic programs and diversified compilation environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120010909A_ABST
    Figure CN120010909A_ABST
Patent Text Reader

Abstract

The invention provides a GNN-based cross-architecture binary program similarity detection method, apparatus and device, and relates to the technical field of information security processing. The method comprises the following steps: acquiring two binary programs to be detected, and disassembling the binary programs into a low-level virtual machine intermediate representation (LLVM) IR; constructing a program diagram based on LLVM IR; then inputting the program diagram into a FastText model to extract an LLVM IR instruction, and taking a corpus created based on the LLVM IR instruction as a vocabulary of the FastText model to perform multiple rounds of training so as to express an instruction mark as a word vector in a continuous vector space, and generating an instruction vector; according to the program diagram and the instruction vector, processing by using a global attention enhanced diagram neural network GNN to generate a diagram embedding vector with a fixed dimension; and calculating the similarity between the graph embedding vectors corresponding to the two binary programs to evaluate the similarity. According to the method, unified program representation of cross-framework binary programs can be realized, high-level semantic features are captured, and the processing efficiency of large-scale program library analysis is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of graph neural network and information security processing technology, and in particular to a GNN-based cross-architecture binary program similarity detection method, device and equipment. Background Art

[0002] Binary program similarity analysis plays an important role in many fields, such as vulnerability detection, patch analysis, network security, software plagiarism detection, software engineering and reverse engineering. This analysis aims to quantify the similarity between binary programs, but this task faces great challenges due to the lack of source code and high-level semantic information and the complexity of binary files.

[0003] Traditional detection methods usually rely on domain-specific feature engineering, which is not only time-consuming and labor-intensive, but also difficult to achieve universality between different hardware architectures. In recent years, natural language processing (NLP) technology has been introduced into this field to solve the cross-architecture problem by treating binary instructions as sequences. However, these methods usually have difficulty capturing high-level semantic information of programs, limiting the comprehensive understanding of complex program behaviors.

[0004] Although existing technologies have made some progress in the field of binary program similarity analysis, such as feature engineering-based extraction methods (such as graph matching, control flow graph CFG analysis) and natural language processing (NLP)-based deep learning methods (such as word embedding, Transformer-based models, etc.), they still have many limitations: (1) Feature extraction is complex and has poor versatility, making it difficult to apply in cross-architecture scenarios. (2) High-level semantic modeling capabilities are insufficient, and the global logical relationship of the program cannot be effectively captured. (3) The processing efficiency is low in large-scale program libraries and cannot meet the actual application needs.

[0005] These limitations seriously affect its effect and applicability in practical applications. In view of this, the applicant has proposed this application after studying the existing technology. Summary of the invention

[0006] The present invention aims to provide a GNN-based cross-architecture binary program similarity detection method, device and equipment to solve the shortcomings of existing methods such as difficulty in achieving unified program representation of cross-architecture binary programs, insufficient high-level semantic modeling capabilities, and low processing efficiency of large-scale program library analysis.

[0007] In order to solve the above technical problems, the present invention is implemented through the following technical solutions: A GNN-based cross-architecture binary program similarity detection method, comprising: Get the two binary programs to be tested and disassemble them into the low-level virtual machine intermediate representation LLVM IR; Build program graph based on LLVM IR; Inputting the program graph into the FastText model to extract LLVM IR instructions, and performing multiple rounds of training based on a corpus created based on the LLVM IR instructions as a vocabulary of the FastText model, so as to represent instruction tokens as word vectors in a continuous vector space, and generate instruction vectors; According to the program graph and the instruction vector, a global attention enhanced graph neural network (GNN) is used for processing to generate a graph embedding vector of fixed dimension; The similarity between the graph embedding vectors corresponding to the two binary programs is calculated to evaluate the similarity.

[0008] Preferably, the program graph is constructed based on LLVM IR, specifically: Extracting the instruction, control flow, data flow and call flow information of LLVM IR and constructing a program graph; wherein the nodes of the program graph are used to represent each instruction, and the edges are used to capture the relationship between the control flow, call flow and data flow between the instructions; Control flow edges reflect the execution order between instructions and are limited to a single function. The control relationship across functions is captured by call edges. Call flow edges are used to capture bidirectional dependencies between function calls; Data flow edges are used to represent data dependencies between instructions. The types of edges are distinguished according to the data type and encoded using one-hot encoding to distinguish various types of data dependencies.

[0009] Preferably, when extracting LLVM IR instructions, the names of global variables, local variables, constants, and functions of LLVM IR are standardized to eliminate unnecessary name differences.

[0010] Preferably, the instruction vector is obtained by summing up the word vectors of each tag in the instruction and converting them into a vector representation of fixed dimension.

[0011] Preferably, the graph neural network includes a GPS layer and an embedding layer; the GPS layer is composed of a plurality of stacked graph convolution modules, each module including a multi-head attention-based graph convolution unit and a ReLU activation function, for generating a fixed-dimensional graph embedding vector for each program graph; The embedding layer is used to encode the edge features of the program graph into vectors of the same dimension as the node features so as to incorporate the edge type information during the graph convolution process.

[0012] Preferably, the graph convolution unit combines a graph isomorphism network GINEConv with edge feature convolution and a multi-head attention mechanism to capture local semantic information and long-distance dependencies of the program graph; GINEConv updates the features of each node by aggregating the features of adjacent nodes and the corresponding edge features. The update formula of GINEConv is: ; in, is the feature of node i, and the feature vector of each node is the corresponding instruction vector; represents the updated features of node i; Represents an aggregate function; represents the regulating factor; is the set of neighbor nodes of node i; is the feature of node j, the neighbor of node i; is the edge feature of nodes i and j; is the activation function; After being processed by multiple layers of GINEConv and multi-head attention mechanisms, all node features are globally pooled to generate a graph embedding vector of fixed dimension.

[0013] Preferably, when calculating the similarity between the embedding vectors corresponding to the two binary programs, the cosine similarity is used to calculate the similarity score; if the similarity score exceeds a set threshold, the two binary programs are determined to be similar.

[0014] Preferably, it also includes using a contrast loss function for training optimization to improve the accuracy of similarity detection. The contrast loss function The expression is: ; in, represents the calculated similarity; Indicates the set similarity threshold; Indicates the label of positive and negative sample pairs. The value of 1 indicates a positive sample pair, and the value of 0 indicates a negative sample pair. max indicates the maximum value.

[0015] The present invention also provides a cross-architecture binary program similarity detection device based on GNN, comprising: The disassembly unit is used to obtain the two binary programs to be tested and disassemble them into the low-level virtual machine intermediate representation LLVM IR; A program graph construction unit, used to construct a program graph based on LLVM IR; An instruction vector generation unit, configured to input the program graph into a FastText model to extract LLVM IR instructions, and perform multiple rounds of training based on a corpus created based on the LLVM IR instructions as a vocabulary of the FastText model, so as to represent instruction tokens as word vectors in a continuous vector space, and generate instruction vectors; A graph neural network unit, configured to process the program graph and the instruction vector using a graph neural network GNN enhanced with global attention to generate a graph embedding vector of fixed dimension; The similarity calculation unit is used to calculate the similarity between the graph embedding vectors corresponding to the two binary programs to evaluate the similarity.

[0016] The present invention also provides a GNN-based cross-architecture binary program similarity detection device, including a processor and a memory, wherein the memory stores a computer program, and the computer program can be executed by the processor to implement a GNN-based cross-architecture binary program similarity detection method as described above.

[0017] The present invention also provides a computer-readable storage medium, on which computer-readable instructions are stored. When the computer-readable instructions are executed by a processor of a device where the computer-readable storage medium is located, a GNN-based cross-architecture binary program similarity detection method as described above is implemented.

[0018] In summary, compared with the prior art, the present invention has the following beneficial effects: The present invention uses LLVM IR representation and program graph construction technology to abstract a cross-architecture intermediate representation layer, shielding the differences between different hardware architectures, providing a unified expression framework for cross-architecture binary program similarity analysis, and is applicable to multiple hardware architectures, thereby realizing cross-architecture program representation.

[0019] The present invention adopts a graph neural network with global attention enhancement, combined with a multi-head attention mechanism and high-quality instruction embedding, to comprehensively capture the high-level semantic features and complex logical dependencies of the program, significantly improves the ability to understand program behavior, and solves the problem of insufficient program semantic modeling in the prior art.

[0020] The present invention optimizes the GNN model by using the Performer-based GlobalAttention mechanism through embedding generation and similarity calculation methods, maps complex program graphs into embedding vectors of fixed dimensions, significantly improves the efficiency of large-scale program library analysis, and solves the problem of low efficiency of large-scale program library analysis.

[0021] The present invention standardizes variables, constants and function names in the preprocessing stage, eliminates unnecessary name differences, and retains the original form of library function calls, thereby enhancing the model's adaptability to symbol stripping programs and diversified compilation environments, and solving the problem of insufficient robustness in symbol stripping programs and diversified compilation environments.

[0022] The present invention provides an efficient, robust and accurate cross-architecture binary program similarity analysis method, which is applicable to multiple fields such as vulnerability detection, patch analysis, network security, software plagiarism detection and reverse engineering. By combining LLVM IR representation, program graph construction, instruction vectorization processing, graph neural network modeling and similarity calculation and other technical means, a deep understanding and efficient analysis of binary programs is achieved, providing strong support for research and application in related fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0024] Figure 1 A schematic diagram of a GNN-based cross-architecture binary program similarity detection method provided in Example 1.

[0025] Figure 2 The overall architecture diagram of a GNN-based cross-architecture binary program similarity detection method provided in Example 1.

[0026] Figure 3 An example flow chart is constructed for the program diagram provided in Example 1.

[0027] Figure 4 This is a flowchart of an example of instruction standardization preprocessing provided in Example 1.

[0028] Figure 5 This is a schematic diagram of the GPS layer architecture design provided in Example 1.

[0029] Figure 6 A schematic diagram of a GNN-based cross-architecture binary program similarity detection device provided in Example 2.

[0030] The present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. DETAILED DESCRIPTION

[0031] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the invention claimed for protection, but merely represents the selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0032] Embodiment 1 Embodiment 1 of the present invention provides a GNN-based cross-architecture binary program similarity detection method, which can be implemented by a GNN-based cross-architecture binary program similarity detection device (hereinafter referred to as the similarity detection device), and in particular, is executed by one or more processors in the similarity detection device.

[0033] In this embodiment, the similarity detection device may be an electronic device equipped with a processor, the processor having a computer program of the GNN-based cross-architecture binary program similarity detection method and the computer program can be executed, such as a computer, a smart phone, a smart tablet, a workstation, etc., which is not limited here.

[0034] Although the existing technology has made some progress in the field of binary program similarity analysis, it still has many limitations. For example, the feature engineering method and call graph matching method are used to extract the function call relationship of the program to construct a function call graph, and use graph isomorphism or subgraph matching algorithms to identify the similarity between programs. For example, the control flow graph analysis method is used to parse program instructions, extract control flow relationships to generate a control flow graph (CFG), analyze its structural features, such as path length, number of branches, etc., and calculate the similarity between programs. For example, the symbolic execution and graph edit distance method is used to extract program path constraints using symbolic execution, and combine the graph edit distance to calculate the behavioral similarity of two programs.

[0035] Traditional methods mainly rely on low-level graph structure analysis (such as control flow graphs, call graphs) or path constraint calculations. These methods are difficult to capture the global semantic features of programs. At the same time, they are highly dependent on specific architecture characteristics when extracting features and lack universal design. For example, the natural language processing (NLP) method lacks high-level semantic modeling. NLP methods often treat binary instructions as simple text sequences for processing, making it difficult to capture the overall structure and logical relationship of the program; and the cross-architecture capability is insufficient. Existing NLP methods (such as jTrans) are usually optimized for specific architectures (such as x86 or ARM) and cannot effectively handle the similarity analysis of cross-architecture programs; in addition, the context expression capability is limited. For example, although the IR2Vec method enhances the expression capability of symbols and flow information, it cannot fully capture the complex interactions between the control flow, data flow, and call flow of the program.

[0036] The NLP method borrows the concept of natural language sequence processing and treats binary instructions as simple words or sentences, while ignoring the complex structured relationships between program instructions (such as control dependencies and data dependencies), resulting in limited depth and accuracy of semantic understanding.

[0037] This paper provides a cross-architecture binary program embedding generation and similarity detection analysis method based on low-level virtual machine intermediate representation (LLVM IR) and global attention enhanced graph neural network (GNN), named Binary2Vec. This method converts binary code into a unified intermediate representation and combines the structural and semantic features of the program with the graph neural network to achieve efficient, robust and accurate binary program similarity analysis.

[0038] like Figure 1-Figure 2 As shown, a GNN-based cross-architecture binary program similarity detection method includes steps S1 to S5.

[0039] S1, obtain two binary programs to be tested and disassemble them into low-level virtual machine intermediate representation LLVM IR.

[0040] like Figure 2 The overall architecture of Binary2Vec shown in the figure includes binary program disassembly, LLVM IR representation generation, program graph construction, instruction vector generation, graph neural network processing, node feature aggregation, cosine similarity calculation, and contrast loss function optimization. These steps together constitute a complete binary program similarity analysis process.

[0041] By disassembling the binary program into LLVM IR representation and building a program graph based on control flow, data flow and call flow, the program graph is modeled using a global attention enhanced graph neural network to generate a fixed-dimensional program embedding vector, thereby achieving efficient binary program similarity analysis.

[0042] First, the input binary program is converted into LLVM IR representation through a disassembly tool (such as llvm-objdump or llvm-mc). This step uses the disassembler in the LLVM tool chain to abstract the binary code of different hardware architectures into a unified intermediate representation. LLVM IR is a structured and strongly typed intermediate language that can mask the differences between different processor architectures and is suitable for a variety of hardware targets. For example, binary files generated under x86 and ARM architectures can be represented in the same LLVM IR form after disassembly, thereby achieving cross-architecture program representation consistency. This design solves the cross-architecture universality problem caused by instruction set differences in the existing technology.

[0043] S2, builds the program graph based on LLVM IR.

[0044] The construction of the program graph is one of the core steps of the present invention. After obtaining the LLVM IR, the instructions, control flow, data flow and call flow information are further extracted to construct the program graph. Figure 3 As shown, each instruction is represented as a node in the graph, and edges are used to capture the control flow, call flow, and data flow relationships between instructions.

[0045] Control flow edges reflect the execution order between instructions, such as sequential instructions are directly connected, branch instructions are linked to their branch targets, and loop conditions are connected to the start node of the loop body. It should be noted that control flow edges are limited to a single function, the control flow graph of each function forms an independent subgraph, and the control relationship across functions is captured separately by call edges.

[0046] Call flow edges are used to capture bidirectional dependencies between function calls by adding an edge from the call instruction node to the entry node of the called function and another edge from the return node of the called function to the call instruction node.

[0047] Dataflow edges represent data dependencies between instructions. The edge types are distinguished according to the data type and encoded using one-hot encoding to ensure that the program graph accurately expresses the logical structure and semantic information of the program. When constructing dataflow edges, we modify the original program (PROGRAML) graph by focusing on the output data of each instruction node A. For each instruction node A, identify its output data item B and find all instruction nodes C that use B as input (there may be multiple such nodes). For each identified node C, add a dataflow edge between A and C, the type of which is determined by the data type of B, such as integer, floating point, or other specific types.

[0048] This step reduces the number of nodes in the original PROGRAML graph, simplifies its structure, and improves the efficiency of its graph network learning. In addition, a detailed type-based data flow edge distinction is introduced and encoded using one-hot encoding. When the data produced by an instruction is not used by other instructions, the data flow edge is connected to a special "external" node. This "external" node represents an entity outside the program, indicating that instruction node A is outputting data B to an external context. This encoding clearly distinguishes various types of data dependencies, ensuring a more efficient and accurate representation of the input-output relationship between instructions while retaining the semantic information of data dependencies.

[0049] Through this construction process, the program graph can fully capture the control flow, calling relationships, and data dependencies of the program while maintaining the simplicity of the graph structure.

[0050] S3, inputting the program graph into the FastText model to extract LLVM IR instructions, and performing multiple rounds of training based on the corpus created based on the LLVM IR instructions as the vocabulary of the FastText model to represent the instruction tokens as word vectors in a continuous vector space to generate instruction vectors.

[0051] Instruction vectorization is a key step in program graph processing.

[0052] After the program graph is constructed, the instruction vector generation phase begins. The goal of this phase is to eliminate unnecessary name differences while retaining key semantic features to provide high-quality input for subsequent graph neural network learning. The instruction vector generation process is as follows: Figure 4 As shown, it is divided into two steps: standardization preprocessing and vectorization.

[0053] In the standardization preprocessing stage, the names of global variables, local variables, constants, and functions in LLVM IR are standardized to eliminate unnecessary name differences. Specifically, global variable identifiers (such as @global_var_X, such as Figure 4In the example above, @global_var_1bc48 is replaced with @global_var, eliminating any differences in global variable names or identifiers. Local variable identifiers (such as %eax or %1, such as Figure 4 %1, %2, %3, %4, %buffer) in %ID are simplified to avoid unnecessary changes in their identifiers. Constant values ​​(such as 123 or -12.34, such as Figure 4 ”42” in ) is uniformly replaced with %CONST to ensure that the comparison is based on the existence of the constant, not its specific value. Function names such as Figure 4 @function_11c4c) in is standardized as @function, so that when comparing similarity, the calling structure of the functions is compared without being affected by the specific name to preserve the key semantic features.

[0054] However, library function calls (such as Figure 4 @fdopen in ) are not subject to standardization. Library functions usually play a fixed, well-defined role in a program and are a key feature of the code. Therefore, library function calls are preserved in their original form and contribute to the program's feature set in subsequent analysis.

[0055] The standardized preprocessed LLVM IR instructions are input into the FastText model as a corpus for the vectorization processing stage.

[0056] The initialization parameters of the FastText model include vector size, window size, and epoch number. First, the FastText model is initialized, then the program graph is loaded and the corresponding LLVM IR instructions are extracted to build a corpus. Based on this corpus, the vocabulary of the model is built and multiple rounds of training are performed (e.g., setting the vector dimension d=256, the sliding window size w=5, the number of negative samples k=10, 5 initial learning rate η=0.025, and training 50 epochs using a linear decay strategy) so that the model learns to represent instruction tags as word vectors in a continuous vector space. The vector of each instruction is calculated by summing the vectors of each tag in the instruction. After the model training is completed, each instruction vector is obtained by summing the vectors of each tag in each instruction. This design eliminates the interference of variable names and constant values ​​on semantic modeling while retaining the key semantic information of the instruction. Through this process, each instruction is converted into a fixed-dimensional vector representation, providing high-quality input features for subsequent graph neural network modeling.

[0057] S4, processing the program graph and the instruction vector using a global attention enhanced graph neural network GNN to generate a graph embedding vector of fixed dimension.

[0058] Graph neural network (GNN) processing is the core technical part of this invention.

[0059] In this step, the program graph and instruction vector are passed as input to the graph neural network for processing. The graph neural network adopts a global attention-enhanced architecture to fully capture the high-level semantic features and complex logical dependencies of the program. The core component of the graph neural network is the GPS layer, and its architecture is as follows: Figure 5 As shown in Figure 2. The GPS layer combines local information propagation with the global attention mechanism to improve the model's ability to understand the graph structure and semantics. The embedding layer is used to encode edge features into vectors of the same dimension as the node features. For each edge, an embedding is generated based on its specific type so that edge type information can be included in the graph convolution process, thereby improving its ability to model relationships in the graph.

[0060] like Figure 5 The overall architecture design diagram of the GPS layer shown in the figure includes the input and output of the GPS layer, the convolution operation of the GINE layer (graph isomorphic network with edge feature convolution), the two-layer MLP perceptron, and the Performer-based GlobalAttention mechanism (i.e., a method that combines the Performer model and the global attention mechanism Global Attention). The program graph is input into the graph neural network and processed by the GPS layer to generate a fixed-dimensional graph embedding vector for each program graph. The GPS layer consists of multiple stacked graph convolution modules, each of which contains a multi-head attention-based graph convolution unit and a ReLU activation function.

[0061] The graph convolution unit combines GINEConv and multi-head attention mechanism to capture local semantic information, model node relationships and long-distance dependencies. The update formula of GINEConv updates the features of each node by aggregating the features of adjacent nodes and the corresponding edge features, thereby fully expressing the complex logical dependencies of the program graph. The update formula of GINEConv is expressed as: ; in, is the feature of node i, and the feature vector of each node is the corresponding instruction vector; represents the updated features of node i; Represents an aggregate function; represents the regulating factor; is the set of neighbor nodes of node i; is the feature of node j, the neighbor of node i; is the edge feature between nodes i and j.

[0062] In this example, the nodes in the graph network represent instructions, and the instruction vector is the feature vector of the node.

[0063] like Figure 5 As shown, The node feature vector of the layer With edge eigenvector After being updated by the GPS layer, and .

[0064] The global attention mechanism introduces the Performer-based Global Attention mechanism to capture long-distance dependencies in the graph and improve the model's ability to understand global semantics. After multiple layers of GINEConv operations and global attention mechanisms, a fixed-size graph embedding vector is generated through a global pooling operation. This design enables complex program graphs to be mapped into fixed-dimensional embedding vectors, thereby capturing their overall structure and semantics.

[0065] In addition, the global attention mechanism introduces multi-head attention calculation to further capture long-distance dependencies in the graph and improve the model's ability to understand the high-level semantic features of the program. After completing the GPS layer processing, a global pooling operation is applied to aggregate the features of all nodes in the graph into a fixed-size vector as the embedded representation of the program. This embedded vector can fully capture the overall structure and semantic features of the program, providing a basis for subsequent similarity calculations.

[0066] S5, calculate the similarity between the graph embedding vectors corresponding to the two binary programs to evaluate the similarity.

[0067] Similarity calculation and loss function optimization are the output links of the present invention.

[0068] like Figure 2 As shown, this embodiment calculates the cosine similarity between the graph embedding vectors corresponding to the two program graphs through cosine similarity to obtain a similarity score to evaluate their similarity. Of course, other similarity calculation methods can also be used, such as Siamese twin network, Euclidean distance, etc. If the similarity score exceeds the set threshold, the two binary programs are judged to be similar; otherwise, they are not similar. Cosine similarity is a commonly used vector similarity measurement method that can quickly compare the directional consistency of two vectors.

[0069] In order to guide the model to distinguish different program graphs, the contrast loss function is used as the training target to improve the accuracy of similarity detection. The expression is: ; in, represents the calculated similarity; Indicates the set similarity threshold; Indicates the label of positive and negative sample pairs. If label=1, it is a positive sample pair, and label=0, it is a negative sample pair. max means taking the maximum value.

[0070] This loss function is used to optimize the similarity between positive and negative samples. In the case of positive sample pairs, the loss function penalizes the similarity The difference from 1, that is, we hope that the similarity of positive samples is as close to 1 as possible. For negative sample pairs, the loss function will penalize samples whose similarity exceeds a certain threshold (margin) to ensure that the similarity of negative samples remains low. Specifically, the loss of negative samples will only occur when their similarity is greater than the set threshold margin, and this loss will be doubled to strengthen the penalty for negative sample pairs. This design helps the model better learn to distinguish between positive and negative samples and improve the accuracy of classification or similarity calculation.

[0071] Specifically, when two program graphs are derived from the same source code, the similarity score is high; when the graphs are derived from different source codes, the similarity score is low. Through the optimization strategy, the model can better distinguish between similar and dissimilar program graphs, thereby improving the accuracy of binary program similarity analysis.

[0072] The Binary2Vec method model of the present invention has wide applicability in practical application scenarios. For example, in the field of vulnerability detection, the binary program to be detected can be analyzed for similarity with the known vulnerability program to quickly identify potential security threats. In the patch analysis scenario, by comparing the binary programs before and after the patch, the impact of the patch on the program behavior can be quantified, thereby evaluating the effectiveness of the patch. In software plagiarism detection, the suspected plagiarized binary program can be analyzed for similarity with the original program to assist in determining whether plagiarism exists. In reverse engineering, the technical solution of the present invention can help analyze the behavior patterns of unknown programs and provide support for security research.

[0073] In order to verify the technical effect of the present invention, binary program samples of various hardware architectures (such as x86, ARM, and MIPS) were selected for experiments. The experimental results show that the present invention shields the differences between different hardware architectures through LLVM IR representation and program graph construction technology, and provides universality and consistency for cross-architecture binary program similarity analysis. The global attention-enhanced graph neural network is used, combined with a multi-head attention mechanism and high-quality instruction embedding, to fully capture the high-level semantic features and complex logical dependencies of the program, and significantly improve the ability to understand program behavior. Through embedding generation and similarity calculation methods, complex program graphs are mapped into graph embedding vectors of fixed dimensions, which significantly improves the efficiency of large-scale program library analysis. In the preprocessing stage, variable, constant, and function names are standardized to eliminate unnecessary name differences, while retaining the original form of library function calls, which enhances the model's adaptability to symbol-stripped programs and diversified compilation environments.

[0074] The present invention uses a graph neural network (GNN) with a global attention mechanism to learn graph embeddings that can effectively express program structure and semantic features. These embeddings can be used as robust representations of programs, and program similarity is evaluated by calculating the cosine similarity between embeddings. Binary2vec generates a graph structure based on the low-level virtual machine intermediate representation (LLVM IR), which is a cross-architecture intermediate representation obtained by disassembling binary programs. Each graph encodes the relationship between instructions (such as control flow, data flow, and call flow) through edge features, and uses instruction embeddings generated by word embedding technology as node features. This design enables Binary2vec to capture both low-level structural features and high-level semantic information of programs, thereby adapting to different architectures and application scenarios. The Binary2vec framework combines advanced concepts of graph learning and natural language processing, effectively improving the ability to understand and compare binary programs, and provides a scalable and efficient solution for cross-architecture binary analysis.

[0075] In summary, compared with the prior art, the present invention has the following beneficial effects: The present invention solves the problem of inconsistent cross-architecture program representation by using LLVM IR representation and program graph construction technology. Specifically, the binary program is disassembled into an architecture-independent intermediate representation, and the program graph is constructed by parsing the control flow, data flow, and call flow relationships between instructions. This representation method shields the differences between different hardware architectures and provides consistency and versatility for cross-architecture binary program similarity analysis.

[0076] This paper adopts a global attention-enhanced graph neural network (GNN) to solve the problem of insufficient program semantic modeling in the existing technology. By combining the graph representation of control flow, data flow and call flow, and introducing a multi-head attention mechanism and high-quality instruction embedding, the high-level semantic features and complex logical dependencies of the program are fully captured, significantly improving the ability to understand program behavior.

[0077] The present invention solves the problem of low efficiency in large-scale program library analysis through embedding generation and similarity calculation methods. Specifically, complex program graphs are mapped into fixed-dimensional embedding vectors, and program similarity is quickly evaluated by calculating the cosine similarity between embedding vectors. At the same time, the Performer-based Global Attention mechanism is used to optimize the GNN model, enabling it to efficiently process large amounts of graph data and achieve efficient program analysis.

[0078] The present invention solves the problem of insufficient robustness in symbolic stripping programs and diversified compilation environments through standardized processing and diversified training. In preprocessing, variable, constant and function names are standardized to eliminate unnecessary name differences; at the same time, the original form of library function calls is retained to retain key semantic features. In addition, the use of diversified data sets for training enhances the generalization ability of the model, making it more stable in different optimization levels and compilation environments.

[0079] In summary, the present invention provides an efficient, robust and accurate cross-architecture binary program similarity analysis method, which is applicable to multiple fields such as vulnerability detection, patch analysis, network security, software plagiarism detection and reverse engineering. By combining LLVM IR representation, program graph construction, instruction vectorization processing, graph neural network modeling and similarity calculation and other technical means, a deep understanding and efficient analysis of binary programs is achieved, providing strong support for research and application in related fields.

[0080] Embodiment 2 like Figure 6 As shown, the second embodiment of the present invention further provides a cross-architecture binary program similarity detection device based on GNN, comprising: The disassembly unit is used to obtain the two binary programs to be tested and disassemble them into the low-level virtual machine intermediate representation LLVM IR; A program graph construction unit, used to construct a program graph based on LLVM IR; An instruction vector generation unit, configured to input the program graph into a FastText model to extract LLVM IR instructions, and perform multiple rounds of training based on a corpus created based on the LLVM IR instructions as a vocabulary of the FastText model, so as to represent instruction tokens as word vectors in a continuous vector space, and generate instruction vectors; A graph neural network unit, configured to process the program graph and the instruction vector using a graph neural network GNN enhanced with global attention to generate a graph embedding vector of fixed dimension; The similarity calculation unit is used to calculate the similarity between the graph embedding vectors corresponding to the two binary programs to evaluate the similarity.

[0081] Embodiment 3 The third embodiment of the present invention also provides a GNN-based cross-architecture binary program similarity detection device, which includes a memory and a processor, wherein the memory stores a computer program, and the computer program can be executed by the processor to implement the GNN-based cross-architecture binary program similarity detection method as described above.

[0082] Embodiment 4 The fourth embodiment of the present invention also provides a computer-readable storage medium, on which computer-readable storage medium is stored computer-readable instructions. When the computer-readable instructions are executed by a processor of a device where the computer-readable storage medium is located, the cross-architecture binary program similarity detection method based on GNN as described above is implemented.

[0083] In several embodiments provided in the embodiments of the present invention, it should be understood that the disclosed apparatus and method can also be implemented in other ways. The apparatus and method embodiments described above are merely schematic. For example, the flowcharts in the accompanying drawings show the possible architecture, functions and operations of the apparatus, method and computer program product according to multiple embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of a code, and the module, program segment or a part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.

[0084] In addition, the functional modules in the various embodiments of the present invention may be integrated together to form an independent part, or each module may exist independently, or two or more modules may be integrated to form an independent part.

[0085] If the function is implemented in the form of a software function module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, an electronic device, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program code. It should be noted that in this article, the term "include", "include" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements includes not only those elements, but also includes other elements that are not explicitly listed, or also includes elements inherent to such process, method, article or device. Without more constraints, an element defined by the phrase "comprising a..." does not exclude the existence of other identical elements in the process, method, article or apparatus comprising the element.

[0086] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms "a", "said" and "the" used in the embodiments of the present invention and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings.

[0087] It should be understood that the term "and / or" used in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.

[0088] The word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)", depending on the context.

[0089] The "first\second" mentioned in the embodiments is only to distinguish similar objects, and does not represent a specific order for the objects. It is understandable that the "first\second" can be interchanged with the specific order or sequence where permitted. It should be understood that the objects distinguished by "first\second" can be interchanged where appropriate, so that the embodiments described herein can be implemented in an order other than those illustrated or described herein.

[0090] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A cross-architecture binary program similarity detection method based on GNN, characterized in that: include: Get the two binary programs to be tested and disassemble them into the low-level virtual machine intermediate representation LLVM IR; Build program graph based on LLVM IR; Inputting the program graph into the FastText model to extract LLVM IR instructions, and performing multiple rounds of training based on a corpus created based on the LLVM IR instructions as a vocabulary of the FastText model, so as to represent instruction tokens as word vectors in a continuous vector space, and generate instruction vectors; According to the program graph and the instruction vector, a global attention enhanced graph neural network (GNN) is used for processing to generate a graph embedding vector of fixed dimension; The similarity between the graph embedding vectors corresponding to the two binary programs is calculated to evaluate the similarity.

2. According to claim 1, a GNN-based cross-architecture binary program similarity detection method is characterized in that ,Build the program graph based on LLVM IR, specifically: Extracting the instruction, control flow, data flow and call flow information of LLVM IR and constructing a program graph; wherein the nodes of the program graph are used to represent each instruction, and the edges are used to capture the relationship between the control flow, call flow and data flow between the instructions; Control flow edges reflect the execution order between instructions and are limited to a single function. The control relationship across functions is captured by call edges. Call flow edges are used to capture bidirectional dependencies between function calls; Data flow edges are used to represent data dependencies between instructions. The types of edges are distinguished according to the data type and encoded using one-hot encoding to distinguish various types of data dependencies.

3. According to claim 1, a GNN-based cross-architecture binary program similarity detection method is characterized in that ,When extracting LLVM IR instructions, the global variables, local variables, constants, and function names of LLVM IR are standardized to eliminate unnecessary name differences.

4. According to claim 1, a GNN-based cross-architecture binary program similarity detection method is characterized in that ,The instruction vector is obtained by summing the word vectors of each token in the instruction and converting it into a ,vector representation of fixed dimension.

5. According to claim 1, a GNN-based cross-architecture binary program similarity detection method is characterized in that ,The graph neural network includes a GPS layer and an embedding layer; the GPS layer is composed of a number of stacked graph convolution modules, each of which contains a multi-head attention-based graph convolution unit and a ReLU activation function, which is used to generate a fixed-dimensional graph embedding vector for each program graph; The embedding layer is used to encode the edge features of the program graph into vectors of the same dimension as the node features so as to incorporate the edge type information during the graph convolution process.

6. A GNN-based cross-architecture binary program similarity detection method according to claim 5, characterized in that ,The graph convolution unit combines the graph isomorphism network GINEConv with edge feature convolution and multi-head attention mechanism to capture the local semantic information and long-distance dependencies of the program graph; GINEConv updates the features of each node by aggregating the features of adjacent nodes and the corresponding edge features. The update formula of GINEConv is: ; in, is the feature of node i, and the feature vector of each node is the corresponding instruction vector; represents the updated features of node i; Represents an aggregate function; represents the regulating factor; is the set of neighbor nodes of node i; is the feature of node j, the neighbor of node i; is the edge feature of nodes i and j; is the activation function; After being processed by multiple layers of GINEConv and multi-head attention mechanisms, all node features are globally pooled to generate a graph embedding vector of fixed dimension.

7. A cross-architecture binary program similarity detection method based on GNN according to claim 1, characterized in that ,When calculating the similarity between the embedding vectors corresponding to two binary programs, the ,cosine similarity is used to obtain the similarity score; if the similarity score exceeds the ,set threshold, the two binary programs are judged to be similar.

8. A cross-architecture binary program similarity detection method based on GNN according to claim 1, characterized in that , also includes the use of contrast loss function for training optimization to improve the accuracy of similarity detection, contrast loss function The expression is: ; in, represents the calculated similarity; Indicates the set similarity threshold; Indicates the label of positive and negative sample pairs. The value of 1 indicates a positive sample pair, and the value of 0 indicates a negative sample pair. max indicates the maximum value.

9. A cross-architecture binary program similarity detection device based on GNN, characterized in that: include: The disassembly unit is used to obtain the two binary programs to be tested and disassemble them into the low-level virtual machine intermediate representation LLVM IR; A program graph construction unit, used to construct a program graph based on LLVM IR; An instruction vector generation unit, configured to input the program graph into a FastText model to extract LLVM IR instructions, and perform multiple rounds of training based on a corpus created based on the LLVM IR instructions as a vocabulary of the FastText model, so as to represent instruction tokens as word vectors in a continuous vector space, and generate instruction vectors; A graph neural network unit, configured to process the program graph and the instruction vector using a graph neural network GNN enhanced with global attention to generate a graph embedding vector of fixed dimension; The similarity calculation unit is used to calculate the similarity between the graph embedding vectors corresponding to the two binary programs to evaluate the similarity.

10. A cross-architecture binary program similarity detection device based on GNN, characterized in that: It includes a processor and a memory, wherein the memory stores a computer program, and the computer program can be executed by the processor to implement a GNN-based cross-architecture binary program similarity detection method as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Binary code similarity comparison technology capable of resisting compilation difference

    CN113010209A

  • Method and system for carrying out similarity detection on cross-architecture binary codes

    CN117034281A

  • Cross-architecture binary code similarity detection method, system, equipment and medium

    CN118885827A

  • Methods and systems for identifying binary code vulnerability

    US20240419806A1

  • Binary Code Similarity Detection System Based on Hard Sample-aware Momentum Contrastive Learning

    US20250110715A1

Cited By

  • Binary program similarity analysis method and system

    CN121187640A

  • Binary program similarity analysis method and system

    CN121187640B