GNN-based Cross-architecture Binary Program Similarity Detection Method, Device and Equipment
By disassembling the binary program into LLVM IR and building a program graph, the graph neural network GNN with global attention-enhanced graph embedding vector is used to solve the problems of unified representation, semantic modeling and low processing efficiency in cross-architecture binary program similarity analysis, and efficient and robust binary program similarity analysis is achieved.
Patent Information
- Application Number
- CN202510486738.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-04-18
AI Technical Summary
In the analysis of binary program similarity, it is difficult to achieve cross-architecture unified program representation, insufficient high-level semantic modeling capabilities, and low processing efficiency of large-scale program library analysis.
By disassembling the binary program into a low-level virtual machine intermediate representation LLVM IR, and building a program graph based on LLVM IR, using the global attention-enhanced graph neural network GNN for processing, generating fixed-dimensional graph embedding vectors, and computing the similarity between graph embedding vectors to evaluate the similarity.
It realizes unified representation of cross-architecture binary programs, significantly improves the ability to understand program behavior, improves the efficiency of large-scale program library analysis, and solves the problems of insufficient semantic modeling and low processing efficiency in the existing technology.
Smart Images

Figure CN120010909B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of graph neural networks and information security processing. Specifically, it relates to a method, device, and equipment for cross-architecture binary program similarity detection based on GNN. Background Art
[0002] Binary program similarity analysis plays an important role in many fields such as vulnerability detection, patch analysis, network security, software plagiarism detection, software engineering, and reverse engineering. This analysis aims to quantify the similarity between binary programs, but due to the lack of source code and high-level semantic information, as well as the complexity of binary files, this task faces great challenges.
[0003] Traditional detection methods usually rely on domain-specific feature engineering, which is not only time-consuming and laborious but also difficult to achieve generality between different hardware architectures. In recent years, natural language processing (NLP) technologies have been introduced into this field to solve the cross-architecture problem by treating binary instructions as sequences. However, these methods usually have difficulty capturing the high-level semantic information of programs, limiting the comprehensive understanding of complex program behaviors.
[0004] Although existing technologies have made certain progress in the field of binary program similarity analysis, such as extraction methods based on feature engineering (such as graph matching, control flow graph CFG analysis), and deep learning methods based on natural language processing (NLP) (such as word embedding, based on Transformer models, etc.), there are still limitations in many aspects: (1) Feature extraction is complex and has poor generality, making it difficult to apply in cross-architecture scenarios. (2) The high-level semantic modeling ability is insufficient, and the global logical relationships of programs cannot be effectively captured. (3) The processing efficiency is low in large-scale program libraries and cannot meet the actual application requirements.
[0005] These limitations seriously affect their effectiveness and applicability in practical applications. In view of this, the applicant has specifically proposed this application after studying the existing technologies. Summary of the Invention
[0006] The present invention aims to provide a method, device, and equipment for cross-architecture binary program similarity detection based on GNN to solve the shortcomings of existing methods, such as difficulty in achieving unified program representation for cross-architecture binary programs, insufficient high-level semantic modeling ability, and low processing efficiency in large-scale program library analysis.
[0007] To solve the above technical problems, the present invention is achieved through the following technical solutions:
[0008] A method for cross-architecture binary program similarity detection based on GNN, comprising:
[0009] Obtain two binary programs to be detected and disassemble them into the low-level virtual machine intermediate representation LLVM IR;
[0010] Construct a program graph based on the LLVM IR;
[0011] Input the program graph into the FastText model to extract LLVM IR instructions, and use the corpus created based on the LLVM IR instructions as the vocabulary of the FastText model for multiple rounds of training to represent the instructions as word vectors in a continuous vector space and generate instruction vectors;
[0012] Process according to the program graph and the instruction vectors using a graph neural network GNN enhanced by global attention to generate a graph embedding vector of a fixed dimension;
[0013] Calculate the similarity between the graph embedding vectors corresponding to the two binary programs to evaluate the similarity.
[0014] Preferably, constructing a program graph based on the LLVM IR is specifically as follows:
[0015] Extract the instruction, control flow, data flow, and call flow information of the LLVM IR to construct a program graph; wherein, the nodes of the program graph are used to represent each instruction, and the edges are used to capture the relationships of the control flow, call flow, and data flow between instructions;
[0016] The control flow edges reflect the execution order between instructions and are limited to a single function. The cross-function control relationships are captured through call edges;
[0017] The call flow edges are used to capture the bidirectional dependency relationships between function calls;
[0018] The data flow edges are used to represent the data dependency relationships between instructions, distinguish the types of edges according to the data types, and use one-hot encoding for encoding to distinguish various types of data dependency relationships.
[0019] Preferably, when extracting LLVM IR instructions, perform standardization operations on the global variables, local variables, constants, and function names of the LLVM IR to eliminate unnecessary name differences.
[0020] Preferably, the instruction vector is obtained by summing the word vectors of each token in the instruction and converting it into a vector representation of a fixed dimension.
[0021] Preferably, the graph neural network includes a GPS layer and an embedding layer; the GPS layer is composed of multiple stacked graph convolution modules, and each module contains a graph convolution unit based on multi-head attention and a ReLU activation function, which is used to generate a graph embedding vector of a fixed dimension for each program graph;
[0022] The embedding layer is used to encode the edge features of the program graph into vectors with the same dimension as the node features, so as to incorporate edge type information during the graph convolution process.
[0023] Preferably, the graph convolution unit combines the Graph Isomorphism Network with Edge Features Convolution (GINEConv) and the multi-head attention mechanism to capture the local semantic information and long-range dependencies of the program graph;
[0024] GINEConv updates the features of each node by aggregating the features of adjacent nodes and the corresponding edge features. The update formula of GINEConv is:
[0025] ;
[0026] Where, is the feature of node i, and the feature vector of each node is the corresponding instruction vector; represents the updated feature of node i; represents the aggregation function; represents the adjustment factor; is the set of neighbor nodes of node i; is the feature of neighbor node j of node i; is the edge feature between nodes i and j; is the activation function;
[0027] After being processed by multiple layers of GINEConv and the multi-head attention mechanism, a global pooling operation is performed on all node features to generate a graph embedding vector with a fixed dimension.
[0028] Preferably, when calculating the similarity between the embedding vectors corresponding to two binary programs, the cosine similarity is used to obtain the similarity score; if the similarity score exceeds the set threshold, it is determined that these two binary programs are similar.
[0029] Preferably, it also includes training optimization using a contrast loss function to improve the accuracy of similarity detection. The contrast loss function has the following expression:
[0030] ;
[0031] Where, represents the calculated similarity; represents the set similarity threshold; represents the label of the positive and negative sample pairs, with a value of 1 indicating a positive sample pair and a value of 0 indicating a negative sample pair; max represents taking the maximum value.
[0032] The present invention also provides a cross-architecture binary program similarity detection device based on GNN, including:
[0033] A disassembly unit, configured to obtain two binary programs to be detected and disassemble them into Low-Level Virtual Machine Intermediate Representation (LLVM IR);
[0034] A program graph construction unit, configured to construct a program graph based on the LLVM IR;
[0035] An instruction vector generation unit, configured to input the program graph into a FastText model to extract LLVM IR instructions, and use the corpus created based on the LLVM IR instructions as the vocabulary of the FastText model for multiple rounds of training, so as to represent the instructions as word vectors in a continuous vector space and generate instruction vectors;
[0036] A graph neural network unit, configured to process according to the program graph and the instruction vectors by using a graph neural network (GNN) enhanced with global attention to generate a graph embedding vector with a fixed dimension;
[0037] A similarity calculation unit, configured to calculate the similarity between the graph embedding vectors corresponding to the two binary programs to evaluate the similarity.
[0038] The present invention also provides a cross-architecture binary program similarity detection device based on GNN, including a processor and a memory. A computer program is stored in the memory and can be executed by the processor to implement a cross-architecture binary program similarity detection method as described above.
[0039] The present invention also provides a computer-readable storage medium. Computer-readable instructions are stored on the computer-readable storage medium. When the computer-readable instructions are executed by a processor of a device where the computer-readable storage medium is located, a cross-architecture binary program similarity detection method as described above is implemented.
[0040] In summary, compared with the prior art, the present invention has the following beneficial effects:
[0041] Through the LLVM IR representation and program graph construction technology, the present invention abstracts an intermediate representation layer across architectures, shields the differences between different hardware architectures, provides a unified expression framework for cross-architecture binary program similarity analysis, is applicable to multiple hardware architectures, and thus realizes cross-architecture program representation.
[0042] The present invention adopts a graph neural network enhanced with global attention, combines the multi-head attention mechanism and high-quality instruction embeddings, comprehensively captures the high-level semantic features and complex logical dependencies of the program, significantly improves the ability to understand program behavior, and solves the problem of insufficient program semantic modeling in the prior art.
[0043] The present invention optimizes the GNN model by using the Performer-based GlobalAttention mechanism through embedding generation and similarity calculation methods, maps complex program graphs into embedding vectors of fixed dimensions, significantly improves the efficiency of large-scale program library analysis, and solves the problem of low efficiency in large-scale program library analysis.
[0044] In the preprocessing stage of the present invention, variables, constants, and function names are standardized, unnecessary name differences are eliminated, and at the same time, the original form of library function calls is retained, enhancing the adaptability of the model to stripped symbolic programs and diverse compilation environments, and solving the problem of insufficient robustness in stripped symbolic programs and diverse compilation environments.
[0045] The present invention provides an efficient, robust, and accurate cross-architecture binary program similarity analysis method, which is applicable to multiple fields such as vulnerability detection, patch analysis, network security, software plagiarism detection, and reverse engineering. By combining technical means such as LLVM IR representation, program graph construction, instruction vectorization processing, graph neural network modeling, and similarity calculation, in-depth understanding and efficient analysis of binary programs are achieved, providing strong support for research and applications in related fields. Brief Description of the Drawings
[0046] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0047] Figure 1 Schematic diagram of a cross-architecture binary program similarity detection method based on GNN provided for Embodiment 1.
[0048] Figure 2 Overall architecture diagram of a cross-architecture binary program similarity detection method based on GNN provided for Embodiment 1.
[0049] Figure 3 Flowchart of an example of program graph construction provided for Embodiment 1.
[0050] Figure 4 Flowchart of an example of instruction standardization preprocessing provided for Embodiment 1.
[0051] Figure 5 Schematic diagram of the GPS layer architecture design provided for Embodiment 1.
[0052] Figure 6Schematic diagram of a cross-architecture binary program similarity detection device provided for Embodiment 2.
[0053] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Specific embodiments
[0054] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention. Therefore, the detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0055] Embodiment 1
[0056] Embodiment 1 of the present invention provides a cross-architecture binary program similarity detection method based on a GNN, which can be implemented by a cross-architecture binary program similarity detection device (hereinafter referred to as the similarity detection device) based on a GNN. Specifically, it is executed by one or more processors in the similarity detection device.
[0057] In this embodiment, the similarity detection device may be an electronic device equipped with a processor. The processor has a computer program of the cross-architecture binary program similarity detection method based on a GNN and the computer program can be executed, such as a computer, a smart phone, a smart tablet, a workstation, etc., which is not limited herein.
[0058] Although certain progress has been made in the field of binary program similarity analysis in the prior art, there are still limitations in many aspects. For example, the method based on feature engineering is adopted, the graph matching method is called, the function call relationship of the program is extracted to construct a function call graph, and the graph isomorphism or subgraph matching algorithm is used to identify the similarity between programs. For example, the control flow graph analysis method is adopted. By parsing program instructions, the control flow relationship is extracted to generate a control flow graph (Control Flow Graph, CFG), and its structural features, such as path length, number of branches, etc., are analyzed to calculate the similarity between programs. For example, the symbolic execution and graph edit distance method is adopted. The symbolic execution is used to extract the program path constraints, and the graph edit distance is combined to calculate the behavioral similarity of two programs.
[0059] Traditional methods mainly rely on low-level graph structure analysis (such as control flow graphs, call graphs) or path constraint calculations. These methods are difficult to capture the global semantic features of programs. At the same time, they highly depend on specific architecture features during feature extraction and lack general design. For example, when using natural language processing (NLP) methods, these methods lack high-level semantic modeling. NLP methods mostly treat binary instructions as simple text sequences for processing, making it difficult to capture the overall structure and logical relationships of programs. Moreover, they have insufficient cross-architecture capabilities. Existing NLP methods (such as jTrans) are usually optimized for specific architectures (such as x86 or ARM) and cannot effectively handle similarity analysis of cross-architecture programs. In addition, their context expression ability is limited. For example, although the IR2Vec method enhances the expression ability of symbolic and flow information, it cannot comprehensively capture the complex interactions between the control flow, data flow, and call flow of programs.
[0060] NLP methods draw on the concept of natural language sequence processing, treating binary instructions as simple words or sentences while ignoring the complex structural relationships (such as control dependencies and data dependencies) between program instructions, resulting in limited depth and accuracy of semantic understanding.
[0061] The present invention provides a cross-architecture binary program embedding generation and similarity detection analysis method based on low-level virtual machine intermediate representation (LLVM IR) and global attention enhanced graph neural network (GNN), named Binary2Vec. This method converts binary code into a unified intermediate representation form and combines graph neural networks to learn the structural and semantic features of programs, thereby achieving efficient, robust, and accurate binary program similarity analysis.
[0062] As Figure 1 - Figure 2 shown, a cross-architecture binary program similarity detection method based on GNN includes steps S1 to S5.
[0063] S1, obtain two binary programs to be detected and disassemble them into low-level virtual machine intermediate representation LLVM IR.
[0064] As Figure 2 shown in the overall architecture of Binary2Vec, it includes disassembly processing of binary programs, generation of LLVM IR representations, construction of program graphs, generation of instruction vectors, processing of graph neural networks, node feature aggregation, calculation of cosine similarity, and optimization of contrast loss functions. These steps together constitute a complete binary program similarity analysis process.
[0065] By disassembling a binary program into LLVM IR representation and constructing a program graph based on control flow, data flow, and call flow, a graph neural network enhanced with global attention is used to model the program graph, generating a program embedding vector with a fixed dimension, thereby achieving efficient binary program similarity analysis.
[0066] First, the input binary program is converted into LLVM IR representation through a disassembler (such as llvm-objdump or llvm-mc). This step utilizes the disassembler in the LLVM toolchain to abstract binary code of different hardware architectures into a unified intermediate representation form. LLVM IR is a structured and strongly typed intermediate language that can shield the differences between different processor architectures and is applicable to multiple hardware targets. For example, binary files generated under the x86 and ARM architectures can be represented in the same LLVM IR form after disassembly, thus achieving cross-architecture program representation consistency. This design solves the cross-architecture generality problem caused by instruction set differences in the prior art.
[0067] S2. Construct a program graph based on the LLVM IR.
[0068] The construction of the program graph is one of the core steps of the present invention. After obtaining the LLVM IR, information such as instructions, control flow, data flow, and call flow is further extracted to construct the program graph. As Figure 3 shown, each instruction is represented as a node in the graph, and the edges are used to capture the control flow, call flow, and data flow relationships between instructions.
[0069] The control flow edges reflect the execution order between instructions. For example, sequential instructions are directly connected, branch instructions are linked to their branch targets, and loop conditions are connected to the starting node of the loop body. It should be noted that the control flow edges are limited to a single function, and the control flow graph of each function forms an independent subgraph, while the cross-function control relationships are captured separately through call edges.
[0070] The call flow edges are used to capture the bidirectional dependency relationships between function calls. An edge is added from the call instruction node to the entry node of the called function, and another edge is added from the return node of the called function to the call instruction node.
[0071] Data flow edges represent the data dependencies between instructions. The types of edges are distinguished according to the data types and encoded using one-hot encoding to ensure that the program graph accurately represents the logical structure and semantic information of the program. When constructing the data flow edges, we modify the original program (PROGRAML) graph by focusing on the output data of each instruction node A. For each instruction node A, identify its output data item B and find all instruction nodes C (there may be multiple such nodes) that use B as input. For each identified node C, add a data flow edge between A and C, and the type is determined by the data type of B, such as integer, floating point, or other specific types.
[0072] This step reduces the number of nodes in the original PROGRAML graph, simplifies its structure, and improves the learning efficiency of its graph network. In addition, a detailed type-based distinction of data flow edges is introduced and encoded using one-hot encoding. When the data produced by an instruction is not used by other instructions, the data flow edge is connected to a special "external" node. This "external" node represents an entity outside the program, indicating that instruction node A is outputting data B to the external context. This encoding clearly distinguishes various types of data dependencies, ensuring a more efficient and accurate representation of the input-output relationships between instructions while preserving the semantic information of the data dependencies.
[0073] Through this construction process, the program graph can comprehensively capture the control flow, call relationships, and data dependencies of the program while maintaining the simplicity of the graph structure.
[0074] S3. Input the program graph into the FastText model to extract LLVM IR instructions, and use the corpus created based on the LLVM IR instructions as the vocabulary of the FastText model for multiple rounds of training to represent the instruction tokens as word vectors in a continuous vector space and generate instruction vectors.
[0075] Instruction vectorization is a key link in program graph processing.
[0076] After the construction of the program graph is completed, it enters the stage of generating instruction vectors. The goal of this stage is to eliminate unnecessary name differences while preserving the key semantic features, providing high-quality input for the subsequent learning of graph neural networks. The process of generating instruction vectors is as Figure 4 shown, which is divided into two steps: normalization preprocessing and vectorization.
[0077] In the normalization preprocessing stage, normalization operations are performed on global variables, local variables, constants, and function names in the LLVM IR to eliminate unnecessary name differences. Specifically, global variable identifiers (such as @global_var_X, as Figure 4The @global_var_1bc48) in is uniformly replaced with @global_var to eliminate any differences in global variable names or identifiers. Local variable identifiers (such as %eax or %1, such as Figure 4 The %1, %2, %3, %4, %buffer) in is simplified to %ID to avoid unnecessary changes in its identifiers. Constant values (such as 123 or -12.34, such as Figure 4 The "42" in is uniformly replaced with %CONST to ensure that the comparison is based on the existence of the constant rather than its specific value. Function names (such as Figure 4 The @function_11c4c) in is standardized to @function, so that when comparing similarities, the call structure of the function can be compared without being affected by the specific name, in order to retain the key semantic features.
[0078] However, library function calls (such as Figure 4 The @fdopen) in is not subject to standardization. Library functions usually play a fixed and clear role in the program and are key features of the code. Therefore, library function calls retain their original form and contribute to the function set of the program in subsequent analyses.
[0079] The standardized preprocessed LLVM IR instructions are input into the FastText model as a corpus for the vectorization process.
[0080] The initialization parameters of the FastText model include the vector size, window size, and number of epochs. First, the FastText model is initialized, then the program graph is loaded and the corresponding LLVM IR instructions are extracted to build a corpus. Based on this corpus, the vocabulary of the model is built and trained for multiple rounds (for example, setting the vector dimension d = 256, the sliding window size w = 5, the negative sampling number k = 10, the initial learning rate η = 0.025, and training for 50 epochs using a linear decay strategy), so that the model learns to represent instruction tokens as word vectors in a continuous vector space. The vector of each instruction is calculated by summing the vectors of each token in the instruction. After the model training is completed, the vector of each instruction is obtained by summing the vectors of each token in each instruction. This design eliminates the interference of variable names and constant values on semantic modeling while retaining the key semantic information of the instructions. Through this process, each instruction is converted into a vector representation of a fixed dimension, providing high-quality input features for subsequent graph neural network modeling.
[0081] S4. According to the program graph and the instruction vectors, use a graph neural network GNN enhanced by global attention for processing to generate a graph embedding vector of a fixed dimension.
[0082] Graph Neural Network (GNN) processing is the core technical part of the present invention.
[0083] In this step, the program graph and the instruction vector are passed as inputs to the graph neural network for processing. The graph neural network adopts an architecture enhanced by global attention to comprehensively capture the high-level semantic features and complex logical dependencies of the program. The core component of the graph neural network is the GPS layer, and its architecture is as Figure 5 shown. The GPS layer combines local information propagation and global attention mechanism to improve the model's understanding ability of graph structure and semantics. The embedding layer is used to encode edge features into vectors with the same dimension as node features. For each edge, an embedding is generated according to its specific type to incorporate edge type information during graph convolution, thereby improving its ability to model relationships in the graph.
[0084] As Figure 5 shown in the overall architecture design schematic diagram of the GPS layer, it includes the input and output of the GPS layer, the convolution operation of the GINE layer (graph isomorphism network with edge feature convolution), two-layer MLP perceptron, and the Performer-based GlobalAttention mechanism (that is, the method combining the Performer model and the global attention mechanism Global Attention). The program graph is input into the graph neural network and processed through the GPS layer to generate a graph embedding vector with a fixed dimension for each program graph. The GPS layer consists of multiple stacked graph convolution modules, and each module contains a graph convolution unit based on multi-head attention and a ReLU activation function.
[0085] The graph convolution unit combines GINEConv and multi-head attention mechanism to capture local semantic information, model node relationships, and long-distance dependencies. The update formula of GINEConv updates the features of each node by aggregating the features of adjacent nodes and the corresponding edge features, thereby fully expressing the complex logical dependencies of the program graph. The update formula of GINEConv is expressed as:
[0086] ;
[0087] where is the feature of node i, and the feature vector of each node is the corresponding instruction vector; represents the updated feature of node i; represents the aggregation function; represents the adjustment factor; is the set of neighbor nodes of node i; is the feature of neighbor node j of node i; is the edge feature between node i and j.
[0088] In this example, the nodes in the graph network represent instructions, and the instruction vector is the feature vector of the nodes.
[0089] As Figure 5 shown, the node feature vector of the th layer and the edge feature vector are updated by the GPS layer and become and .
[0090] The global attention mechanism introduces the Performer-based Global Attention mechanism to capture long-range dependencies in the graph and improve the model's ability to understand global semantics. After being processed by multiple layers of GINEConv operations and the global attention mechanism, a fixed-size graph embedding vector is generated through global pooling operations. This design enables complex program graphs to be mapped into fixed-dimensional embedding vectors, thereby capturing their overall structure and semantics.
[0091] In addition, the global attention mechanism introduces multi-head attention calculation to further capture long-range dependencies in the graph and improve the model's ability to understand high-level semantic features of the program. After completing the GPS layer processing, global pooling operations are applied to aggregate the features of all nodes in the graph into a fixed-size vector as the embedding representation of the program. This embedding vector can comprehensively capture the overall structure and semantic features of the program, providing a basis for subsequent similarity calculation.
[0092] S5. Calculate the similarity between the graph embedding vectors corresponding to two binary programs to evaluate the similarity.
[0093] Similarity calculation and loss function optimization are the output links of the present invention.
[0094] As Figure 2 shown, in this embodiment, the cosine similarity between the graph embedding vectors corresponding to two program graphs is calculated through cosine similarity to obtain a similarity score to evaluate their similarity. Of course, other similarity calculation methods can also be used, such as Siamese twin networks, Euclidean distance, etc. If the similarity score exceeds the set threshold, it is determined that the two binary programs are similar; otherwise, they are not. Cosine similarity is a commonly used vector similarity measurement method that can quickly compare the direction consistency of two vectors.
[0095] To guide the model to distinguish different program graphs, a contrastive loss function is used as the training objective to improve the accuracy of similarity detection. The expression of the contrastive loss function is:
[0096] ;
[0097] Among them, represents the calculated similarity; represents the set similarity threshold; represents the label of positive and negative sample pairs. If label = 1, it is a positive sample pair; if label = 0, it is a negative sample pair; max means taking the maximum value.
[0098] This loss function is used to optimize the similarity between positive and negative samples. In the case of positive sample pairs, the loss function will penalize the difference between the similarity and 1, that is, it is hoped that the similarity of positive samples is as close to 1 as possible. For negative sample pairs, the loss function will penalize those samples whose similarity exceeds a certain threshold (i.e., margin), ensuring that the similarity of negative samples remains low. Specifically, the loss of negative samples will only occur when their similarity is greater than the set threshold margin, and this loss will be doubled to strengthen the penalty for negative sample pairs. This design helps the model better learn to distinguish positive and negative samples and improve the accuracy of classification or similarity calculation.
[0099] Specifically, when two program graphs are derived from the same source code, the similarity score is high; when the graphs come from different source codes, the similarity score is low. Through the optimization strategy, the model can better distinguish similar and dissimilar program graphs, thereby improving the accuracy of binary program similarity analysis.
[0100] The Binary2Vec method model of the present invention has wide applicability in actual application scenarios. For example, in the field of vulnerability detection, the binary program to be detected can be subjected to similarity analysis with known vulnerable programs to quickly identify potential security threats. In the patch analysis scenario, by comparing the binary programs before and after patching, the impact of the patch on the program behavior can be quantified, thereby evaluating the effectiveness of the patch. In software plagiarism detection, the binary program suspected of plagiarism can be subjected to similarity analysis with the original program to assist in determining whether there is a plagiarism behavior. In reverse engineering, the technical solution of the present invention can help analyze the behavior patterns of unknown programs and provide support for security research.
[0101] To verify the technical effects of the present invention, binary program samples of various hardware architectures (such as x86, ARM, MIPS) were selected for experiments. The experimental results show that the present invention, through the LLVM IR representation and program graph construction technology, shields the differences between different hardware architectures, and provides generality and consistency for cross-architecture binary program similarity analysis. By using a graph neural network (GNN) with global attention mechanism, combined with multi-head attention mechanism and high-quality instruction embeddings, it comprehensively captures the high-level semantic features and complex logical dependencies of the program, significantly improving the ability to understand program behavior. Through the embedding generation and similarity calculation methods, the complex program graph is mapped into a graph embedding vector of a fixed dimension, significantly improving the efficiency of large-scale program library analysis. Standardizing variable, constant, and function names in the preprocessing stage eliminates unnecessary name differences, while retaining the original form of library function calls, enhancing the model's adaptability to stripped symbol programs and diverse compilation environments.
[0102] The present invention utilizes a graph neural network (GNN) with a global attention mechanism to learn graph embeddings that can effectively represent program structure and semantic features. These embeddings can serve as a robust representation of the program, and program similarity is evaluated by calculating the cosine similarity between the embeddings. Binary2vec generates a graph structure based on the low-level virtual machine intermediate representation (LLVM IR), which is a cross-architecture intermediate representation obtained by disassembling binary programs. Each graph encodes the relationships between instructions (such as control flow, data flow, and call flow) through edge features, and uses instruction embeddings generated by word embedding technology as node features. This design enables Binary2vec to capture both the low-level structural features and high-level semantic information of the program, thus adapting to different architectures and application scenarios. The Binary2vec framework combines advanced concepts of graph learning and natural language processing, effectively enhancing the ability to understand and compare binary programs, and providing a scalable and efficient solution for cross-architecture binary analysis.
[0103] In summary, compared with the prior art, the present invention has the following beneficial effects:
[0104] The present invention solves the problem of inconsistent cross-architecture program representation by using the LLVM IR representation and program graph construction technology. Specifically, the binary program is disassembled into an architecture-independent intermediate representation, and a program graph is constructed by parsing the control flow, data flow, and call flow relationships between instructions. This representation shields the differences between different hardware architectures and provides consistency and generality for cross-architecture binary program similarity analysis.
[0105] The present invention uses a graph neural network (GNN) enhanced with global attention to solve the problem of insufficient program semantic modeling in the prior art. By combining the graph representations of control flow, data flow, and call flow, and introducing a multi-head attention mechanism and high-quality instruction embeddings, it comprehensively captures the high-level semantic features and complex logical dependencies of the program, significantly improving the ability to understand program behavior.
[0106] The present invention solves the problem of low efficiency in analyzing large-scale program libraries through an embedding generation and similarity calculation method. Specifically, it maps the complex program graph into an embedding vector of a fixed dimension, and quickly evaluates program similarity by calculating the cosine similarity between the embedding vectors. At the same time, it uses the Performer-based Global Attention mechanism to optimize the GNN model, enabling it to efficiently process a large amount of graph data and achieve efficient program analysis.
[0107] The present invention solves the problem of insufficient robustness in stripped symbolic programs and diverse compilation environments through normalization processing and diverse training. In preprocessing, it normalizes variables, constants, and function names to eliminate unnecessary name differences; at the same time, it retains the original form of library function calls to retain key semantic features. In addition, it uses a diverse dataset for training to enhance the generalization ability of the model, making it more stable in different optimization levels and compilation environments.
[0108] In summary, the present invention provides an efficient, robust, and accurate cross-architecture binary program similarity analysis method, which is applicable to multiple fields such as vulnerability detection, patch analysis, network security, software plagiarism detection, and reverse engineering. By combining technical means such as LLVM IR representation, program graph construction, instruction vectorization processing, graph neural network modeling, and similarity calculation, it realizes in-depth understanding and efficient analysis of binary programs, providing strong support for research and applications in related fields.
[0109] Embodiment 2
[0110] As Figure 6 shown, the second embodiment of the present invention also provides a cross-architecture binary program similarity detection device based on GNN, including:
[0111] A disassembly unit for obtaining two binary programs to be detected and disassembling them into low-level virtual machine intermediate representation LLVM IR;
[0112] A program graph construction unit for constructing a program graph based on LLVM IR;
[0113] An instruction vector generation unit, configured to input the program graph into a FastText model to extract LLVM IR instructions, and use the corpus created based on the LLVM IR instructions as the vocabulary of the FastText model for multiple rounds of training, so as to represent the instruction tokens as word vectors in a continuous vector space and generate instruction vectors;
[0114] A graph neural network unit, configured to process according to the program graph and the instruction vectors by using a graph neural network GNN enhanced with global attention to generate a graph embedding vector with a fixed dimension;
[0115] A similarity calculation unit, configured to calculate the similarity between the graph embedding vectors corresponding to two binary programs to evaluate the similarity.
[0116] Embodiment III
[0117] The third embodiment of the present invention further provides a GNN-based cross-architecture binary program similarity detection device, which includes a memory and a processor. A computer program is stored in the memory, and the computer program can be executed by the processor to implement the GNN-based cross-architecture binary program similarity detection method as described above.
[0118] Embodiment IV
[0119] The fourth embodiment of the present invention further provides a computer-readable storage medium. Computer-readable instructions are stored on the computer-readable storage medium. When the computer-readable instructions are executed by the processor of the device where the computer-readable storage medium is located, the GNN-based cross-architecture binary program similarity detection method as described above is implemented.
[0120] In several embodiments provided by the embodiments of the present invention, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device and method embodiments described above are merely illustrative. For example, the flowcharts in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the part of the module, program segment, or code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0121] In addition, each functional module in various embodiments of the present invention can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.
[0122] If the above functions are implemented in the form of software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, an electronic device, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes. It should be noted that in this article, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device including the said element.
[0123] The terms used in the embodiments of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The singular forms "a", "the", and "said" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise.
[0124] It should be understood that the term "and / or" used herein is merely a description of an associated relationship between associated objects, indicating that three relationships may exist. For example, A and / or B may represent: the sole existence of A, the simultaneous existence of A and B, and the sole existence of B. Additionally, the character " / " herein generally represents an "or" relationship between the associated objects before and after.
[0125] Depending on the context, the word "if" as used herein can be interpreted as "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detected (stated condition or event)" can be interpreted as "when determined" or "in response to determining" or "when detected (stated condition or event)" or "in response to detecting (stated condition or event)".
[0126] The "first / second" mentioned in the embodiments is merely to distinguish similar objects and does not represent a specific order for the objects. It can be understood that the "first / second" can be interchanged in their specific order or sequence when permitted. It should be understood that the objects distinguished by the "first / second" can be interchanged appropriately so that the embodiments described herein can be implemented in an order other than those illustrated or described herein.
[0127] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, various modifications and variations can be made to the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A cross-architecture binary program similarity detection method based on GNN, characterized in that: include: Get the two binary programs to be tested and disassemble them into the low-level virtual machine intermediate representation LLVM IR; Build a program graph based on LLVM IR, specifically: Extracting the instruction, control flow, data flow and call flow information of LLVM IR and constructing a program graph; wherein the nodes of the program graph are used to represent each instruction, and the edges are used to capture the relationship between the control flow, call flow and data flow between the instructions; Control flow edges reflect the execution order between instructions and are limited to a single function. The control relationship across functions is captured by call edges. Call flow edges are used to capture bidirectional dependencies between function calls; Data flow edges are used to represent data dependencies between instructions. The edge types are distinguished according to the data type and encoded using one-hot encoding to distinguish various types of data dependencies. Inputting the program graph into the FastText model to extract LLVM IR instructions, and performing multiple rounds of training based on a corpus created based on the LLVM IR instructions as a vocabulary of the FastText model, so as to represent instruction tokens as word vectors in a continuous vector space, and generate instruction vectors; According to the program graph and the instruction vector, a global attention enhanced graph neural network GNN is used for processing to generate a fixed-dimensional graph embedding vector; wherein the graph neural network includes a GPS layer and an embedding layer; the GPS layer is composed of a plurality of stacked graph convolution modules, each module including a multi-head attention-based graph convolution unit and a ReLU activation function, for generating a fixed-dimensional graph embedding vector for each program graph; The embedding layer is used to encode the edge features of the program graph into a vector of the same dimension as the node features so as to incorporate the edge type information in the graph convolution process; The graph convolution unit combines the graph isomorphism network GINEConv with edge feature convolution and a multi-head attention mechanism to capture the local semantic information and long-distance dependencies of the program graph; GINEConv updates the features of each node by aggregating the features of adjacent nodes and the corresponding edge features. The update formula of GINEConv is: ; in, is the feature of node i, and the feature vector of each node is the corresponding instruction vector; represents the updated features of node i; Represents an aggregate function; represents the regulating factor; is the set of neighbor nodes of node i; is the feature of node j, the neighbor of node i; is the edge feature of nodes i and j; is the activation function; After being processed by multiple layers of GINEConv and multi-head attention mechanism, all node features are globally pooled to generate a graph embedding vector of fixed dimension; The similarity between the graph embedding vectors corresponding to the two binary programs is calculated to evaluate the similarity.
2. According to claim 1, a GNN-based cross-architecture binary program similarity detection method is characterized in that ,When LLVM IR instructions are extracted, the names of global variables, local variables, constants, and functions of LLVM IR are standardized to eliminate unnecessary name differences.
3. According to claim 1, a GNN-based cross-architecture binary program similarity detection method is characterized in that ,The instruction vector is obtained by summing the word vectors of each token in the instruction and converting it into a ,vector representation of fixed dimension.
4. According to claim 1, a GNN-based cross-architecture binary program similarity detection method is characterized in that ,When calculating the similarity between the embedding vectors corresponding to two binary programs, the ,cosine similarity is used to obtain the similarity score; if the similarity score exceeds the ,set threshold, the two binary programs are judged to be similar.
5. According to claim 1, a GNN-based cross-architecture binary program similarity detection method is characterized in that , also includes the use of contrast loss function for training optimization to improve the accuracy of similarity detection, contrast loss function The expression is: ; in, represents the calculated similarity; Indicates the set similarity threshold; Indicates the label of positive and negative sample pairs. The value of 1 indicates a positive sample pair, and the value of 0 indicates a negative sample pair. max indicates the maximum value.
6. A GNN-based cross-architecture binary program similarity detection device, used to implement a GNN-based cross-architecture binary program similarity detection method as described in any one of claims 1-5, characterized in that: include: The disassembly unit is used to obtain the two binary programs to be tested and disassemble them into the low-level virtual machine intermediate representation LLVM IR; A program graph construction unit, used to construct a program graph based on LLVM IR; An instruction vector generation unit, configured to input the program graph into a FastText model to extract LLVM IR instructions, and perform multiple rounds of training based on a corpus created based on the LLVM IR instructions as a vocabulary of the FastText model, so as to represent instruction tokens as word vectors in a continuous vector space, and generate instruction vectors; A graph neural network unit, configured to process the program graph and the instruction vector using a graph neural network GNN enhanced with global attention to generate a graph embedding vector of fixed dimension; The similarity calculation unit is used to calculate the similarity between the graph embedding vectors corresponding to the two binary programs to evaluate the similarity.
7. A cross-architecture binary program similarity detection device based on GNN, characterized in that: It includes a processor and a memory, wherein a computer program is stored in the memory, and the computer program can be executed by the processor to implement a GNN-based cross-architecture binary program similarity detection method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Binary code similarity comparison technology capable of resisting compilation difference
CN113010209A
Method and system for carrying out similarity detection on cross-architecture binary codes
CN117034281A