Multi-feature fusion binary code similarity detection method based on large language model
Through the multi-feature fusion method of large language model and GraphCNN combined with HNSW and Siamese networks, the robustness and efficiency of binary code similarity detection are solved, and efficient and accurate detection of binary code is achieved.
Patent Information
- Application Number
- CN202510369219.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-11
AI Technical Summary
Existing binary code similarity detection methods are poorly robust in the face of code obfuscation and compiler optimization, and are inefficient in large-scale database retrieval, making it difficult to meet the real-time detection requirements.
The large language model is used to extract semantic features, combine GraphCNN to generate structural feature embedding, and efficient search is performed through the HNSW algorithm. Finally, the Siamese network is used for fine comparison to realize binary code similarity detection of multi-feature fusion.
Improves the accuracy and robustness of binary code similarity detection, and can keenly identify subtle differences caused by compilation optimization or code obfuscation, maintaining efficient computing performance.
Smart Images

Figure FDA0005330863910000022 
Figure FDA0005330863910000023 
Figure FDA0005330863910000031
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of information security and relates to a multi-feature fusion binary code similarity detection method based on a large language model. Background Art
[0002] In recent years, with the increasing demand for in-depth analysis of binary code in fields such as software reverse engineering, malware detection, intellectual property protection, and network security, binary code similarity detection has become one of the hot topics in the research community. In these application scenarios, binary code similarity detection plays an irreplaceable role, especially in identifying malicious code variants, combating software piracy, monitoring code injection attacks, and vulnerability mining, significantly improving the accuracy and efficiency of security protection and auditing.
[0003] However, the parsing of binary code faces many problems. Traditional binary similarity detection methods mostly rely on static analysis techniques and manually extracted features, such as symbol matching, instruction sequence comparison, and simple control flow graph matching. However, these methods often have many deficiencies: on the one hand, manually extracted features are difficult to comprehensively capture the deep semantics and complex control flow structures of the code, resulting in poor robustness when facing code obfuscation and compiler optimization; on the other hand, traditional methods are inefficient in large-scale database retrieval and difficult to meet the requirements of real-time detection.
[0004] The present invention proposes a multi-feature fusion binary code similarity detection method based on a large language model, aiming to effectively detect similar binary code segments. This method uses a large language model fine-tuned by prompt engineering and a GraphCNN graph convolutional neural network to bidirectionally capture the subtle syntax and global structure information of functions. The Hierarchical Navigable Small World Graphs (HNSW) algorithm used in the vector library search process also greatly improves the efficiency of large-scale similarity retrieval. At the same time, the fused Siamese network improves the accuracy of binary code similarity detection, has higher robustness and accuracy compared with traditional methods, and can better cope with the complex challenges faced in modern binary code analysis. Summary of the Invention
[0005] The purpose of the present invention is to provide a multi-feature fusion binary code similarity detection method based on a large language model to solve the limitations of the existing binary similarity detection methods described above.
[0006] To achieve the above purpose, the present invention provides the following solutions:
[0007] The present invention provides a multi-feature fusion binary code similarity detection method based on a large language model, including the following steps:
[0008] S1: Use IDAPro to perform reverse engineering on the binary file, extract the function features of the binary file and perform preprocessing to obtain the assembly code of each function and the CFG control flow graph;
[0009] S2: Put the obtained assembly code corresponding to the binary function into a large language model that has been fine-tuned in advance using prompt engineering to generate its semantic feature embedding, and at the same time put the CFG of the function into the GraphCNN graph convolutional neural network to generate the structural feature embedding;
[0010] S3: Connect the semantic feature embedding and the structural feature embedding to obtain a synthetic embedding representing this binary function. Include the embedding of the query function, and store the generated synthetic embeddings corresponding to all functions in the local offline vector database;
[0011] S4: Apply the HNSW algorithm to the synthetic embedding of the query function to find the top K most similar vectors in the local offline vector database;
[0012] S5: Finally, pass this query function embedding and the top K vectors into the Siamese network for comparison to find the truly most similar binary function among them.
[0013] Further, in S1, performing reverse engineering on the binary file and performing preprocessing specifically includes:
[0014] Use IDAPro to perform reverse engineering on the binary file and extract the function features of the binary file. In this process, IDAPro splits out each function to obtain the assembly language code corresponding to each function, and extracts key information such as the hash value of the function through preprocessing operations. At the same time, organize and unify the specific information of these functions to construct the CFG control flow graph of each function;
[0015] Among them, the key information of the function extraction includes the hash value sHash, the offset sea, the function name name, the number of nodes node_size, and the number of edges edge_size. Then use the networkx module in python to create an empty graph g, and use the obtained number of function nodes node_size and the number of edges edge_size to construct the control flow graph of the function in the empty graph g. At the same time, store the hash value sHash, the offset sea, and the function name name in the control flow graph information for subsequent data processing and extraction.
[0016] Further, in S2, generating the semantic feature embedding and generating the structural feature embedding specifically includes:
[0017] Input the preprocessed function assembly code into a large language model fine-tuned by prompt engineering to generate feature embeddings that can accurately reflect the semantic information of the function. At the same time, input the CFG corresponding to the function into a GraphCNN (Graph Convolutional Neural Network), and generate structural feature embeddings through graph structure learning. This parallel process can not only capture the subtle syntactic features contained in the code but also obtain the complex control flow structure information inside the function, achieving efficient encoding of the binary code in both semantic and structural dimensions;
[0018] To accurately construct the semantic feature vector of a binary function, we used prompt engineering to fine-tune the large language model and designed a prompt template for specific tasks of assembly code: "Please extract the semantic feature vector that describes the function logic and key information based on the following assembly code: {assembly code}" to clearly guide the model to focus on the semantic details in the code and thus construct the semantic feature vector of the function;
[0019] During the fine-tuning process, we also introduced a contrastive loss function to optimize the expression of semantic embeddings. Specifically, assume f i represents the semantic embedding generated by the i-th sample, and f i + represents its corresponding positive sample embedding. The loss function can be defined as: where sim(·,·) represents the cosine similarity calculation function, and τ is the temperature hyperparameter. This formula effectively improves the discriminative ability of the embedding vector in semantic expression by pulling the distance between positive samples closer and pushing the distance between negative samples farther apart;
[0020] The GraphCNN model used to obtain the structural feature embeddings adopts a 5-layer structure, where 4 layers actually perform graph convolution (neighbor information aggregation and feature transformation). Inside each graph convolution layer, a 5-layer MLP (Multi-Layer Perceptron) is used for non-linear transformation. The feature dimension of all hidden layers is 128. Both neighbor aggregation and graph pooling adopt the summation method without introducing additional epsilon weights, so that self-loops are directly reflected in the adjacency matrix. Finally, the linear mapping output of the multi-layer features is superimposed after passing through a dropout layer that discards with a probability of 10% to obtain the final prediction.
[0021] Furthermore, in S3, the semantic feature embedding and the structural feature embedding are integrated, and the synthetic embedding is stored in the local offline vector database. Specifically:
[0022] The semantic feature embedding obtained in S2 is concatenated with the structural feature embedding for fusion to form a comprehensive feature vector representing the overall information of the binary function. This fusion strategy effectively integrates the semantic and structural information, making the representation of each function more comprehensive. All the generated comprehensive embeddings, including the vector of the query function, will be stored in the local offline vector database for subsequent fast and efficient similarity retrieval.
[0023] Furthermore, in S4, the method for using the HNSW algorithm to search for the top K similar vectors in the local offline vector database is as follows:
[0024] Then, using the comprehensive embedding of the query function obtained in S3, the HNSW algorithm is used to efficiently retrieve the local offline vector database. The HNSW algorithm is based on the Approximate Nearest Neighbor (ANN) technology and can quickly find the top K candidate embeddings that are most similar to the query vector in large-scale vector data. This step greatly reduces the computational cost while ensuring the accuracy of the retrieval process, providing a high-quality candidate set for subsequent fine-grained comparison.
[0025] In the HNSW algorithm, a new node is inserted into the 0 to lever layer with a probability where lever is randomly generated (exponential decay distribution). The higher-layer nodes are sparser for fast navigation, and the lower-layer nodes are denser for precise search. In each layer, a heuristic algorithm is used to select the nearest M neighbors to construct short connections and long connections: the currently searched candidate nodes candidates are maintained by the max heap heapq, and the nearest ef_construction candidates are retained. Then, after sorting the candidate pool in ascending order of distance by sorted(neighbors, key=lambda x: -x[0]), only the first M neighbors are retained to avoid over-connection. At the same time, when inserting, the bidirectional connection between the new node and its neighbors is truncated if len(...) > M: pop(), ensuring the sparsity and navigation efficiency of the graph structure in each layer.
[0026] The search first selects an entry node in the highest layer (the sparsest layer) and gradually searches downward. Second, in each layer, starting from the current node, move along the node closest to the query vector until no closer node can be found, and then enter the denser graph in the lower layer. Repeat the second step until the bottom layer to return the final result. Use the L2 Euclidean distance to search for the top K vectors with the closest Euclidean distance in the vector database. The Euclidean distance is calculated where A and B represent two vectors, and d represents the number of dimensions of the vector.
[0027] Further, in S5, the specific process of using the Siamese network to find the exact most similar function is as follows:
[0028] The comprehensive embedding of the query function and the topK candidate embeddings filtered out in S4 are jointly input into the Siamese network composed of two three-layer multi-layer perceptrons with exactly the same structure. Through the deep comparison mechanism of the Siamese network, the system can further analyze and quantify the subtle differences between candidate functions, thereby accurately identifying the truly most similar binary function. The refined comparison at this stage not only further improves the accuracy of similarity detection but also ensures the robustness and reliability of the entire system in practical applications;
[0029] The Siamese network receives the embeddings generated by two internal neural networks and produces the Euclidean distance as the output. The two multi-layer perceptron neural networks share the same parameters and are jointly iteratively optimized using the loss function of stochastic gradient descent, which remains the same during the training process. The loss function is as follows: Given n pairs of extracted binary function vectors (E1, E2), when they are similar, each pair is assigned a label y i = +1, otherwise y i = -1. Through this loss function, it is ensured that the embedding E of a specific binary function i is closer to the embeddings of all its similar binary functions, thereby determining the truly most similar binary function.
[0030] The present invention has achieved the following beneficial technical effects compared with the prior art:
[0031] 1. The present invention provides a method for fine-tuning a large language model using prompt engineering to generate binary function semantic embeddings. By leveraging the powerful semantic learning ability of the large language model, it effectively extracts accurate semantic features from assembly code, enabling the generated embeddings to not only capture the core logic of the code but also keenly identify subtle differences brought about by compilation optimization or code obfuscation. This process greatly reduces the dependence on manual feature engineering and provides a solid semantic foundation for subsequent processing, overall improving the accuracy and robustness of detection.
[0032] 2. The multi-feature fusion binary code similarity detection method based on large language models provided by the present invention extracts function semantic embeddings through a large language model fine-tuned by prompt engineering, and combines the control flow graph structure embeddings generated by GraphCNN to achieve multi-level representations across syntax and program logic. It innovatively combines the efficient approximate nearest neighbor search of HNSW with a two-branch Siamese network. First, it quickly filters the candidate set through a hierarchical graph index, and then performs fine similarity measurement through the parameter-sharing Siamese network, breaking through the limitations of single-feature analysis while maintaining computational efficiency. This solution first combines the deep semantic understanding ability of large language models with the topological modeling advantages of graph neural networks, and balances detection accuracy and system overhead through a two-stage retrieval-validation mechanism, providing a new multi-dimensional analysis paradigm for binary code comparison. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 It is a flowchart of the multi-feature fusion binary code similarity detection method based on large language models provided by an embodiment of the present invention.
[0034] Figure 2 It is an architecture diagram of the multi-feature fusion binary code similarity detection method based on large language models provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0035] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0036] The core of the present invention is to provide a multi-feature fusion binary code similarity detection method based on large language models to solve the problem of limitations in existing binary similarity detection methods.
[0037] In order to enable those skilled in the art to better understand the solution of the present invention, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0038] Figure 1 It is a flowchart of a multi-feature fusion binary code similarity detection method based on large language models provided by an embodiment of the present invention, as Figure 1 shown, including the following steps:
[0039] S1: Use IDAPro to perform reverse engineering on the binary file, extract the function features of the binary file and perform preprocessing to obtain the assembly code and CFG control flow graph of each function;
[0040] Use IDAPro to perform reverse engineering on binary files and extract the function features of binary files. During this process, IDAPro splits each function, obtains the corresponding assembly language code for each function, and extracts key information such as the hash value of the function through preprocessing operations. At the same time, it organizes and unifies the specific information of these functions to construct the CFG (Control Flow Graph) of each function;
[0041] Among them, the key information extracted from the function includes the hash value sHash, offset sea, function name name, number of nodes node_size, and number of edges edge_size. Then, use the networkx module in python to create an empty graph g, and use the obtained number of function nodes node_size and number of edges edge_size to construct the control flow graph of the function in the empty graph g. At the same time, store the hash value sHash, offset sea, and function name name in the control flow graph information for subsequent data processing and extraction.
[0042] S2: Put the assembly code corresponding to the obtained binary function into the large language model fine-tuned in advance using prompt engineering to generate its semantic feature embedding, and at the same time put the CFG of the function into the GraphCNN (Graph Convolutional Neural Network) to generate structural feature embedding;
[0043] Input the preprocessed function assembly code into the large language model finely tuned using prompt engineering to generate feature embeddings that can accurately reflect the semantic information of the function. At the same time, input the CFG corresponding to the function into the GraphCNN graph convolutional neural network to generate structural feature embeddings through graph structure learning. This parallel process can not only capture the subtle syntax features contained in the code but also obtain the complex control flow structure information inside the function, realizing the efficient encoding of the binary code in both semantic and structural dimensions;
[0044] To accurately construct the semantic feature vector of the binary function, we use prompt engineering to fine-tune the large language model and design a prompt template for specific tasks of assembly code: "Please extract the semantic feature vector describing the function logic and key information according to the following assembly code: {assembly code}", to clearly guide the model to focus on the semantic details in the code and thus construct the semantic feature vector of the function;
[0045] During the fine-tuning process, we also introduce a contrastive loss function to optimize the expression of semantic embeddings. Specifically, assume f i represents the semantic embedding generated by the i-th sample, and f i + represents its corresponding positive sample embedding. The loss function can be defined as: Among them, sim(·,·) represents the cosine similarity calculation function, and τ is the temperature hyperparameter. This formula effectively improves the discriminative ability of the embedded vector in semantic expression by shortening the distance between positive samples and lengthening the distance between negative samples;
[0046] The GraphCNN model used to obtain the structural feature embedding adopts a 5-layer structure. Among them, 4 layers actually perform graph convolution (neighbor information aggregation and feature transformation). Inside each graph convolution layer, a 5-layer MLP (Multi-Layer Perceptron) is used for non-linear transformation. The feature dimension of all hidden layers is 128. Both neighbor aggregation and graph pooling adopt the summation method without introducing additional epsilon weights, enabling self-loops to be directly reflected in the adjacency matrix. Finally, the linear mapping output of the multi-layer features is superimposed after being discarded by the dropout layer with a probability of 10% to obtain the final prediction.
[0047] S3: Then connect the semantic feature embedding and the structural feature embedding to obtain the composite embedding representing this binary function. Include the embedding of the query function, and store the composite embeddings corresponding to all generated functions in the local offline vector database;
[0048] Connect and fuse the semantic feature embedding obtained in S2 with the structural feature embedding to form a comprehensive feature vector representing the overall information of this binary function. This fusion strategy effectively integrates the information from both semantic and structural aspects, making the representation of each function more comprehensive. All generated composite embeddings, including the vector of the query function, will be stored in the local offline vector database for subsequent fast and efficient similarity retrieval.
[0049] S4: Apply the HNSW algorithm to the embedding of the synthesized query function to find the top K most similar vectors in the local offline vector database;
[0050] Then use the composite embedding of the query function obtained in S3 and adopt the HNSW algorithm to efficiently retrieve the local offline vector database. The HNSW algorithm is based on the Approximate Nearest Neighbor (ANN) technology and can quickly find the top K candidate embeddings most similar to the query vector in large-scale vector data. This step greatly reduces the computational cost while ensuring the accuracy of the retrieval process and provides a high-quality candidate set for subsequent fine-grained comparison;
[0051] In the HNSW algorithm, the new node has a probability of It is inserted into the 0 to lever layer, where lever is randomly generated (exponential decay distribution). Higher-level nodes are sparser for fast navigation, and lower-level nodes are denser for precise search. In each layer, a heuristic algorithm is used to select the nearest M neighbors to construct short connections and long connections: maintain the currently searched candidate nodes candidates through the max heap heapq, retain the ef_construction candidates with the closest distance, and then sort the candidate pool in ascending order of distance by sorted(neighbors, key=lambda x: -x[0]), and only retain the top M neighbors to avoid over-connection. At the same time, when inserting, truncate the bidirectional connection between the new node and its neighbors if len(...) > M to ensure the sparsity and navigation efficiency of the graph structure in each layer;
[0052] The search first selects an entry node in the highest layer (the sparsest layer) and gradually searches downward. In the second step, in each layer, starting from the current node, move along the node closest to the query vector until no closer node can be found, and then enter the denser graph in the lower layer. Repeat the second step until the bottom layer to return the final result. Use the L2 Euclidean distance to search for the top K vectors with the closest Euclidean distance in the vector database, and calculate the Euclidean distance where A and B represent two vectors, and d represents the number of dimensions of the vectors.
[0053] S5: Finally, embed this query function and pass the top K vectors into the Siamese network for comparison to find the truly most similar binary function among them.
[0054] Pass the comprehensive embedding of the query function and the top K candidate embeddings selected in S4 into the Siamese network composed of two three-layer multi-layer perceptrons with exactly the same structure. Through the depth comparison mechanism of the Siamese network, the system can further analyze and quantify the subtle differences between candidate functions, so as to accurately identify the truly most similar binary function. The fine comparison at this stage not only further improves the accuracy of similarity detection, but also ensures the robustness and reliability of the entire system in practical applications;
[0055] Among them, the Siamese network receives the embeddings generated by two internal neural networks and produces the Euclidean distance as the output. The two multi-layer perceptron neural networks share the same parameters and are jointly iteratively optimized using the loss function of stochastic gradient descent, which remains the same during the training process. The loss function is as follows: For the given n pairs of extracted binary function vectors (E1, E2), each pair is assigned a label y when they are similar i = +1, otherwise y i = -1. Through this loss function, ensure the embedding E of a specific binary functioni Closer to all its similar binary function embeddings, thereby determining the truly most similar binary function.
[0056] The above has introduced in detail the multi-feature fusion binary code similarity detection method based on the large language model provided by the present invention. The various embodiments in the specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same and similar parts between the various embodiments, reference can be made to each other.
[0057] The content not described in detail in the specification of the present invention belongs to the prior art well-known to those skilled in the art.
Claims
1. A multi-feature fusion binary code similarity detection method based on large language models, characterized in that, It includes the following steps: S1: Use IDA Pro to perform reverse engineering on the binary file, extract the function features of the binary file and perform preprocessing to obtain the assembly code of each function and the CFG control flow graph; S2: Put the obtained assembly code corresponding to the binary function into the large language model fine-tuned in advance using prompt engineering to generate its semantic feature embedding. At the same time, put the CFG of the function into the GraphCNN graph convolutional neural network to generate the structural feature embedding; S3: Connect the semantic feature embedding and the structural feature embedding to obtain the synthetic embedding representing this binary function. Include the embedding of the query function. Store the synthetic embeddings corresponding to all the generated functions in the local offline vector database; S4: Apply the Hierarchical Navigable Small World Graphs (HNSW) algorithm to the synthetic embedding of the query function to find the top K most similar vectors in the local offline vector database; S5: Finally, pass this query function embedding and the top K vectors to the Siamese network for comparison to find the truly most similar binary function among them.
2. The method for detecting the similarity of multi-feature fusion binary codes based on a large language model according to claim 1, wherein In S1: Use IDA Pro to perform reverse engineering on the binary file and extract the function features of the binary file. In this process, IDA Pro splits out each function to obtain the assembly language code corresponding to each function, and extracts key information such as the hash value of the function through preprocessing operations. At the same time, organize and unify the specific information of these functions to construct the CFG control flow graph of each function; Among them, the key information extracted from the function includes the hash value sHash, the offset sea, the function name name, the number of nodes node_size, and the number of edges edge_size. Then use the networkx module in python to create an empty graph g, and use the obtained number of function nodes node_size and the number of edges edge_size to construct the control flow graph of the function in the empty graph g. At the same time, store the hash value sHash, the offset sea, and the function name name in the control flow graph information for subsequent data processing and extraction.
3. The multi-feature fusion binary code similarity detection method based on a large language model according to claim 1, wherein In S2: Input the preprocessed function assembly code into the large language model fine-tuned through prompt engineering to generate feature embeddings that can accurately reflect the semantic information of the function. At the same time, input the CFG corresponding to the function into the GraphCNN graph convolutional neural network to generate structural feature embeddings through graph structure learning. This parallel process can not only capture the subtle syntactic features contained in the code but also obtain the complex control flow structure information inside the function, realizing the efficient encoding of the binary code in both semantic and structural dimensions; To accurately construct the semantic feature vector of a binary function, we adopted prompt engineering to fine-tune the large language model and designed a prompt template for specific tasks of assembly code: "Please extract the semantic feature vector that describes the function logic and key information from the following assembly code: {assembly code}", to clearly guide the model to focus on the semantic details in the code and thus construct the semantic feature vector of the function; During fine-tuning, we also introduce a contrastive loss function to optimize the expression of semantic embedding. Specifically, assuming f i represents the semantic embedding generated by the i-th sample, f i + represents the corresponding positive sample embedding, and the loss function can be defined as: Among them, sim(·,·) represents the cosine similarity calculation function, and τ is the temperature hyperparameter. This formula effectively improves the discriminative ability of the embedding vector in semantic expression by shortening the distance between positive samples and increasing the distance between negative samples. The GraphCNN model used to obtain the structural feature embedding adopts a 5-layer structure, among which 4 layers actually perform graph convolution (neighbor information aggregation and feature transformation). Inside each graph convolution layer, a 5-layer MLP (Multi-Layer Perceptron) is used for non-linear transformation. The feature dimension of all hidden layers is 128. Both neighbor aggregation and graph pooling adopt the summation method without introducing additional epsilon weights, so that self-loops are directly reflected in the adjacency matrix. Finally, the linear mapping output of the multi-layer features is superimposed after being discarded by the dropout layer with a probability of 10% to obtain the final prediction.
4. The method for detecting the similarity of multi-feature fusion binary codes based on a large language model according to claim 1, characterized in that In S3: The semantic feature embedding obtained in S2 is connected with the structural feature embedding for fusion to form a comprehensive feature vector representing the overall information of the binary function. This fusion strategy effectively integrates the information from both semantic and structural aspects, making the representation of each function more comprehensive. All the generated comprehensive embeddings, including the vector of the query function, will be stored in the local offline vector database for subsequent fast and efficient similarity retrieval.
5. The method for detecting binary code similarity with multi-feature fusion based on a large language model according to claim 1, characterized in that In S4: Using the comprehensive embedding of the query function obtained in S3, the HNSW algorithm is used to efficiently retrieve the local offline vector database. The HNSW algorithm is based on the Approximate Nearest Neighbor (ANN) technology and can quickly find the topK candidate embeddings most similar to the query vector in large-scale vector data. This step greatly reduces the computational cost while ensuring the accuracy of the retrieval process and provides a high-quality candidate set for subsequent fine-grained comparison; In the HNSW algorithm, a new node is inserted into the 0 to lever layer with probability where lever is randomly generated (exponentially decaying distribution). Nodes in higher layers are sparser for fast navigation, and nodes in lower layers are denser for precise search. In each layer, a heuristic algorithm is used to select the nearest M neighbors to build short and long connections: the currently searched candidate nodes candidates are maintained by a max heap heapq, and the ef_construction nearest candidates are retained. Then, after sorting the candidate pool in ascending order of distance by sorted(neighbors, key=lambda x: -x[0]), only the first M neighbors are retained to avoid over-connection. At the same time, when inserting, the bidirectional connection between the new node and its neighbors is truncated if len(...) > M: pop(), ensuring the sparsity and navigation efficiency of the graph structure in each layer; The search first selects an entry node at the highest level (the sparsest level) and gradually searches downward. In the second step, starting from the current node in each layer, it moves along the node closest to the query vector until no closer node can be found. Then it enters the denser graph in the lower layer and repeats the second step until the bottom layer to return the final result. The L2 Euclidean distance is used to search for the top K vectors with the closest Euclidean distance in the vector database. Euclidean distance calculation Where A and B represent two vectors, and d represents the number of dimensions of the vectors.
6. The method for detecting binary code similarity with multi-feature fusion based on a large language model according to claim 1, characterized in that In S5: The comprehensive embedding of the query function and the topK candidate embeddings selected in S4 are passed into a Siamese network composed of two three-layer multi-layer perceptrons with exactly the same structure. Through the deep comparison mechanism of the Siamese network, the system can further analyze and quantify the subtle differences between candidate functions, thereby accurately identifying the truly most similar binary functions. The fine-grained comparison at this stage not only further improves the accuracy of similarity detection but also ensures the robustness and reliability of the entire system in practical applications; The Siamese network receives the embeddings produced by two internal neural networks and produces the Euclidean distance as the output. Two multi-layer perceptron neural networks share the same parameters and are jointly iteratively optimized using the loss function of stochastic gradient descent, which remains the same during the training process. The loss function is as follows: For a given n pairs of extracted binary function vectors (E1, E2), when they are similar, each pair is assigned a label y i = +1, otherwise y i = -1. Through this loss function, it is ensured that the embedding E of a specific binary function i is closer to the embeddings of all its similar binary functions, thereby determining the truly most similar binary function.
Citation Information
Cited By
Software security detection method and system based on code lines
CN121188778A
Code reuse recommendation method and system based on multi-language unified analysis and vectorization retrieval
CN121255180A