Binary program vulnerability detection method and system based on large decompilation model and EAST features

By decompiling the big model and EAST features based on the method, the problems of syntax semantic information loss and gradient descent in binary similarity analysis are solved, and efficient and accurate detection of vulnerability in binary program is achieved.

CN120162791APending Publication Date: 2025-06-17Chinese People's Liberation Army Cyberspace Force Information Engineering University
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510207613.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The existing binary similarity analysis has problems with missing syntax semantic information and gradient descent in the tree neural network model in vulnerability detection, resulting in insufficiency of detection and high false positive rate.

Method used

A binary program vulnerability detection method based on decompiled large models and enhanced abstract syntax tree (EAST) features is adopted. By fine-tuning the decompiled large models and extracting enhanced abstract syntax tree features, and similarity calculation is performed in combination with the Siamese network to evaluate the vulnerabilities of binary programs.

Benefits of technology

It improves the accuracy and efficiency of non-obfuscated binary program vulnerability detection, solves the gradient disappearance problem in the tree neural network architecture, and significantly improves the reliability of the detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120162791A_ABST
    Figure CN120162791A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer program detection, in particular to a binary program vulnerability detection method and system based on a large decompilation model and EAST features. A vulnerability function data set and a system call function data set are used for performing fine adjustment on a pre-trained large decompilation model LLM4Decompact; performing grammar and semantic optimization on the pseudo code which is output by the anti-compiler Ghidra and is subjected to the decompilation by using the LLM4Decompile; extracting a corresponding enhanced abstract syntax tree EAST from the decompilation pseudo-code after grammar and semantic optimization of the decompilation large model after fine tuning, and performing feature coding on the enhanced abstract syntax tree by using an EASTNN model; and inputting the feature vectors of the coded non-obfuscated binary program to be detected and the vulnerability function data set into a Siamese network for similarity calculation, and evaluating the vulnerability of the non-obfuscated binary program according to a similarity calculation result. According to the method, by using the fine-tuned decompilation large model and the enhanced abstract syntax tree feature coding, the accuracy and efficiency of non-confusion binary program vulnerability detection are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer program detection, and particularly relates to a binary program vulnerability detection method and system based on a decompiled large model and EAST features. Background Art

[0002] With the continuous increase in the complexity and scale of software development, and the gradual popularity of open source projects and software life cycle development rules, the development between different software has gradually tended to be modularized. Although this modular development technology has greatly reduced the development cost of new software and saved the development cycle, it has also led to software developed using the same vulnerable module components having the same vulnerabilities, and the number and types of software vulnerabilities have also been increasing rapidly. According to the 2024 Open Source Security and Risk Analysis Report released by Synopsys, the report points out that the use of open source components is universal. Currently, 96% of code libraries contain open source components, 77% of source code and files also come from open source. At the same time, there are also numerous vulnerabilities and risks in open source components. Among them, 84% of the open source code libraries used to evaluate security risks have at least one known vulnerability, and 74% of these code libraries contain high-risk vulnerabilities, which is a significant increase compared to 48% in the statistics of 2022. Finally, the report points out that 87% of the open source code libraries in the computer hardware and semiconductor industries contain high-risk vulnerabilities, indicating that the computer software and hardware industries have extensive vulnerability risks. Binary code vulnerability detection, as an important means to ensure the security of software systems, has become a key research direction in the field of information security.

[0003] Traditional vulnerability detection methods, such as symbolic execution, taint analysis, and dynamic detection, although achieving good results in certain specific scenarios, often face problems such as low efficiency, high false positive rate, and poor cross-platform adaptability when dealing with large-scale binary code analysis. With the increasingly complex means of attackers, new vulnerability detection methods are urgently needed. In recent years, vulnerability detection methods based on binary similarity analysis have attracted extensive attention. The core idea of binary similarity analysis is to judge whether there are the same or similar functions, structures, or security defects by comparing the similarity between different binary code segments. Compared with traditional source code analysis-based vulnerability detection methods, binary similarity analysis methods have advantages such as not relying on source code, wide application range, and strong ability to detect hidden vulnerabilities. Since binary code is essentially compiled machine code, information such as source code comments, structures, and programming language features will be lost during the compilation process. Therefore, vulnerability detection for binary code is not only more complex but also needs to solve problems such as code optimization, compiler differences, and platform differences. Summary of the Invention

[0004] The present invention aims to solve the problems of the lack of syntax and semantic information in current binary similarity analysis for vulnerability detection and the gradient descent problem in the tree neural network model, and proposes a binary program vulnerability detection method and system based on a decompilation large model and EAST features. By using the fine-tuned decompilation large model and enhanced abstract syntax tree feature encoding, the accuracy and efficiency of non-obfuscated binary program vulnerability detection are improved, and at the same time, the problem of gradient disappearance in the tree neural network architecture is solved.

[0005] To achieve the above object, the technical solution adopted is as follows:

[0006] The present invention provides a binary program vulnerability detection method based on a decompilation large model and EAST features, including:

[0007] Using the vulnerability function dataset and the system call function dataset to fine-tune the pre-trained decompilation large model LLM4Decompile, so that LLM4Decompile optimizes the syntax and semantics of the decompiled pseudocode output by the decompiler Ghidra, and obtains the fine-tuned decompilation large model through iterative training;

[0008] Extracting the corresponding enhanced abstract syntax tree EAST from the decompiled pseudocode with optimized syntax and semantics by the fine-tuned decompilation large model, and performing feature encoding on the enhanced abstract syntax tree using the EASTNN model;

[0009] Inputting the feature vectors of the encoded non-obfuscated binary program to be detected and the vulnerability function dataset into the Siamese network for similarity calculation, and evaluating the non-obfuscated binary program vulnerability according to the similarity calculation result.

[0010] According to the binary program vulnerability detection method based on the decompilation large model and EAST features of the present invention, further, the process of fine-tuning the pre-trained decompilation large model LLM4Decompile is as follows:

[0011] Load the pre-trained decompilation large model M LLM And initialize the fully connected layer for large model fine-tuning preparation;

[0012] Start the large model iterative training, sampling a batch of samples B v from the vulnerability function dataset D s and the system call function dataset D v and B s , and each sample consists of a vulnerability function and a system call function;

[0013] Use the decompiler Ghidra to decompile the original binary input to obtain the pseudocode C v ,C s ;

[0014] Using the model forward propagation method, the decompiled pseudo-code C v and C s are input into M LLM , and the decompilation result after vulnerability and system function syntax and semantics optimization is obtained;

[0015] The loss function is calculated, and then the model is updated. The model parameters are updated using the gradient descent method. After multiple iterations are completed, the fine-tuned decompilation large model M fine-tuned is output.

[0016] According to the binary program vulnerability detection method based on the decompilation large model and EAST features of the present invention, further, during the fine-tuning process of the large model, the loss function calculation includes:

[0017] For the syntax and semantics optimization of the vulnerability function, calculate the cross-entropy loss between the prediction result P LLM of M v and the vulnerability source code annotation dataset Y v corresponding to P; for the syntax and semantics optimization of the system call function, calculate the cross-entropy loss between the prediction result P v of M LLM and the system call source code annotation dataset Y s corresponding to P s , and combine the two loss functions through weight assignment as the total loss. s

[0018] According to the binary program vulnerability detection method based on the decompilation large model and EAST features of the present invention, further, extracting the corresponding enhanced abstract syntax tree EAST from the decompiled pseudo-code after syntax and semantics optimization of the fine-tuned decompilation large model specifically includes: first, performing preprocessing operations on the decompiled pseudo-code of the large model; then, performing code normalization processing on the decompiled pseudo-code of the large model; finally, constructing an enhanced abstract syntax tree after normalization processing and syntax and semantics completion.

[0019] According to the binary program vulnerability detection method based on the decompilation large model and EAST features of the present invention, further, the preprocessing operation on the decompiled pseudo-code of the large model includes removing invalid characters, basic syntax completion, and placeholder marking; the code normalization processing on the decompiled pseudo-code of the large model includes variable type completion and inference, and function return type completion and inference.

[0020] According to the binary program vulnerability detection method based on the decompilation large model and EAST features of the present invention, further, the feature encoding of the enhanced abstract syntax tree using the EASTNN model includes:

[0021] ​By taking the first bifurcation point of the Enhanced Abstract Syntax Tree (EAST) as the segmentation point, the EAST divided into different parts, namely p1…p n , and then input into the function-level encoder, the output vectors e1…e t Finally, an EAST encoding vector V is generated through a bidirectional gated recurrent unit and pooling.

[0022] According to the binary program vulnerability detection method based on decompilation large model and EAST features of the present invention, further, the process of inputting the feature vectors of the encoded non-obfuscated binary program to be detected and the vulnerability function data set into the Siamese network for similarity calculation is as follows: Siamese first uses EASTNN to encode the EAST syntax tree features T1 and T2 into fixed-length feature vectors V1 and V2, and then inputs the feature vectors V1 and V2 into the cosine function COS(V1, V2) for similarity calculation to obtain a specific score s.

[0023] According to the binary program vulnerability detection method based on decompilation large model and EAST features of the present invention, further, a similarity determination threshold is set as θ. When the similarity calculation score s≥θ, it is determined to be similar and the non-obfuscated binary program to be detected has a vulnerability; when the similarity calculation score s<θ, it is determined to be dissimilar and the non-obfuscated binary program to be detected does not have a vulnerability.

[0024] Further, the present invention also provides a binary program vulnerability detection system based on decompilation large model and EAST features, which is used to implement the binary program vulnerability detection method based on decompilation large model and EAST features as described above. The system includes a large model fine-tuning module, an EAST feature encoding module, and a similarity calculation module, where:

[0025] The large model fine-tuning module is used to fine-tune the pre-trained decompilation large model LLM4Decompile using the vulnerability function data set and the system call function data set, so that LLM4Decompile performs syntax and semantic optimization on the decompiled pseudo-code output by the decompiler Ghidra, and obtains the fine-tuned decompilation large model through iterative training;

[0026] The EAST feature encoding module is used to extract the corresponding Enhanced Abstract Syntax Tree (EAST) from the decompiled pseudo-code with optimized syntax and semantics by the fine-tuned decompilation large model, and perform feature encoding on the Enhanced Abstract Syntax Tree using the EASTNN model;

[0027] The similarity calculation module is used to input the feature vectors of the encoded non-obfuscated binary program to be detected and the vulnerability function data set into the Siamese network for similarity calculation, and evaluate the vulnerability of the non-obfuscated binary program according to the similarity calculation result.

[0028] The beneficial effects achieved by adopting the above technical solutions are as follows:

[0029] 1. The present invention uses the decompiler Ghidra to perform decompilation processing on binary programs, abandons the traditional intermediate disassembly representation with missing semantic and structural information, and at the same time uses a vulnerability function dataset and a system call function dataset to fine-tune the decompilation large model; by inputting the dataset into the fine-tuned decompilation large model, the syntax and semantic information of the Ghidra decompilation result is optimized, making the decompilation result close to the syntax and semantic structure of the original source code.

[0030] 2. The present invention proposes an enhanced abstract syntax tree (EAST) extraction method. By preprocessing and code normalization operations on the decompiled pseudocode, the irregular pseudocode processed by ordinary decompilation tools is parsed, compared with the source code program for normalization, and the abstract syntax tree (AST) features are extracted. EAST further optimizes the syntax and semantic information compared with the original AST syntax tree features, and the structure is more complete.

[0031] 3. The present invention designs an enhanced abstract syntax tree neural network model (EASTNN) for efficiently processing the EAST structure. This neural network is an improvement based on the abstract syntax tree neural network model (ASTNN), which can efficiently process the EAST multimodal feature tree, encode it, and represent it in a feature vectorized manner, improving the efficiency of feature encoding and vectorization, and basically solving the problem of gradient disappearance often existing in the tree-shaped neural network architecture. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings of the embodiments of the present invention will be briefly introduced below. Among them, the drawings are only used to show some embodiments of the present invention, rather than limiting all embodiments of the present invention thereto.

[0033] Figure 1 is a framework diagram of the binary program vulnerability detection method based on the decompilation large model and EAST features in the embodiment of the present invention;

[0034] Figure 2 is a fine-tuning process diagram of the decompilation large model LLM4Decompile in the embodiment of the present invention;

[0035] Figure 3 is an EAST extraction process diagram of the decompiled pseudocode after optimizing the large model in the embodiment of the present invention;

[0036] Figure 4 is an architecture diagram of the ASTNN and EASTNN models in the embodiment of the present invention;

[0037] Figure 5It is the architecture diagram of the Siamese network in the embodiments of the present invention;

[0038] Figure 6 It is the bar-line hybrid analysis diagram of the firmware vulnerability detection task results in the embodiments of the present invention;

[0039] Figure 7 It is the heat map of the correlation analysis between CVE vulnerabilities and OpenSSL versions in the embodiments of the present invention;

[0040] Figure 8 It is the similarity analysis and vulnerability detection task time performance evaluation diagram in the embodiments of the present invention. Detailed implementation manners

[0041] In the following, the exemplary solutions of the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings of the specific embodiments of the present invention. Unless otherwise defined, the technical terms or scientific terms used in the present invention shall have the ordinary meanings understood by those of ordinary skill in the art.

[0042] This embodiment discloses a binary program vulnerability detection method - Zeus based on a decompilation large model and EAST features, as Figure 1 shown, which includes the following steps:

[0043] Step S101: Use the vulnerability function dataset and the system call function dataset to fine-tune the pre-trained decompilation large model LLM4Decompile, so that LLM4Decompile optimizes the syntax and semantics of the decompiled pseudocode output by the decompiler Ghidra, and obtains the fine-tuned decompilation large model through iterative training.

[0044] The existing LLM4Decompile effectively improves the accuracy and practicability of binary code decompilation, solves many problems that are difficult to handle by traditional disassembly / decompilation tools, and is particularly excellent in the readability of the decompiled code and the optimization of code semantics. In order to further improve the accuracy of the decompilation optimization of semantics and syntax of LLM4Decompile in vulnerability detection, this solution uses the vulnerability function dataset and the system call function dataset commonly used in software programs for targeted fine-tuning, so that the large model can supplement and predict the function names of relevant vulnerability functions and common system call names when optimizing semantic expressions, and improve the structural integrity and semantic comprehensiveness of the decompiled code.

[0045] The core of fine-tuning LLM4Decompile is to perform secondary expansion on the model based on the pseudo-code generated by Ghidra decompilation, combined with vulnerability function data and system call function data. To enable LLM4Decompile to perform better in vulnerability detection and system call prediction tasks, the publicly available vulnerability function dataset Vul-Dataset and system call function dataset Syscall-Dataset were collected, and the Ftun-Dataset large model fine-tuning dataset was formed in an 8:2 ratio. By training the LLM4Decompile large model specifically on the Ftun-Data dataset, it can perform targeted decompiled code semantic optimization and completion for vulnerability detection tasks and similarity analysis tasks. The simple fine-tuning process of the LLM4Decompile large model is as Figure 2 shown.

[0046] Algorithm 1 presents the fine-tuning process of the decompiled large model. First, the parameters required by the model are input, mainly including the vulnerability function dataset D v , the system call function dataset D s , the original decompiled large model M LLM , the decompiler Ghidra(X), the number of training iterations E, and the learning rate l r . Then, the pre-trained decompiled large model M LLM is loaded and the fully connected layer is initialized for fine-tuning preparation. Then, the algorithm iterative training begins. For each epoch, first, sample batching is performed. A batch of samples B v and B s are sampled from D v and D s . Each sample consists of the corresponding vulnerability function and system call function, and the original binary input is decompiled using the decompiler Ghidra to obtain the pseudo-code C v ,C s . Then, using the model forward propagation method, the decompiled pseudo-code C v and C s are input into M LLM to obtain the decompiled result with optimized syntax and semantics of the vulnerability and system functions. Finally, the loss function is calculated. For the optimization of the syntax and semantics of the vulnerability function, the cross-entropy loss between the prediction result P LLM of M v and the corresponding vulnerability source code annotation dataset Y v is calculated; for the optimization of the syntax and semantics of the system call function, the cross-entropy loss between the prediction result P v of M LLM and the corresponding system call source code annotation dataset Y s is calculated. Finally, for the system call function syntax and semantics optimization, the cross-entropy loss between the prediction result P s of M sThe cross-entropy loss combines two loss functions through weight allocation as the total loss (Line 12), where α and β are the weights of the total loss. The weight allocation for α and β is determined according to the ratio of the vulnerability function dataset and the system call function dataset in the large model fine-tuning dataset. That is, α = 0.8 and β = 0.2 are set. Such weight allocation reflects the importance of semantic optimization of vulnerability functions in the vulnerability detection task, and semantic optimization is performed for its high-weight loss. Then the model is updated, and the model parameters are optimized through backpropagation. The gradient descent method is used to update the model parameters, and the learning rate l r is used to control the step size of the model parameter update during the gradient descent process. The Adam adaptive learning rate optimizer is used to set the value of lr, and is the gradient of the loss function L with respect to the model parameters, which can be calculated during the implementation of the backpropagation algorithm. Finally, after E iterations, the fine-tuned decompiled large model M fine-tuned is output. The fine-tuned decompiled large model M fine-tuned can perform syntax and semantic optimization and default complementation on vulnerability functions and system call functions in the decompiled pseudo-code.

[0047] By fine-tuning the decompiled large model, the decompiled pseudo-code after Ghidra decompilation is further optimized, and the optimized decompiled function set is input into the subsequent tasks, improving the efficiency of subsequent EAST extraction and encoding. At the same time, it also provides higher-quality training datasets and test datasets for the similarity calculation and learning of the Siamese model.

[0048]

[0049] Step S102: Extract the corresponding enhanced abstract syntax tree EAST from the decompiled pseudo-code after syntax and semantic optimization of the fine-tuned decompiled large model, and perform feature encoding on the enhanced abstract syntax tree using the EASTNN model.

[0050] (1) Enhanced abstract syntax tree EAST feature extraction

[0051] This solution proposes a method for extracting features of an enhanced abstract syntax tree (EAST). Through a series of preprocessing and code normalization operation steps, for the problems of partial semantic parsing and missing syntax structures still existing in the pseudo-C code after decompilation by the large model, through certain preprocessing and normalization operations, an enhanced abstract syntax tree that conforms to the standardized syntax and semantic expression, namely EAST, is generated. At the same time, it provides a high-quality feature set for the subsequent EAST feature encoding work, further improving the detection performance of Zeus.

[0052] ① Preprocess the decompiled code

[0053] The invalid characters and incomplete syntax structures still present in the decompiled code of the large model need to be preprocessed first, which mainly includes the following three steps. Removing invalid characters: Cleaning up redundant whitespace, non-printable characters, and comments to ensure code format standardization. Completing basic syntax: Completing basic symbols such as semicolons and parentheses at the end of each line of code to avoid parsing errors. Placeholder marking: Marking placeholders for missing key syntax elements (such as variable types or function return types), such as __UNKNOWN_TYPE, for subsequent repair.

[0054] ② Code normalization processing

[0055] Regarding the issues of non-standard structures in the decompiled code of the large model, such as missing variable types and incomplete function declarations, the code is normalized using regular expressions to ensure compliance with the standard C language syntax rules, mainly including the following two steps. Variable type completion and inference: If a variable is missing or the non-standard part exists as a variable declaration with the type placeholder __UNKNOWN_TYPE after being optimized by the large model decompilation, its type is inferred as int. This inference is based on default settings, especially when there is insufficient context information to infer the actual type. Function return type completion and inference: If the missing part or non-standard part is a function definition and the return type is void, further check whether the function body contains a return value statement. If a return statement exists, the return type is inferred as int; otherwise, void is retained.

[0056] ③ Output EAST

[0057] Finally, an enhanced abstract syntax tree after normalization processing and syntax and semantics completion is output. As Figure 3 shown, it is the EAST extraction process of the decompiled pseudocode after being optimized by the large model. The structure of the decompiled code after being processed through the above steps is more complete and the syntax is accurate, which can more accurately represent the semantic structure of the original program source code and can provide a more accurate EAST feature dataset for subsequent binary similarity analysis calculation and learning.

[0058] (2) EAST feature encoding based on the EASTNN model

[0059] Currently, the main method for parsing and encoding ordinary Abstract Syntax Trees (ASTs) is to use Tree Neural Networks (TNNs). Given an AST tree, TNNs learn its vector representation by recursively computing node embeddings from bottom to top. The most representative methods of TNNs mainly include Recurrent Neural Networks (RNNs), Tree-based Convolutional Neural Networks (Tree-based CNNs), and Tree Long Short-Term Memory Networks (Tree-LSTMs). These methods can better capture semantic and syntactic information. However, these tree-based neural models are vulnerable to the vanishing gradient problem, and traversing and encoding the entire syntax tree in a bottom-up manner or using the sliding window technique may lose long-term context information. Secondly, to simplify and improve efficiency, these methods either transform the syntax tree into or directly view the syntax tree as a complete binary tree, which destroys the original syntax structure of the source code and even makes the syntax tree grow deeper. The transformed deep syntax tree further weakens the ability of the neural model to capture more real and complex semantics. At the same time, the extracted AST is relatively large, making it prone to long-term dependence problems.

[0060] Based on the above problems with TNN for AST syntax tree parsing and encoding, an efficient neural network for AST encoding processing - ASTNN - is used and improved to learn and encode EAST features. The original working mode of this model is to take the AST as input and split each AST into a series of statement trees (STs) through a preorder traversal algorithm, that is, to continue splitting at the AST abstract syntax tree level into a tree structure at the statement level for each function. All ST trees are encoded into vectors by a statement-level encoder, denoted as e1…e t . Then, a Bidirectional Gated Recurrent Unit (Bi-GRU) is used to simulate the naturalness of the statements. Finally, the hidden state of the Bi-GRU is sampled into a single vector V through pooling as the encoding vector of the function-level AST syntax tree, as shown in Figure 4 (a).

[0061] Since the dataset has been supplemented, filtered, and optimized in terms of semantics, syntax, and structural information in the early semantic optimization stage of the decompilation large model and the EAST extraction stage, there is no need to further split and learn at the statement level like the original ASTN model. At the same time, considering that the extracted EAST features are at the function level and the input model granularity is small enough, performing statement-level splitting on EAST will result in additional performance and time overhead. Therefore, based on ASTNN, the statement-level splitting step will be discarded, and model encoding and learning will be directly performed on EAST, as shown in Figure 4 (b). By taking the first bifurcation point of the Enhanced Abstract Syntax Tree (EAST) as the splitting point, the EAST split into different parts is denoted as p1…p n, and then input the function-level encoder, and the output vectors e1…e t Finally, the EAST encoding vector V is generated through a bidirectional gated recurrent unit and pooling. The improved ASTNN model is named EASTNN, that is, the enhanced abstract syntax tree neural network. EASTNN not only realizes the accurate parsing and encoding of the EAST syntax tree at the function level, enabling the model to learn the syntax, semantic knowledge, and structural information in EAST, but also reduces the performance and time overhead of the original ASTNN parsing and encoding, and avoids problems such as gradient descent that often exist in the tree structure encoding of the TNN model.

[0062] Step S103: Input the feature vectors of the encoded non-confused binary program to be detected and the vulnerability function data set into the Siamese network for similarity calculation, and evaluate the non-confused binary program vulnerability according to the similarity calculation result.

[0063] Siamese is a neural network with two branches sharing weights, used to learn and measure the similarity of input data pairs. The feature extraction of the input is respectively performed through two sub-networks, and the similarity analysis is carried out by using a mathematical similarity metric function to generate a similarity metric result. Its design architecture ensures the consistency and efficiency in the similarity calculation process. In this solution, the time complexity and computational cost of the similarity calculation work can be reduced by using the Siamese neural network architecture to calculate the similarity of the vectors encoded by EASTNN.

[0064] The Siamese network architecture designed in this solution integrates two identical EASTNN models to calculate the similarity between the encoded vectors, and uses the cosine function for similarity calculation. The architecture of the Siamese structure M(T1,T2) is as Figure 5 shown. The Siamese architecture consists of two identical EASTNNs, which share the same parameters. In the similarity calculation process, Siamese first encodes the EAST syntax tree features T1 and T2 into fixed-length feature vectors V1 = EASTNN(T1), V2 = EASTNN(T2) by using EASTNN, and then inputs the feature vectors V1 and V2 into the cosine function COS(V1,V2) for similarity calculation to obtain a specific score s. The two EASTNN models use the same input to remain the same during training, and they are jointly and iteratively optimized using a loss function with stochastic gradient descent.

[0065] For the construction of the loss function as shown in formula (1), by giving n pairs of extracted binary function EAST features (T1,T2), calculate the specific score s of the similarity degree of their encoded vectors (V1,V2), and set the similarity determination threshold as θ. When the similarity calculation score s≥θ, it is determined as similar and the label y is assignedi = 1 indicates that there is a vulnerability in the binary program to be detected that is not obfuscated; when the similarity calculation score s < θ, it is determined to be dissimilar and it is assigned the label y i = 0 indicates that there is no vulnerability in the binary program to be detected that is not obfuscated, where the θ threshold parameter will be determined in the experiment. The loss function optimizes the encoded vector to be detected as close as possible to all vectors similar to it and outputs a more accurate similarity calculation result.

[0066]

[0067] Correspondingly to the above method, this embodiment also discloses a binary program vulnerability detection system based on a decompilation large model and EAST features, including a large model fine-tuning module, an EAST feature encoding module, and a similarity calculation module, where:

[0068] The large model fine-tuning module is used to fine-tune the pre-trained decompilation large model LLM4Decompile using the vulnerability function dataset and the system call function dataset, so that LLM4Decompile performs syntax and semantic optimization on the decompiled pseudo-code output by the decompiler Ghidra, and the fine-tuned decompilation large model is obtained through iterative training.

[0069] The EAST feature encoding module is used to extract the corresponding enhanced abstract syntax tree EAST from the decompiled pseudo-code after syntax and semantic optimization by the fine-tuned decompilation large model, and perform feature encoding on the enhanced abstract syntax tree using the EASTNN model.

[0070] The similarity calculation module is used to input the feature vectors of the encoded binary program to be detected that is not obfuscated and the vulnerability function dataset into the Siamese network for similarity calculation, and evaluate the vulnerability of the binary program that is not obfuscated according to the similarity calculation result.

[0071] To verify the effectiveness of this solution, further explanation will be given below in combination with experimental data.

[0072] This experiment mainly conducts a comprehensive performance and usability evaluation of Zeus. For this purpose, the present invention adopts a variety of different measurement metrics, more comprehensively characterizing the detection capabilities of different methods, and adds six classical methods in this research field as experimental baselines for comparison. In addition, three large-scale evaluation datasets are constructed, serving as the training, testing, and validation datasets for similarity analysis experiments and vulnerability detection experiments, as well as the fine-tuning dataset for decompiling large models, respectively, to measure the effectiveness and performance of the method of the present invention from multiple perspectives. Since the experiment conducts similarity analysis and non-obfuscated vulnerability detection at the disassembly level, with a relatively high abstraction level for binary program codes, for the differential compilation of the same semantic program under different CPU instruction set architectures, different compiler versions and types, and different compilation optimization levels, the syntactic structure and semantic differences at the source code abstraction level after decompilation optimization are basically small, which can also be verified by the comparative experimental results. Therefore, in the research, non-obfuscated binary programs are defined as: binary programs compiled under different CPU instruction set architectures, different compiler versions and types, and compilation optimization levels. Obfuscated binary programs are defined as: binary programs compiled under different obfuscation techniques and obfuscation programs.

[0073] (I) Experimental Environment and Dataset Settings

[0074] (1) Experimental Environment Settings

[0075] To meet the requirements of compiling large-scale binary programs and training and testing deep learning models, the experiment runs on a high-performance server cluster. The main computing node is equipped with 2 × Intel Xeon Platinum 8368Q processors (38 cores, 2.6 GHz), 1 TB of DDR4 memory, 8 TB of NVMe SSD storage, and 8 × NVIDIA A100 Tensor Core GPUs (80 GB HBM2e). The experiment uses the Ubuntu 22.04 LTS operating system, integrating tools such as Ghidra 10.2 and LLVM / Clang 15.0 for binary code generation, disassembly, and analysis. The deep learning framework uses PyTorch 2.0.1, paired with CUDA 12.2 and cuDNN 8.8 to optimize training and inference efficiency; data storage uses the distributed Hadoop file system (HDFS) to support the management of large-scale binary programs and functions. The PyTorch distributed data parallel technology is used to accelerate model training, which can process approximately 3,000 functions per second. This environment has excellent computing performance and flexibility, and can efficiently support tasks related to binary program similarity analysis and vulnerability detection.

[0076] (2) Setting of Experimental Related Metrics

[0077] In the evaluation of the work, the similarity evaluation of the function pairs extracted from the dataset is denoted as s, and the threshold for judging whether it is a homologous function is θ. If the similarity score s of the input function pair is greater than or equal to θ, then the function pair is regarded as a positive result, that is, a homologous function; otherwise, it is regarded as a negative result, that is, a non-homologous function. And for the positive and negative results judged by the model, they are further subdivided according to the dataset division, that is, the following four task results are defined.

[0078] ·TP: For homologous pairs, when the similarity score s is greater than or equal to θ, it is judged as True Positive (TP).

[0079] ·FN: For homologous pairs, when the similarity score s is less than θ, it is judged as False Negative (FN).

[0080] ·FP: For non-homologous pairs, when the similarity score s is greater than or equal to θ, it is judged as False Positive (FP).

[0081] ·TN: For non-homologous pairs, when the similarity score s is less than θ, it is judged as True Negative (TN).

[0082] At the same time, according to the evaluation effects of the evaluation indicators in previous studies, 5 indicators were finally selected for comprehensive evaluation. Among them, the AUC indicator was selected to evaluate the experimental results in the similarity analysis task, and the MRR and Recall@topK were selected to evaluate the experimental results in the vulnerability detection task. The specific evaluation indicator settings are shown in formulas (2)-(6).

[0083] ·TPR: That is, True Positive Rate, which measures the ability of the model to identify positive examples, represents the proportion of positive examples correctly predicted by the model among all actual positive examples, and represents the accuracy of homologous function detection at the threshold θ. Its calculation method is shown in formula (2).

[0084] TPR = TP / (TP + FN) #(2)

[0085] ·FPR: That is, False Positive Rate, which measures the proportion of negative examples mispredicted as positive examples by the model, represents the error rate of the model on negative examples, and represents the accuracy of non-homologous function detection at the threshold θ. Its calculation method is shown in formula (3).

[0086] FPR = FP / (FP + TN) #(3)

[0087] · AUC: The AUC is the area under the Receiver Operating Characteristic (ROC) curve, which represents the ability of a classifier to distinguish between positive and negative samples. The larger the AUC value, the better the performance of the classifier. Its value ranges from [0, 1], where 1 represents a perfect classifier and 0.5 represents random guessing. The AUC is usually obtained by integrating the ROC curve, and the calculation method is shown in formula (4).

[0088]

[0089] · MRR: It is used to measure the performance of a vulnerability detection system, which represents the average of the reciprocals of the ranks of the first correct results returned by the system. The higher the rank, the higher the score, and the calculation method is shown in formula (5). Where rank i is the rank of the first relevant result for query i, and n is the total number of queries.

[0090]

[0091] · Recall@top-K: The Recall metric measures what proportion of positive examples a model can find from actual positive examples, which is a comprehensive performance of the model on positive examples. And Recall@top-K is the recall rate for the first k results in the vulnerability detection task, that is, it represents the proportion of correct targets included in the first k prediction results, as shown in formula (6).

[0092]

[0093] (3) Setting the experimental baseline comparison method

[0094] FASER, TREX, Gemini, VulSeeker, SAFE, and Diaphora.

[0095] (4) Construction of the experimental dataset

[0096] For the experimental dataset, it is set as a non-confusing dataset for similarity analysis tasks, a non-confusing dataset for vulnerability detection tasks, and a fine-tuning dataset for decompiling large model fine-tuning. Since the decompiled semantic optimization large model performs high-level abstract syntax and semantic completion tasks for non-confusing vulnerability functions, the construction of this task dataset only divides the dataset compilation under different architectures, without considering different compilers, compilation optimization methods, function inlining, etc. Among them, the similarity analysis dataset conducts dataset compilation and comparative experiment detection on the similarity analysis detection performance under cross-architecture methods on four common CPU instruction set architectures: ARM, X86, X64, and MIPS. The vulnerability detection data uses the collected common vulnerability software package programs and compiles them on the above four architectures. At the same time, real firmware software statistical datasets are also added. The decompiled large model fine-tuning dataset is based on the widely publicized vulnerability dataset and system call dataset. To comprehensively compare the experimental results, a broad benchmark is established based on several works with good experimental performance in previous studies [6, 8, 9, 10, 12, 18]. The experimental design includes a total of 3 task datasets, 3 experimental tasks, and 5 experimental result metrics, and experimental comparisons are made with 6 benchmark methods.

[0097] A. Similarity Analysis Task Dataset Sim-Data

[0098] Through the statistics and collection of the datasets used for similarity analysis tasks in the past few years, it is formed based on the work datasets of Marcelli et al. and BINKIT, supplemented with program software package datasets and system call program datasets commonly used in past research and daily environments. It contains a total of 36,284 functions extracted from 12,879 different binary software programs and firmware programs, and 614,642 pairs of homologous functions and 614,642 pairs of non-homologous functions are generated through compilation on four common architectures: X86, X64, MIPS, and ARM. To ensure that the evaluation effect of the model performance is reasonable enough, the dataset is divided into a training set, a validation set, and a test set using a ratio of 7:1:2. This means that 70% of the function pairs are used to train the model, 10% are used to verify the similarity analysis effect of the model, and the remaining 20% are used to test and evaluate the performance of the model. By dividing the dataset into a training set, a validation set, and a test set, the performance of the model can be evaluated on an unknown dataset to measure its performance in identifying homologous functions.

[0099] B. Vulnerability Detection Task Dataset VDVul-Data

[0100] The dataset for the vulnerability detection task is mainly based on the datasets in Marcelli et al., BINKIT, and YANG et al., and incorporates the vulnerability software package datasets corresponding to common vulnerabilities in the past decade as statistically collected from vulnerability collection websites such as NVD, CNVD, and CVE. Vulnerabilities from different collection sources are uniformly queried and represented using CVE numbers. As shown in Table 2, the software information with vulnerabilities in the first column is collected from open-source software widely used in IoT firmware, including OpenSSL, Tcpdump, MySQL, Docker, Apache Htppd, Busybox, Dnsmap, Lighhttpd, Nginx, and Samba. The second column lists the number of software vulnerabilities. In the third column, the time range of vulnerability disclosure is listed, and the CVE number range of the collected software vulnerabilities is set to those CVE vulnerabilities that have been used by multiple IoT firmwares or have significant harm during the 10-year period from 2014 to 2024. The fourth column describes the main vulnerability types of the collected CVE vulnerabilities. The fifth column lists the software version ranges affected by the collected CVE vulnerabilities, which are integrated according to the union, and the subsequent column lists the software statistics of the software versions corresponding to the collected CVE vulnerabilities.

[0101] Table 2 Sources of the Vulnerability Detection Dataset

[0102]

[0103]

[0104] The vulnerability dataset is established based on the impact of CVE vulnerabilities on software types in Table 2 in the past decade. The statistical quantity of each software with a vulnerability impact range is compiled under four architectures: X64, X86, ARM, and MIPS. Finally, the vulnerability functions corresponding to the CVE numbers under each architecture are extracted. Among them, the CVE vulnerability functions under the X64 architecture are approximately 6.32M, those under the X86 architecture are approximately 6.31M, those under the ARM architecture are approximately 6.30M, and those under the MIPS architecture are approximately 6.30M. The final vulnerability dataset is named VDVul-Dataset, and a total of approximately 25.23M functions corresponding to CVE vulnerability numbers under different architectures are extracted. As shown in Table 3, the reason for the difference in the number of extracted functions after compiling the same software package under different architectures is mainly due to partial software package compilation failures on their architectures caused by compilation environments, compilation programs, and other reasons such as compilation optimization techniques, which fall within the normal error range.

[0105] Table 3 Vulnerability Detection Dataset

[0106]

[0107] C. LLM4Decompile Large Model Fine-tuning Task Dataset FTun-Data

[0108] For the LLM4Decompile large model fine-tuning task, two datasets, namely public vulnerability functions and system call functions, are collected for the task objective to form the FTun-Dataset dataset, and semantic optimization fine-tuning of the large model is carried out based on this. For the public vulnerability function dataset, in order to avoid overfitting of the detection results after model fine-tuning caused by reusing the above-mentioned vulnerability detection task dataset, the vulnerability dataset Vul-Dataset is constructed by collecting the well-known public software vulnerability function datasets Juliet Test Suite, Draper VDISC, and ComVul on the Internet and performing a union operation. For the system call function dataset Syscall-DataSet, the general system call function dataset ComSysCall is collected locally, and the system call sequence dataset ADFA-LD for malware detection publicly available on the Internet is collected to form a union. The specific names of the collected datasets and the proportion of construction are shown in Table 4. Since the decompilation semantic optimization large model performs high-level abstract syntax and semantic completion tasks for non-obfuscated vulnerability functions, the construction of this task dataset only divides the datasets for compilation under different architectures, without considering different compilers, compilation optimization methods, function inlining, etc. The composition of the dataset FTun-DataSet is shown in Table 4.

[0109] Table 4 LLM4Decompile Decompilation Large Model Fine-tuning Dataset

[0110]

[0111]

[0112] (2) Experimental Design and Analysis of Experimental Results

[0113] The method designed in the present invention is for vulnerability detection of non-obfuscated binary programs. Its essence is still binary similarity analysis. Therefore, in the experimental task design, the settings of the experimental work in YANG et al. are referred to. A total of three experiments are designed to evaluate the performance of the method of the present invention in multiple aspects, namely, the binary program similarity analysis experiment and the vulnerability detection experiment, which respectively evaluate the performance of the method in the ordinary similarity analysis scenario and the performance in the vulnerability detection scenario. Finally, the selection of the similarity threshold parameter and the time performance are also experimentally evaluated. And since the method of the present invention conducts experiments at the decompilation level and has a high degree of abstraction, the situations that may affect the experimental analysis performance at the assembly level or binary code level, such as cross-compiler and cross-compilation optimization, are not discussed. Similarly, for the similarity analysis and vulnerability detection of non-obfuscated programs, the non-obfuscated situation is not within the scope of discussion of the present invention. Further research on the vulnerability detection of obfuscated programs will be carried out later.

[0114] (1) Similarity Analysis Experiment and Performance Evaluation

[0115] The focus of this task is to evaluate the homology analysis ability of the model of the present invention, that is, the similarity comparison and classification ability. The main purpose is to classify function pairs as homologous or non-homologous pairs. This task uses three indicators to evaluate performance: TPR, FPR, and AUC value. TPR and FPR are usually used to measure the performance of binary classification models, while AUC provides an overall performance measure of the model's discrimination ability. The performance of Zeus under the same architecture is shown in Table 5. The AUC value of the homologous detection of binary functions under the same architecture basically reaches about 0.99, and the average detected AUC value is also as high as 0.995, indicating that the similarity analysis detection performance of the model under the same architecture basically reaches the current state-of-the-art level. In the cross-architecture detection, as shown in Table 6, the Zeus method achieves the highest performance among all cross-architecture detection methods. Its average detected AUC value reaches 0.987, and all methods achieve the best performance in the cross-architecture detection of X86-X64, proving that the structures of the X86 architecture and the X64 architecture are more similar after compiling binary programs. At the same time, the performance of Diaphora and VulSeeker is poor in cross-architecture detection, proving that the cross-architecture detection scenario and its impact on the detection performance of their methods may not have been considered in their design stage.

[0116] Table 5 Experimental Results of Similarity Analysis Task under the Same Architecture

[0117]

[0118] Table 6 Experimental Results of Similarity Analysis Task under Cross-Architecture

[0119]

[0120]

[0121] (2) Vulnerability Detection Experiment and Performance Evaluation

[0122] The vulnerability detection experiment mainly evaluates the performance of the method of the present invention in the non-obfuscated program vulnerability dataset. For the Zeus method designed by the present invention, three experiments of non-obfuscated program vulnerability detection are carried out under the same architecture, cross-architecture and real demand scenarios to comprehensively evaluate the performance of Zeus in vulnerability detection. Among them, a baseline method is designed in the cross-architecture as a performance comparison for Zeus. In the non-obfuscated program vulnerability detection experiment under the real demand scenario, the real firmware programs containing or mostly containing the types and versions of software packages with CVE vulnerabilities extracted are extracted, unpacked, compiled through the same architecture / cross-architecture, and finally input into Zeus for vulnerability detection. The three experiments comprehensively can reflect the performance of Zeus in non-obfuscated program vulnerability detection.

[0123] A. Same-architecture Vulnerability Detection Task

[0124] Table 7 shows the detection performance of the method of the present invention under the same architecture. By dividing the dataset of the same type of CVE vulnerabilities in different software packages under the same architecture in the vulnerability dataset into a training set and a test set according to a ratio of 8:2, the detection performance of different CVE vulnerabilities under the same architecture is tested. Among them, the average detection performances of the method of the present invention, Zeus, on the metrics of MRR, Recall@Top1 and Recall@Top10 are 0.980, 0.980 and 0.987 respectively. These three metrics comprehensively reflect the ranking of the correct results of Zeus's vulnerability detection under the same architecture, and the detection performance basically reaches about 0.98.

[0125] Table 7 Experimental Results of the Same-architecture Vulnerability Detection Task

[0126]

[0127] B. Cross-architecture Vulnerability Detection Task

[0128] In terms of the vulnerability detection performance across architectures, for the same type of CVE vulnerabilities under the same software package, the architecture types of the training set and the test set are determined according to the construction of the vulnerability datasets for different architectures in Table 3. At the same time, the CVE vulnerabilities with losses during the construction of the datasets for different architectures are filtered to ensure the fairness of the datasets in cross-architecture model training and testing, thereby reducing unnecessary errors. As shown in Table 8, in the cross-architecture vulnerability detection experiment, six baseline methods are selected as the comparison methods. These six methods are also classic methods in similarity detection or vulnerability detection and all have high performance. At the same time, six cross-architecture comparison experiments are designed under four architectures X86, X64, ARM, and MIPS commonly used in different real-world environments, including X86-ARM, X86-X64, X86-MIPS, ARM-X64, ARM-MIPS, and X64-MIPS.

[0129] According to the data analysis reflected in Tables 8 - 10, it can be seen that Zeus has comprehensively higher detection performance than other baseline methods in the six cross-architecture experiments. The cross-architecture average detection performance values of Zeus in the three performance metrics of MRR, Recall@Top1, and Recall@Top10 are 0.908, 0.898, and 0.940 respectively. By horizontally analyzing each method in the table, it can be found that the performance metric values of different methods under X86-X64 are higher than those of other methods. This may benefit from the similarity in compilation between the X86 and X64 architectures, as well as the compatibility of most software with these architectures in the real-world environment. The second-best performing is X86-MIPS, followed by ARM-MIPS, X64-MIPS, ARM-X64, and the worst-performing is X86-ARM. This may be because the compilation effect of its architecture on the dataset varies greatly, resulting in changes in its vulnerability characteristics. Vertically analyzing Table X, it can be found that except for Zeus having the best performance, FASER also has excellent cross-architecture performance. However, TREX, Gemini, VulSeeker, SAFE, and Diaphora all have poor performance in cross-architecture vulnerability detection. The possible reasons are that some methods did not consider cross-architecture vulnerability detection during design, or the methods are only applicable to homologous analysis through binary similarity comparison and have not been experimented with in vulnerability detection, or some other dataset or compilation factors.

[0130] Table 8 Evaluation of the MRR Metric for the Vulnerability Detection Task under Cross-Architecture

[0131]

[0132] Table 9 Evaluation of the Recall@top1 Metric for the Vulnerability Detection Task under Cross-Architecture

[0133]

[0134] Table 10 Evaluation of Recall@top10 Index for Vulnerability Detection Tasks across Architectures

[0135]

[0136] C. Firmware Vulnerability Detection Tasks in the Real Environment

[0137] Based on the same-architecture and cross-architecture vulnerability detection experiments, a firmware program vulnerability detection experiment in the real environment was set up. By extracting the software programs in the real firmware and inputting them into Zeus for detection value statistics to measure its detection performance in the real environment. Since there are many types of CVE vulnerabilities to be counted, in this experiment, only the CVE vulnerabilities with the greatest harm and the widest affected versions in each software package were extracted. As shown in Table 11, the CVE numbers, the ranges of affected software versions, and the vulnerability type information that have a greater impact on each software package are listed. A total of six different types of CVE vulnerabilities were extracted, with one CVE vulnerability for each software, and most of them are high-risk vulnerability types.

[0138] For the selected firmware programs to be detected, by counting the products of common firmware manufacturers in reality and the firmware devices that often use the collected software packages, a total of 10 types of firmware devices were obtained, a total of 1522 firmware packages were extracted, a total of 2952 binary programs were unpacked from the firmware, and a total of 5327551 functions were extracted from the binary programs, corresponding to the first to fourth columns of Table 12 respectively. And for the collected firmware programs, the number of them that conform to the collected software package dataset was counted. Among them, there are 1197 OpenSSL software packages, 1293 Busybox, 305 Dnsmasq, 144 Lighttp, 178 Tcpdump, 124 Nginx, and 134 Docker. The number of software packages extracted from each firmware corresponds to the data in the fifth to eleventh columns of Table 12 respectively. Since most of the firmware in Table 12 has a small number of software packages that can be collected and extracted, it may not be able to form a good detection effect. Therefore, finally, the five types of firmware with the largest number of extracted software packages, namely Netgear, TP-Link, Hikvision, Cisco, and ASUS, were selected for subsequent detection experiments.

[0139] Table 11 Sources of CVE Vulnerabilities in Firmware Vulnerability Detection Tasks

[0140]

[0141] Table 12 Statistical Information on Firmware Vulnerability Detection Tasks

[0142]

[0143] The detection objective of Experiment C is to test whether Zeus can detect the specified CVE vulnerabilities contained in the firmware program and its detection performance. By inputting the software packages collected for each firmware into Zeus for decompilation semantic optimization, feature extraction and encoding, and similarity analysis calculation, Zeus finally determines whether it is a CVE vulnerability according to the detection threshold. Table 13 shows the vulnerability detection results, where O represents Origin, that is, the number of specified CVE - numbered vulnerabilities existing in the software packages extracted from the original firmware. This part is pre - marked and counted according to the publicly available vulnerability information during the software extraction stage and is used to finally measure the detection performance of the method. F represents Find, that is, the number of vulnerabilities with the specified CVE number detected by Zeus. F / O represents Find / Origin, that is, the ratio of the number of CVE vulnerabilities detected by the model to the number of originally marked CVE vulnerabilities in the software package, which is the detection rate and is also the index for Experiment C to measure the performance of Zeus. The higher the detection rate, the better the detection performance of the method of the present invention in the real environment. Through horizontal analysis of Table 13, Zeus has a relatively high detection rate for each specified CVE vulnerability under the specified five firmwares. The average detection rates of Netgear, TP - Link, Hikvision, Cisco, and ASUS are 96.6%, 95.1%, 96.2%, 95.2%, and 92.2% respectively. And from vertical analysis, for the detection rate of the same specified CVE vulnerability, each firmware also has a relatively high detection rate. The average detection rates of Netgear, TP - Link, Hikvision, Cisco, and ASUS are 96.3%, 94.8%, 92.9%, 96.9%, 97.6%, and 94.4% respectively. Even for some firmwares, the detection rate reaches 100% for some specified CVE vulnerability detections. This may be because the number of software packages extracted is small or the CVE vulnerability features contained are more obvious.

[0144] Table 13 Experimental Results of Firmware Vulnerability Detection Task

[0145]

[0146] Meanwhile, the detailed analysis of the detection quantities of the six specified CVE vulnerabilities in Experiment C is shown in a mixed bar - line chart Figure 6 As shown, the bar chart can reflect the statistical and comparative situations of the number of CVE vulnerabilities detected by each firmware. For the bar chart, the X - axis represents the CVE vulnerability number, and the Y - axis represents the corresponding number of CVE vulnerabilities detected by Zeus Figure 6the left main axis in [reference], where Netgear detected the largest number of CVEs in CVE-2021-3449. The main reason is that there are more OpenSSLs with this CVE vulnerability in the corresponding firmware. The number of vulnerabilities detected for CVE-2021-45078 and CVE-2021-41581 in Hikvision firmware is 0. By comparing the original vulnerability situation in Table 13 and the software situation extracted in Table 12, the main reason is that Hikvision basically did not use the software packages corresponding to the vulnerabilities or the corresponding software package versions. And Figure 6 The line chart in [reference] can reflect the number of each CVE vulnerability detected by Zeus, that is, the number of corresponding software packages with this vulnerability. It can be seen from the line chart that the number of CVE vulnerabilities corresponding to the Busybox software is the largest, while the number of CVE vulnerabilities corresponding to the Lighttpd software is the smallest. The overall line trend is basically proportional to the number of CVEs originally included in each software in Table 13.

[0147] To further analyze the performance of Zeus in vulnerability detection, software packages of different versions of OpenSSL in the firmware were extracted and vulnerability detection was performed through Zeus. The main detection targets are the 15 CVEs that have had a greater impact on OpenSSL in the past decade collected in Table 2, and the OpenSSL software version - CVE vulnerability ID was analyzed to further explore the vulnerability distribution under different versions of OpenSSL in different firmwares. As Figure 7 Shown in [reference] is the heat map analysis of the detection results. Among them, subgraphs (a), (b), (c), (d), and (e) respectively represent the analysis results of the OpenSSL vulnerability versions corresponding to different firmware manufacturers. The X-axis represents different OpenSSL versions, and the Y-axis represents the CVE IDs corresponding to 15 CVEs. Each square in each subgraph represents the number of OpenSSL versions corresponding to each CVE vulnerability. The darker the color of the square, the larger the number it contains. The numerical metrics corresponding to different colors are reflected on the vertical axis on the right side of each subgraph.

[0148] Specifically to Figure 7Analyze each subfigure. Subfigure (a) shows that the OpenSSL 1.1.1d version is widely used by Netgear, resulting in a large number of CVE vulnerabilities. Among them, vulnerabilities such as CVE-2021-3711, CVE-2021-3712, CVE-2020-1971, and CVE-2022-0778 have a general impact on all the OpenSSL versions counted. For the TP-Link firmware in subfigure (b), its main used version also focuses on OpenSSL 1.1.1d, and the overall vulnerability analysis and the influence of different vulnerabilities are basically the same as those of Netgear. In subfigure (c), the overall vulnerability distribution mainly concentrates in the range of CVE-2020-1971 to CVE-2021-3711, and the vulnerability distribution within this range basically covers all versions of OpenSSL. The number of detected vulnerabilities of CVE-2021-3712 in the OpenSSL 1.1.1i version reaches Figure 7 the highest among all subfigures. These phenomena may be because the wide use of Hikvision in life makes it the main target of CVE vulnerability mining. In subfigures (d) and (c), the overall number of vulnerabilities shown for Cisco and ASUS is relatively small. In subfigure (d), Cisco has a relatively large number of CVE vulnerabilities in the versions from 1.1.1i to 1.1.1k, and the vulnerability influence is extensive in the range of CVE-2020-1971 to CVE-2021-3711. In subfigure (e), the overall number of detected vulnerabilities for ASUS is less, and they are concentrated in the CVE-2021-3712 vulnerability numbered in the OpenSSL 1.1.1e and OpenSSL 1.1.1f versions. Through the analysis of each subfigure, it can be found that Figure 7 the vulnerability distribution of each firmware's OpenSSL version basically concentrates in the range of CVE-2020-1971 to CVE-2021-3711. This is because the vulnerabilities in this range are basically medium-risk or above, with greater damage and a wider distribution of versions. The vulnerabilities in other ranges are mostly low-risk or have a smaller distribution of versions. Among the OpenSSL software versions, 1.1.1d, 1.1.1i, and 1.1.1e have a relatively large number of vulnerability distributions. And CVE-2021-3712 has the highest number of vulnerability distributions in all the analyzed subfigures, with the largest quantity and the widest distribution among all the counted vulnerabilities and software versions.

[0149] (3) Selection of similarity threshold parameter and time performance evaluation

[0150] A. Selection of similarity threshold parameter θ

[0151] In similarity analysis and calculation, the θ parameter is mainly used to measure the similarity scoring results. Therefore, the selection of θ will affect the overall detection and analysis effect of the model. A sensitivity analysis experiment was conducted on the selection of θ, and the results are shown in Table 14. As θ increases, the similarity detection threshold of the model also continues to rise, and the statistical performance parameters for the similarity analysis task and the vulnerability detection task also continue to improve. When θ is around 0.6, the performance of the similarity analysis task SimTask reaches the highest, with SameArch.AUC = 0.995 and DiffArch.AUC = 0.987. At the same time, the three statistical parameters of the vulnerability task VulTask, namely MRR, Recall@Top1, and Recall@Top10, also reach 0.987, 0.898, and 0.940 respectively. After θ = 0.6, as the value of θ increases, the performance of the statistical indicators of the SimTask gradually decreases. For the VulTask, since its statistical indicators are calculated based on the ranking of the model detection results and the correct detection rate of the top K (Top1 or Top10) detections, the higher the θ, the higher the detection rate and ranking of the detected vulnerabilities, and the corresponding values of the three performance indicators will also be higher. However, too high a value of θ will also cause the false negative rate of the model to rise rapidly. Therefore, considering the performance data of the two tasks and the requirements for the SimTask and VulTask in the actual environment, the final selection is θ = 0.6 as the threshold value for measuring the similarity score of the Siamese calculation.

[0152] Table 14 Experimental analysis results of different similarity threshold parameters

[0153]

[0154] B. Time performance evaluation of similarity analysis and vulnerability detection tasks

[0155] For the evaluation of the time performance of the method of the present invention, by comparing the training and evaluation of the selected baseline methods on the same dataset, the average time consumption statistics of the similarity analysis task and the vulnerability task are collected respectively. Since the fine-tuning stage of the decompiled large model aims to improve the quality of the dataset and the diversity of the dataset preprocessing of other methods, the time of the preprocessing stage that does not affect the detection performance of the trained model is not counted. Only the average time consumption of the feature extraction stage, the feature encoding stage, and the similarity calculation stage of each method in the two tasks is counted as the time performance value of the method for comparison, as Figure 8 shown, where Figure 8(a) Intuitively shows the time consumption comparison of each method under two tasks using a horizontal bar chart. It can be found that the method Zeus of the present invention has the lowest time consumption under both tasks, approximately 56s and 42s, mainly because its high-level abstraction at the decompilation level makes its feature extraction, encoding, and similarity calculation more concise, while other baseline methods basically show the time consumption costs as presented in their corresponding papers. And in Figure 8 (b), by using a radar chart, it can be intuitively reflected that among all methods, the time consumption cost of the vulnerability detection task exceeds that of the similarity analysis task. The time consumption of the vulnerability detection task is distributed in the range of 50s to 350s, while the time consumption of the similarity analysis task is distributed in the range of 0s to 350s. This distribution also reflects that in real-world scenarios, vulnerability detection often faces cross-comparison of various vulnerability features and sorting of vulnerability detection results, which often takes longer, rather than just performing homology analysis in the similarity analysis task.

[0156] The present invention proposes the Zeus method, a vulnerability detection method based on decompilation semantic optimization of a large model and EAST feature encoding, aiming to perform cross-architecture similarity analysis and vulnerability detection on large-scale non-obfuscated binary programs at the decompilation level. Zeus realizes the semantic and syntactic optimization of the dataset after Ghidra decompilation by performing semantic optimization fine-tuning on the decompilation large model. Then, by extracting EAST features from the optimized dataset, the semantic and syntactic structure information of the features is further improved. After that, it is input into the EASTNN model for encoding learning, and finally input into the Siamese model for similarity calculation. Experiments show that Zeus has high performance in large-scale similarity analysis tasks, vulnerability detection tasks, and firmware vulnerability detection in real-world scenarios, and can achieve cross-architecture / same-architecture similarity analysis, cross-architecture / same-architecture vulnerability detection for existing non-obfuscated binary programs, and efficient vulnerability detection and analysis for real firmware programs. At the same time, the time consumption is also much lower than that of other baseline methods.

[0157] Finally, it should be noted that the above-described embodiments are only specific implementation manners of the present invention, used to illustrate the technical solutions of the present invention, rather than to limit it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent substitution on some of the technical features; and these modifications or substitutions do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be determined by the protection scope of the claims.

Claims

1. A binary program vulnerability detection method based on a decompilation large model and EAST features, characterized in that: include: Use the vulnerability function dataset and the system call function dataset to fine-tune the pre-trained decompilation model LLM4Decompile, so that LLM4Decompile can optimize the syntax and semantics of the decompiled pseudocode output by the decompiler Ghidra, and obtain the fine-tuned decompilation model through iterative training; Extract the corresponding enhanced abstract syntax tree EAST from the decompiled pseudocode after fine-tuning and optimizing the syntax and semantics of the decompiled large model, and perform feature encoding on the enhanced abstract syntax tree using the EASTNN model; The feature vectors of the encoded non-obfuscated binary program to be detected and the vulnerability function data set are input into the Siamese network for similarity calculation, and the vulnerability of the non-obfuscated binary program is evaluated based on the similarity calculation results.

2. The binary program vulnerability detection method based on the decompilation large model and EAST features according to claim 1 is characterized in that: The process of fine-tuning the pre-trained decompiled large model LLM4Decompile is as follows: Load the pre-trained decompiled large model M LLM And initialize the fully connected layer to prepare for fine-tuning of the large model; Start iterative training of the large model, starting from the vulnerability function dataset D v and system call function dataset D s Sample a batch of samples B v and B s ,Each sample consists of a vulnerability function and a system call function; Use the decompiler Ghidra to decompile the original binary input and get the pseudo code C v ,C s ; Use the model forward propagation method to decompile the pseudo code C v and C s Enter M LLM , get the decompilation results after optimizing the grammar and semantics of vulnerabilities and system functions; Calculate the loss function, then update the model, use the gradient descent method to update the model parameters, and output the fine-tuned decompiled large model M after completing multiple iterations. fine-tuned .

3. The binary program vulnerability detection method based on the decompilation large model and EAST features according to claim 2 is characterized in that: During the fine-tuning of a large model, the loss function calculation includes: For the syntax and semantic optimization of the vulnerability function, calculate M LLM Prediction result P v With P v Corresponding vulnerability source code annotation dataset Y v Cross entropy loss; for system call function syntax and semantic optimization, calculate M LLM Prediction result P s With P s Corresponding system call source code annotation dataset Y s The cross entropy loss of , combines the two loss functions as the total loss through weight distribution.

4. The binary program vulnerability detection method based on the decompilation large model and EAST features according to claim 1 is characterized in that: The extraction of the corresponding enhanced abstract syntax tree EAST from the decompiled pseudocode after the fine-tuning and optimization of the syntax and semantics of the decompiled large model specifically includes: first, preprocessing the pseudocode decompiled from the large model; then, normalizing the pseudocode decompiled from the large model; and finally, constructing an enhanced abstract syntax tree that has been normalized and completed in syntax and semantics.

5. The binary program vulnerability detection method based on the decompilation large model and EAST features according to claim 4 is characterized in that: Preprocessing operations on the pseudocode after decompilation of the large model include removing invalid characters, basic syntax completion and placeholder marking; code standardization processing on the pseudocode after decompilation of the large model includes variable type completion and inference, as well as function return type completion and inference.

6. The binary program vulnerability detection method based on the decompilation large model and EAST features according to claim 1 is characterized in that: The feature encoding of the enhanced abstract syntax tree using the EASTNN model includes: By using the first bifurcation point of the enhanced abstract syntax tree EAST as the split point, it is split into different parts of EAST, namely p1…p n , and then input the function-level encoder, the output vector e1…e t The EAST encoding vector V is finally generated through a bidirectional gated recurrent unit and pooling.

7. The binary program vulnerability detection method based on the decompilation large model and EAST features according to claim 1 is characterized in that: The feature vectors of the encoded non-obfuscated binary program and vulnerability function data set to be detected are input into the Siamese network for similarity calculation. The process is as follows: Siamese first uses EASTNN to encode the EAST syntax tree features T1 and T2 into fixed-length feature vectors V1 and V2, and then inputs the feature vectors V1 and V2 into the cosine function COS(V1, V2) for similarity calculation to obtain the specific score s.

8. The binary program vulnerability detection method based on the decompilation large model and EAST features according to claim 7 is characterized in that: The similarity judgment threshold is set to θ. When the similarity calculation score s≥θ, it is judged to be similar and the non-obfuscated binary program to be detected has a vulnerability; when the similarity calculation score s<θ, it is judged to be dissimilar and the non-obfuscated binary program to be detected does not have a vulnerability.

9. A binary program vulnerability detection system based on a decompilation large model and EAST features, characterized in that: The system is used to implement the binary program vulnerability detection method based on the decompiled large model and EAST features as described in any one of claims 1 to 8, and comprises a large model fine-tuning module, an EAST feature encoding module and a similarity calculation module, wherein: The large model fine-tuning module is used to fine-tune the pre-trained decompilation large model LLM4Decompile using the vulnerability function dataset and the system call function dataset, so that LLM4Decompile optimizes the syntax and semantics of the decompiled pseudo code output by the decompiler Ghidra, and obtains the fine-tuned decompilation large model through iterative training; The EAST feature encoding module is used to extract the corresponding enhanced abstract syntax tree EAST from the decompiled pseudo code after the fine-tuning and semantic optimization of the decompiled large model, and perform feature encoding on the enhanced abstract syntax tree using the EASTNN model; The similarity calculation module is used to input the feature vectors of the encoded non-obfuscated binary program to be detected and the vulnerability function data set into the Siamese network for similarity calculation, and evaluate the vulnerability of the non-obfuscated binary program according to the similarity calculation results.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Cited By

  • A training method of a binary code anti-obfuscation expert large model

    CN122778355A