Software vulnerability automatic repair method based on bidirectional pre-training
Through the software vulnerability automatic repair method based on two-way pre-training, the multi-head self-attention mechanism of BPE word segmentation and relative position encoding is used to solve the problem of insufficiently accurate processing of long sequences in the existing technology, and efficient and accurate vulnerability repair is achieved, and the security and stability of the system are improved.
Patent Information
- Application Number
- CN202411859946.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art is not accurate enough when processing long sequences, resulting in errors in the model when focusing on marking positions and unable to effectively fix software vulnerabilities.
The software vulnerability automatic repair method based on two-way pre-training is adopted. By crawling the vulnerability and its corresponding patches, data cleaning and BPE word segmentation are used for pre-processing, and the built vulnerability repair model is input after vectorization conversion for repair. The model is built on the T5 model, using stacked encoder blocks, multi-head self-attention mechanisms encoded by relative position, and feedforward neural networks to generate vulnerability repair codes through the bundle search algorithm.
It improves the accuracy and efficiency of vulnerability repair, enhances the applicability and universality of the model, can more accurately understand and handle the relationship between markers, and significantly improves the security and stability of the system.
Smart Images

Figure CN120012096A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer technology, and in particular relates to software source code vulnerability repair, and specifically is a software vulnerability automatic repair method based on bidirectional pre-training. Background Art
[0002] With the rapid development of information technology, software plays an important role in all aspects of the world's economy, military, society, etc. At the same time, potential security issues in software are becoming an emerging global challenge. Software vulnerabilities are one of the root causes of security issues. Highly skilled hackers can exploit software vulnerabilities to do many harmful things at their will, such as stealing users' private information and stopping critical equipment. According to statistics released by the CVE (Common Vulnerabilities and Exposures) organization, the number of software vulnerabilities discovered in 2000 was less than 4,600, while the current number of vulnerabilities has exceeded 25,000. According to the vulnerability and threat trend report of SKYBOXSECURITTY, the number of vulnerabilities has surged in the past five years, breaking all previous records. The frequency and destructiveness of cyber attacks are increasing. Malicious attacks and behaviors are becoming more and more complex, and are using more powerful tools and more advanced strategies. These attacks not only affect companies and governments, but also individuals. Therefore, effective methods must be identified to prevent the consequences of these attacks. Usually, attackers directly invade remote computers by exploiting existing vulnerabilities in hardware, software, and networks, posing serious risks to cyberspace security. To combat hackers and prevent systems from being hacked, defenders must find vulnerabilities and fix them before attackers do. Due to the dramatic increase in the number of detected vulnerabilities and the complexity of modern software systems, manually fixing security vulnerabilities is extremely time-consuming and labor-intensive for security experts. Studies have shown that 50% of vulnerabilities have a life cycle of more than 438 days. Delayed vulnerability patching may lead to continued attacks on software systems, causing economic losses to users. System developers have studied various methods for effectively detecting and fixing vulnerabilities. For software projects, software maintenance costs account for more than 50% of the total project cost. Fixing software vulnerabilities accounts for 20% of all maintenance activities, and the efficiency of fixing vulnerabilities will directly affect the cost of system development and maintenance.
[0003] Current popular vulnerability repair methods still have certain limitations. For example, the VRepair model uses word-level tokenization and replication mechanisms to handle Out-Of-Vocabulary (OOV), which limits the model's ability to generate new tokens that have never appeared in vulnerable functions, and uses absolute position encoding, which limits the ability of the self-attention mechanism to learn the relative position information of code tokens. Absolute position encoding may not be accurate enough when processing long sequences, causing the model to make mistakes when paying attention to token positions, such as focusing on the wrong token. Summary of the invention
[0004] The present invention provides a method for automatically repairing software vulnerabilities based on bidirectional pre-training, so as to solve the problem in the prior art that the method is not accurate enough when processing long sequences, resulting in errors in the model when focusing on the marked position.
[0005] In order to achieve the above object, the technical solution provided by the present invention is: a method for automatically repairing software vulnerabilities based on bidirectional pre-training, comprising the following steps:
[0006] Step 1: Crawl vulnerabilities and their corresponding patches as data sets;
[0007] Step 2: Use data cleaning and BPE segmentation to preprocess the data in the dataset;
[0008] Step 3: Use Word Embedding to vectorize the preprocessed data;
[0009] Step 4: Input the vectorized data into the constructed vulnerability repair model and repair the vulnerability code.
[0010] Furthermore, the specific steps of using the BPE method to segment the data set in the above step 2 are: first, the vulnerability data set is divided into single characters, and then the pair of characters with the highest frequency is replaced with another character in turn until the number of cycles is completed.
[0011] Furthermore, in the above step 4, the vulnerability repair model is constructed based on the T5 model, including the following steps:
[0012] 4.1. Build a vulnerability repair model based on the CodeT5 pre-trained language model;
[0013] 4.2. We use stacked encoder blocks, combined with a multi-head self-attention mechanism and a feedforward neural network for relative position encoding, to capture the relative position information and semantic relationship of the input sequence through residual connections and layer normalization.
[0014] 4.3. By stacking twelve layers of decoder blocks, combining the mask mechanism to limit the prediction range, and using the cross-entropy loss to optimize the model, it only focuses on the context tags when generating vulnerability fix code, and monitors the training process through the validation set to select the best model;
[0015] 4.4. Use a 12-layer Transformer encoder and decoder, combined with a linearly scheduled AdamW optimizer and a Softmax layer, to convert the generated vector into a probability distribution, and fine-tune the learning rate and loss function;
[0016] 4.5. Using the beam search algorithm, by setting the beam width parameter β, the first β vulnerability repair candidates with the highest conditional probability are selected at each time step, and the search is terminated when the EOS mark is generated.
[0017] Furthermore, in 4.2 above, each encoder block consists of two subcomponents, a multi-head self-attention layer with relative position encoding, followed by a feed-forward neural network.
[0018] Furthermore, in 4.2 above, the multi-head self-attention mechanism is summarized as:
[0019]
[0020]
[0021] Where: Q, K, V represent the input query, key and value matrices, W i represents the trainable weight matrix for linear transformation, Concat represents the concatenation of the outputs of multiple heads, Head represents the number of attention heads, and W O Used for linear projection to the expected dimension after concatenation.
[0022] Furthermore, in 4.3 above, the cross entropy loss (H(p,q) = -∑ x∈X p(x)logq(x)) to update the model and optimize the difference between each position in the predicted sequence and the true sequence, where X is the set of classes, p is the true probability distribution, and q is the predicted probability distribution.
[0023] Compared with the existing methods, the present invention has the following advantages:
[0024] 1. The present invention pre-processes the vulnerability functions and their corresponding repair patches in the vulnerability source code by using word tokenization to obtain a more suitable and less redundant vocabulary as the input of the vulnerability repair model, which can improve the repair accuracy, optimize the repair efficiency, and enhance the applicability and versatility.
[0025] 2. In step 4, the present invention proposes a model for vulnerability repair, and adopts relative position coding in position coding, so that the model can better capture the relative position and distance information between markers, improve the accuracy of the model, and enable the model to more accurately understand and process the relationship between markers when generating vulnerability repairs.
[0026] 3. The technical points of the present invention work together to increase the scale and diversity of training data, use the BPE subword-level tokenization method, and use relative position encoding, so that the model can learn a wider range of contextual information, reduce the size of the vocabulary, and effectively improve the model's ability to process new tags, so that the model can more accurately understand and process the relationship between tags when generating vulnerability repairs. Compared with traditional vulnerabilities, the repair efficiency and accuracy are high, which can effectively improve the security and stability of the system and reduce labor costs; because the vulnerability repair process is fast, it can effectively respond to the rapidly changing threat environment, improve the overall anti-attack ability of the software system, shorten the vulnerability exposure period, reduce potential risks, and make the software development process faster and more reliable. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 It is a flowchart of the present invention;
[0028] Figure 2 The figure is a general structure diagram of the vulnerability code automatic repair method in the present invention;
[0029] Figure 3 It is a comparative experimental diagram of the present invention;
[0030] Figure 4 This is an ablation experiment diagram of the present invention. DETAILED DESCRIPTION
[0031] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below with reference to the accompanying drawings and embodiments.
[0032] Example 1, the present invention provides a method for automatically repairing software vulnerabilities based on bidirectional pre-training, see Figure 1 , which includes the following steps:
[0033] Step 1: Crawl vulnerabilities and their corresponding patches as a dataset:
[0034] Use crawler scripts to select vulnerability data from open source vulnerability databases. The largest open source vulnerability databases, CVE and NVD, have statistics on various types of vulnerability information, including vulnerability descriptions, risk levels, impact levels, disclosure times, and vulnerability source codes. This information is very important for subsequent vulnerability repair research.
[0035] Step 2: Use data cleaning and BPE word segmentation to pre-process the data: Since the data obtained by the crawler is redundant and lacks hierarchical logical relationships, it is not conducive to subsequent repair work, and the non-essential parts of the code may mislead the repair results of the model. Therefore, the present invention uses data cleaning and BPE word segmentation methods to pre-process the data.
[0036] Step 2.1 Data cleaning:
[0037] The vulnerability source code in the open source software project has the characteristics of multi-source heterogeneity due to the difference in the acquisition method, which will result in some collected codes only having changes in the project version number, while the code itself has not been modified. At the same time, other interference factors may appear. For example, the vulnerability source code link corresponding to some CVE-IDs should have the vulnerability source code, but after the web crawler crawls the path, it is found that the obtained vulnerability code file is blank, or some vulnerability information data contains garbled characters. In addition, because there are some special vulnerabilities without specific vulnerability source code links, you can manually find the corresponding source code and patch files of the vulnerabilities in GitHub or its official website.
[0038] Step 2.2 Data segmentation:
[0039] In order to improve the performance and processing capability of the model, the present invention uses BPE (word tokenization) to perform word segmentation preprocessing. Figure 2 ,When using the BPE method to segment the dataset: first divide the vulnerability dataset into single characters, and then replace the most frequent pair of characters with another character in turn, until the number of cycles ends.
[0040] The specific process is as follows:
[0041] ① Obtain the corpus, for example: "FloydHub is the fastest way to build, train and deploy deep learning models. Build deep learning models in the cloud. Traindeep learning models."
[0042] ② Split, add suffixes, and count word frequency
[0043] Table 1 Statistical word frequency
[0044]
[0045] ③Build a vocabulary and count character frequencies
[0046] Table 2: Building vocabulary
[0047]
[0048] ④ Taking the above table as an example, replace the characters d and e with de, and then iterate in sequence.
[0049] Table 3 First iteration
[0050]
[0051] Continue iterating until the preset subwords vocabulary size is reached or the next most frequent byte pair has a frequency of 1.
[0052] In this way, a more suitable vocabulary is obtained, which may contain some combinations that are not words, but are meaningful in themselves. This step determines how to split the word by generating a merge operation, and then performs a merge operation based on the subword vocabulary. In order to adapt to the code generation task, in this embodiment: " <s> "and"< / s> "Markers to indicate the beginning of a sequence (BOS) and the end of a sequence (EOS)." <pad>" token is used to pad the input sequences to the same length when needed. In addition, four special tokens are added to the vocabulary (i.e., " <startloc> "," <endloc> "," <modstart> "," <modend>") as additional vocabulary IDs so that they are not split into subcomponents during tokenization and ensure that the generated vectors are more generalizable.
[0053] Step 3: Use Word Embedding for vectorized representation:
[0054] Use Word Embedding to convert the vectors and generate an embedding vector of [1x768] for each subword token, and combine them into a matrix to represent the meaningful relationship between a given code token and other code tokens. In order to capture the semantic meaning of the code tokens, use the word embedding vector pre-trained on the same corpus as the pre-trained tokenizer above. And use relative position embedding to capture the position of each code token in the function and add it to the key matrix and value matrix.
[0055] Step 4: Input the vectorized data into the constructed vulnerability repair model and repair the vulnerability code, including the following steps:
[0056] Step 4.1 Build a vulnerability repair model based on the CodeT5 pre-trained language model:
[0057] The matrix generated in step 3 is input into the encoder stack, and the output of the last encoder is input into each decoder; the output of the decoder stack is then input into a linear layer with softmax activation to generate the probability distribution of the vocabulary; finally, beam search is used on the probability distribution of the vocabulary to generate the final candidate as the prediction result.
[0058] 4.2 uses stacked encoder blocks, combined with a multi-head self-attention mechanism and a feedforward neural network for relative position encoding, to capture the relative position information and semantic relationship of the input sequence through residual connections and layer normalization to enhance the model's ability to represent code tags:
[0059] The present invention uses a stack of twelve layers of encoder blocks to derive the encoder hidden states used by the decoder. Similar to the original Transformer encoder, each encoder block starts with layer normalization, where the activation values are only rescaled and no additive bias is applied. Each encoder block consists of two subcomponents: a multi-head self-attention layer with relative position encoding, followed by a feedforward neural network. Each subcomponent in each encoder (i.e., self-attention and FFNN) has a residual connection around it, followed by a layer normalization step. The multi-head self-attention mechanism calculates the relevance score of each code token using a dot product operation, where each token interacts with itself and other tokens in turn. The multi-head self-attention mechanism relies on three main vectors: query, key, and value. The query is a representation of the current code token, which is used to score it based on all other tokens stored in the key vector. The attention score of each token is obtained by taking the dot product of all query vectors and the key vector. The attention score is then normalized to a probability through the Softmax function to obtain the attention weight. Finally, the value vector can be updated by taking the dot product of the value vector and the attention weight vector.
[0060] Different from using absolute position encoding layer and word embedding layer to capture the position information in the input sequence, the present invention uses relative position encoding to represent the relative position and distance between tokens in the input sequence (i.e., relation-aware self-attention mechanism). In the present invention, the self-attention mechanism is scaled dot product self-attention with relative position encoding. The self-attention operation is calculated using four matrices, namely Q, K, V, and P. The relative position information P is provided as an additional component to the Key matrix and the Value matrix, and the specific method is as follows:
[0061]
[0062] Where P is the edge representation of the two inputs used to determine the position information between tokens in the dot product operation. Instead of using absolute position encoding, relative position encoding generates different learned embeddings according to the offset between K and Q in the self-attention operation to capture the relative information between tokens.
[0063] In order to capture richer semantic meaning of the input sequence, the present invention uses a multi-head self-attention mechanism to implement self-attention, so that the model jointly pays attention to information from different code representation subspaces at different positions. For Q, K and V of d dimensions, the present invention divides these vectors into h heads, each with d / h dimensions. After completing all self-attention operations, each head is reconnected and input into a fully connected feedforward neural network containing two linear transformations and activated by ReLU. The multi-head self-attention mechanism can be summarized by the following formula:
[0064] MultiHead(Q,K,V)=Concat(Head 1 , ..., Head h )W O
[0065]
[0066] Where: Q, K, V represent the input query, key and value matrices, W i represents the trainable weight matrix for linear transformation, Concat represents the concatenation of the outputs of multiple heads, Head represents the number of attention heads, and W O Represents the linear transformation weight matrix of the output.
[0067] 4.3 By stacking twelve layers of decoder blocks, combining the mask mechanism to limit the prediction range, and using cross-entropy loss to optimize the model, it only focuses on context tags when generating vulnerability repair code, and monitors the training process through the validation set to select the best model:
[0068] Implement a stack of twelve layers of decoder blocks based on the hidden state generation bug fix provided by the last encoder block in the stack of twelve layers of encoder blocks in step 4.2. Similar to the original Transformer decoder, each decoder block starts with layer normalization, just like in the encoder block. Same as the encoder block, each subcomponent in each decoder has a residual connection around it and is followed by a layer normalization step. A mask mechanism is used during the training phase of the generative model to restrict the model from predicting the next token without paying attention to the subsequent context. Therefore, the model only pays attention to the previous tokens during the generation process.
[0069] During training, the cross entropy loss (H(p,q) = -∑ x∈X p(x)logq(x)) to update the model and optimize the difference between each position in the predicted sequence and the true sequence, where X is the set of classes, p is the true probability distribution, and q is the predicted probability distribution. To obtain the best fine-tuning weights, the training process is monitored epoch by epoch using the validation set, and the best model is selected based on the best loss value on the validation set (not the test set).
[0070] 4.4 Use a 12-layer Transformer encoder and decoder, combined with a linearly scheduled AdamW optimizer and a Softmax layer, to convert the generated vector into a probability distribution, and fine-tune the learning rate and loss function to improve the vulnerability repair effect:
[0071] The linear layer is a fully connected neural network that projects the vector produced by the decoder stack into a larger logits vector with a number of units equal to the number of unique tokens in the vocabulary. The subsequent Softmax layer then converts this value into a probability distribution that sums to 1, which will be used in step 3 to generate the final output.
[0072] For the vulnerability repair model architecture, 12 Transformer encoder blocks, 12 Transformer decoder blocks, 768 hidden layer size, and 12 attention heads were used. During fine-tuning, the learning rate was set to 2e-5 and a linear schedule was used, i.e., the learning rate decayed linearly throughout the training process. The present invention uses the AdamW optimizer for back propagation to update the model and minimize the loss function.
[0073] 4.5 Using the beam search algorithm, by setting the beam width parameter β, select the first β vulnerability repair candidates with the highest conditional probability at each time step, and terminate the search when the EOS mark is generated:
[0074] A beam search is used on the probability distribution of the vocabulary to select multiple vulnerability fix candidates for the input sequence at each time step based on conditional probability. The number of fix candidates depends on a parameter setting called beam width (β). The top β fix candidates with the highest probability are selected using a best-first search strategy at each time step. The beam search terminates when an EOS marker (i.e. "") is generated. The final candidate is output as the prediction result.
[0075] At this point, the vulnerability has been fixed.
[0076] The perfect prediction percentage (% Perfect Predictions) method is used to evaluate the vulnerability repair effect.
[0077]
[0078] And comparative experiments and ablation experiments are used to verify the technical effects of the present invention.
[0079] Experimental content: The present invention first uses a crawler script to crawl vulnerability files and their corresponding patch files from CVE and NVD public data, and uses manual screening to supplement the missing vulnerabilities and their corresponding patch files in the crawled vulnerabilities to obtain the data set of the present invention. The data set is segmented by BPE to generate a segmentation function (i.e., a list of subword code tags for each function) and input into the repair model. The model first performs word embedding (WordEmbedding) on the input vulnerability function to generate an embedding vector and combines it into a matrix. Then, the matrix is input into the model encoder stack, and the output of the last encoder is input into each of the other decoders, and then the output of the decoder stack is input into a linear layer with softmax activation to generate a probability distribution of the vocabulary. Finally, a beam search is used on the probability distribution of the vocabulary to generate the final candidate as the prediction result.
[0080] The training set, validation set, and test set each have 5936 (70%), 839 (10%), and 1706 (20%) samples, respectively, with a ratio of 7:1:2.
[0081] Experimental results and analysis:
[0082] In order to evaluate the performance of the model effect of the present invention, the model of the present invention is compared with the methods of the other two models:
[0083] ① VRepair uses a basic encoder-decoder Transformer model for vulnerability repair. VRepair is first trained on a dataset with vulnerability repairs annotated, and then fine-tuned on the vulnerability dataset to generate vulnerability repairs.
[0084] ②CodeBERT is an encoder-only Transformer-based model that is pre-trained on a large code library called CodeSearchNet, which was developed by Microsoft Research. CodeBERT consists of a 12-layer Transformer encoder block and a 6-layer Transformer decoder block for generation tasks.
[0085] For each model, we use a unified dataset to evaluate the accuracy. Specifically, we use a beam search with a width of 50 to generate 50 fix candidates for each vulnerable function in the test dataset. Therefore, the perfect prediction percentage can be calculated as the total number of correct predictions divided by the total number of functions in the test dataset. The final comparison results are shown in Figure 2. Figure 3 shown.
[0086] The present invention uses ablation experiments to study the accuracy of each component for the present invention model. Specifically, there are four methods: ① pre-training + BPE + T5 (AutoRepair) ② pre-training + word-level segmentation + T5 ③ no pre-training + BPE + T5 ④ no pre-training + word-level segmentation + T5. The final results are as follows Figure 4 As shown in the figure, it can be seen that in the present invention, the pre-training component contributes 14% to the prediction accuracy. In AutoRepair, the BPE component contributes 9% to the perfect prediction percentage.
[0087] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.< / modend> < / modstart> < / endloc> < / startloc> < / pad>
Claims
1. A method for automatically repairing software vulnerabilities based on bidirectional pre-training, characterized in that: The following steps are included: Step 1: Crawl vulnerabilities and their corresponding patches as data sets; Step 2: Use data cleaning and BPE segmentation to preprocess the data in the dataset; Step 3: Use Word Embedding to vectorize the preprocessed data; Step 4: Input the vectorized data into the constructed vulnerability repair model and repair the vulnerability code.
2. The method for automatically repairing software vulnerabilities based on bidirectional pre-training according to claim 1, characterized in that: The specific steps of the step 2 of using the BPE method to segment the data set are: first, the vulnerability data set is divided into single characters, and then the pair of characters with the highest frequency is replaced with another character in turn until the number of cycles is completed.
3. A method for automatically repairing software vulnerabilities based on bidirectional pre-training according to claim 1 or 2, characterized in that: In step 4, the vulnerability repair model is constructed based on the T5 model, including the following steps: 4.
1. Build a vulnerability repair model based on the CodeT5 pre-trained language model; 4.
2. We use stacked encoder blocks, combined with a multi-head self-attention mechanism and a feedforward neural network for relative position encoding, to capture the relative position information and semantic relationship of the input sequence through residual connections and layer normalization. 4.
3. By stacking twelve layers of decoder blocks, combining the mask mechanism to limit the prediction range, and using the cross-entropy loss to optimize the model, it only focuses on the context tags when generating vulnerability fix code, and monitors the training process through the validation set to select the best model; 4.
4. Use a 12-layer Transformer encoder and decoder, combined with a linearly scheduled AdamW optimizer and a Softmax layer, to convert the generated vector into a probability distribution, and fine-tune the learning rate and loss function; 4.
5. Using the beam search algorithm, by setting the beam width parameter β, the first β vulnerability repair candidates with the highest conditional probability are selected at each time step, and the search is terminated when the EOS mark is generated.
4. The method for automatically repairing software vulnerabilities based on bidirectional pre-training according to claim 3, characterized in that: As described in 4.2, each encoder block consists of two subcomponents, a multi-head self-attention layer with relative position encoding, followed by a feed-forward neural network.
5. The method for automatically repairing software vulnerabilities based on bidirectional pre-training according to claim 4, characterized in that: In 4.2, the multi-head self-attention mechanism is summarized as: MultiHead(Q,K,V)=Concat(Head1,...,Head h )W O Where: Q, K, V represent the input query, key and value matrices, W i represents the trainable weight matrix for linear transformation, Concat represents the concatenation of the outputs of multiple heads, Head represents the number of attention heads, and W O Used for linear projection to the expected dimension after concatenation.
6. A method for automatically repairing software vulnerabilities based on bidirectional pre-training according to claim 4 or 5, characterized in that: In 4.3, the cross entropy loss (H(p,q) = -∑ x∈X p(x)logq(x)) to update the model and optimize the difference between each position in the predicted sequence and the true sequence, where X is the set of classes, p is the true probability distribution, and q is the predicted probability distribution.