Reinforcement learning training method for improving code structure perception capability of vulnerability repair model

By combining the comprehensive reward function of CodeBLEU, T5Score and Tree Edit Distance (TED), and designing the Actor and Critic network, the problem of insufficient code structure perception of the vulnerability repair model is solved, and the grammatical structure accuracy and semantic matching of the generated patches are significantly improved.

CN120597981APending Publication Date: 2025-09-05LIAONING UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510657069.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing vulnerability repair models are not capable of understanding code syntax and semantic structure, resulting in the generated patches not fully complying with the grammatical rules of the programming language or failing to accurately fix actual problems in the code.

Method used

A multi-angle structural similarity reward function is adopted, and a comprehensive reward function is constructed by combining CodeBLEU, T5Score and Tree Edit Distance (TED). By integrating reinforcement learning with pre-trained models, the Actor and Critic network is designed. The PPO and Advantage Actor-Critic framework are used for policy optimization. The KL divergence penalty term is combined to constrain the model output distribution, thereby improving the grammatical structure accuracy and semantic matching of the generated patches.

Benefits of technology

The grammatical structure accuracy and semantic matching of patches generated by the vulnerability repair model have been significantly improved, ensuring the functional correctness and code standardization of the generated code.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120597981A_ABST
    Figure CN120597981A_ABST
Patent Text Reader

Abstract

The invention relates to a reinforcement learning training method for improving the code structure perception capability of a vulnerability repair model, and belongs to the field of software security and deep reinforcement learning, and the method comprises the following steps: 1, designing a multi-angle structure similarity award function; step 2, designing a Critic model; and step 3, designing a reinforcement learning process. Through the method, the code structure sensing capability of the repair model is enhanced, so that the patch closer to a real label is generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention discloses a reinforcement learning training method for improving the code structure perception ability of a vulnerability repair model, belonging to the field of software security and deep reinforcement learning. Background Art

[0002] Traditional, manual approaches to addressing software vulnerabilities face numerous challenges. Manual patching is time-consuming and error-prone. Furthermore, manual intervention is not only difficult but can also introduce new vulnerabilities. With the advancement of deep learning and the proliferation of open-source models, researchers have begun applying deep learning methods to automated code vulnerability remediation. The most typical approach treats code as natural language and approaches the vulnerability remediation task as a sequence-to-sequence task: inputting the vulnerable code into a "translation" model, which outputs a sequence of patches. Common neural machine translation (NMT) models typically parse source code into a sequence of tokens and apply a word embedding model to each token, mapping the sequence into a vector space. An encoder then learns a deep representation, and a decoder decodes the code to output the corresponding patch. While deep learning models have made progress in understanding token relationships in code, their application to automated program remediation still faces several challenges. The primary issue is that while these models can capture some correlations between code tokens, they often lack a deep understanding of structural information, such as the code's syntax and semantics. This means that while the models may recognize certain patterns in the code, the patches they generate may not fully conform to the grammatical rules of the programming language or accurately fix the actual problem in the code.

[0003] RLHF (Reinforcement Learning from Human Feedback) is a fine-tuning method that combines supervised learning and reinforcement learning. It aims to optimize model output through human preference data to make it more consistent with human values ​​and actual needs. Its core process is divided into three steps: first, candidate answers are generated based on pre-trained models (such as the GPT series), and human feedback data on the quality of the answers is collected through manual labeling or comparative ranking; then, these feedbacks are used to train the reward model to convert human subjective preferences into a quantifiable scoring function; finally, the original model is fine-tuned through reinforcement learning algorithms (such as PPO) to maximize the score of the reward model, while constraints such as KL divergence are used to prevent the model from deviating too much from the initial state. RLHF performs outstandingly in scenarios such as dialogue systems and content generation. For example, ChatGPT significantly improves the security, relevance, and logic of answers through this method. Its advantages lie in its ability to handle complex and ambiguous optimization objectives (such as a "natural conversational feel") and reduce reliance on large amounts of manually annotated data. However, it also faces challenges: human feedback can be subjective and inconsistent, potentially introducing implicit biases; the accuracy of the reward model directly impacts the effectiveness of fine-tuning, and over-optimization can lead to rigid model output or violate the original intent. Furthermore, RLHF requires multi-stage training and extensive computing resources, placing high demands on algorithm design and engineering implementation. However, it remains one of the core technologies currently used to align AI behavior with human intent. Summary of the Invention

[0004] The purpose of the present invention is to provide a reinforcement learning training method for improving the code structure perception ability of the vulnerability repair model. This method solves the technical problem that the existing vulnerability repair model has insufficient code structure perception ability and improves the structural similarity between the patches generated by the repair model and the real labels.

[0005] The technical solution adopted by the present invention is as follows: a reinforcement learning training method to improve the code structure perception ability of the vulnerability repair model:

[0006] Step 1: Design a multi-angle structural similarity reward function:

[0007] 1.1) Calculate the positive reward function R + ;

[0008] Positive reward function R + The calculation method is:

[0009] 1.1.1) CodeBLEU

[0010] CodeBLEU = α·BLEU + β·BLEU weight +γ·Match ast +δ·Matchdf (1)

[0011] Among them, α, β, γ, δ are weighted coefficients, BLEU is the standard BLEU, BLEU weight is weighted BLEU, Match ast is the AST subtree matching ratio, Match df The data flow matching ratio;

[0012] 1.1.2) T5Score

[0013] 1.1.2.1) Token representation: Assume that the synthetic repair code after generating the repair patch is represented as x = {x1, x2, ..., x k}, the reference code is Use SASTCodeT5 encoder to encode x and Generate context embedding x={x1,x2,…,x k}and

[0014] 1.1.2.2) Compute similarity: Compute the similarity between the context embeddings represented by the vectors, x i and The cosine similarity is By normalizing, the calculation is simplified to the inner product

[0015] 1.1.2.3) Calculate precision and recall: The complete T5Score uses greedy matching to match each token in x with The tokens in are matched to calculate the recall rate, and the Each token in is matched with a token in x to calculate the accuracy. Each token is matched to the most similar token in the other sentence. Then the BLEU in CodeBLEU is used. weight The weighted method w(x i ), that is, the keyword weight is set to 2.5 times the weight of other words, and the precision and recall rates are shown in the formulas:

[0016]

[0017] Among them, R T5 Recall rate, P T5 Indicates accuracy;

[0018] 1.1.2.4) Scaling: In order to scale the cosine similarity between -1 and 1 to between 0 and 1, calculate R in the previous step T5 、P T5 The average value of R T5 The average value is a, PT5 The mean value of is b, and after scaling it is:

[0019]

[0020] 1.1.2.5) Calculating T5Score

[0021]

[0022] 1.1.3) Tree Edit Distance TED

[0023] The dynamic programming algorithm is used to calculate the minimum number of node deletions, insertions, and replacements required to go from one tree to another, and TED is introduced into the abstract syntax tree AST evaluation.

[0024] In TED, we define trees and forests, edit operations and edit scripts, and tree edit distance. For a tree on the alphabet V = {a, b, c, d, e}, it is expressed as Editing operations include deletion, replacement, and insertion, which are expressed as formulas 7, 8, and 9 respectively:

[0025]

[0026] Editing Scripts Defined as a series of editing operations δ1,…,δ T , applied to trees The above expression is: in The operation is The final tree edit distance is:

[0027]

[0028] When the editing operation fails due to unreasonable location, the cost c(δ i ,x,y)=0; when the edit operation includes a keyword node, c(δ i ,x,y)=2; normal editing operation, c(δ i ,x,u)=1, finally, the AST tree edit distance is expressed as formula 11:

[0029]

[0030] in Expressed as the total cost of editing the script:

[0031]

[0032] Finally, the positive reward function R + =CodeBLEU+ε·T5Score+TED AsT,Since the score of T5Score is between 0-1, ε is used to compensate the proportion of T5Score in CodeBLEU, ε=50.

[0033] 1.2) Calculate the penalty term P KL ;

[0034] The penalty term is calculated as:

[0035] The penalty term P is used to penalize the situation when the Actor output distribution deviates too far from the Reference model. The penalty term is calculated using the KL divergence, which is an asymmetric measure of the difference between two probability distributions. The formula is as follows:

[0036]

[0037] Among them, B code is the vulnerable function, Rep is the generated repair patch, is the probability of each patch token generated by the Actor model, π Ref (Rep|B code ) Similarly. In the corresponding relationship between vulnerability repair and reinforcement learning, B code is the initial state. With the selection of strategy, the state changes to B code With the continuous combination of repair patches, the action space of the strategy is the vocabulary.

[0038] 1.3) Calculate the comprehensive reward function R φ , the comprehensive reward function formula is:

[0039] R φ =R + -P KL

[0040] Step 2: Create a critic model and use the MSE error as the loss function to update the model. The formula is:

[0041]

[0042] in, is the target state value of the Critic model, through the time difference TD error, Use the comprehensive reward function R at time t φ,t and the old state value v at time t+1 old (s t+1 ) is used for calculation, γ is the conversion factor, the value is 0.95, v new (s t ) is the new predicted state value at time t;

[0043] The loss function of the Critic model is:

[0044] The Critic model comes from the encoder part of the SASTCodeT5 model, and the input is the state s at time t t , by vulnerable code B code It is composed of the repair patch Rep and outputs the value v(s) of each token in the current state through the fully connected layer. For the Critic model, the MSE error is used as the loss function to update the model, as shown in Formula 4-15:

[0045]

[0046] in, is the target value of the Critic model. For the target value at time t, the time difference error is used to calculate the value v at time t+1 when collecting experience. old (s t+1 ) and the current reward value at time t to approximate the target value at time t, because at time t, the current true reward has been calculated. If the prediction is divided into two parts: the current prediction and the future prediction, when the true value of the current prediction is known, the current true value is used to replace the current prediction, reducing the error of the total prediction;

[0047]

[0048] Therefore, Equation 15 becomes:

[0049]

[0050] Step 3: Create an Actor model. The Actor model is a stacked model based on the transformer architecture. It uses the CodeT5 model proposed by Google. The formula is:

[0051]

[0052] Among them, φ represents the parameters of the Actor model, ∈ is a hyperparameter that controls the clipping range, usually taken as 0.15, and clip(a,b,c) means limiting a to the interval [b,c]; A t represents the advantage function at time t, min(·,·) is the minimum operation;

[0053] The loss function of the Actor model is:

[0054] Based on the policy gradient algorithm, the training goal of the Actor model is to maximize the expected value of the state. The state value is calculated by multiplying the probability distribution of the policy output at the current time t by the action state value at the current time:

[0055]

[0056] For the action state value, if calculated according to the above formula, no matter what action is taken, a positive value will be obtained. As a result, the Actor model is always updated in this direction. Therefore, a baseline standard is used as a measure for the action value function. Above this baseline, the value is positive; below this baseline, the value is negative. This ensures normal training of the model. The value function after adding the baseline is called the advantage function:

[0057] A t =Q(s t ,a t )-V(s t ) (19)

[0058] The loss function of the Actor model becomes:

[0059]

[0060] In PPO, after importance sampling, the on-policy strategy is changed to the off-policy strategy. Under on-policy, one trajectory is sampled each time. After the actor is updated, the parameters have changed and the sampled data cannot be used anymore, so new sampling is required. Under off-policy, another strategy is used to interact with the environment for sampling, and the data is used to update the actor's parameters. Thus, a batch of data is used to update the actor multiple times. The core of importance sampling is as follows:

[0061]

[0062] That is, the expectation of the function f(x) on the distribution of p(x) is transformed into The expectation on the distribution q(x) is transformed from Formula 20 to Formula 23 by combining Formula 22:

[0063]

[0064] The PPO-Clip method is used to further optimize the loss of Formula 23:

[0065]

[0066] in, ∈ is a hyperparameter that controls the clipping range, and clip(a,b,c) means limiting a to the interval [b,c].

[0067] Step 4: Exploitation Interact with the environment to gather experience and learn:

[0068] 4.1) Collection Process: The vulnerable code is patched by the actor in inference mode and merged with the source code through patch location. The reward value is then given by the comprehensive reward function in step 1. The vulnerable code and the patch are simply concatenated and input into the models in steps 2 and 3 and the reference model. The Actor model's output Actor logits, the Critic model's state value, and the Reference model's output Ref logits are obtained to form a complete trajectory experience.

[0069] 4.2) Learning process:

[0070] Using the collected trajectory experience, online reinforcement learning is performed. Through the back-propagation algorithm, 100 epochs are trained, and the batch size of each epoch is 1. After 100 epochs of training, the final Actor model obtained is the final repair model.

[0071] The beneficial effects created by the present invention are: by integrating reinforcement learning with pre-training models, the grammatical structure accuracy and semantic matching of patches generated by the vulnerability repair model are significantly improved. Its core advantages are reflected in: a comprehensive reward function constructed based on CodeBLEU, T5Score and tree edit distance (TED), which not only captures the semantic similarity of the high-dimensional word vector output by the pre-trained encoder through T5Score, but also innovatively introduces TED with adjustable cost weights to accurately quantify the structural differences between the generated AST and the reference AST, strengthening grammatical alignment from multiple dimensions; at the same time, the deviation of the Actor output distribution from the pre-trained Reference model is constrained by the KL divergence penalty term, combined with the Critic network constructed based on the vulnerability repair model encoder, to achieve stable and efficient strategy optimization under the PPO and Advantage Actor-Critic framework, and ultimately significantly improve the structural similarity and code standardization of the patch and the real repair solution on the basis of ensuring the functional correctness of the generated code. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] Figure 1 This is the training framework diagram of the present invention

[0073] Figure 2 Graph of the comprehensive reward function;

[0074] Figure 3 This is the structure diagram of the Critic model;

[0075] Figure 4 This is the flow chart of the reinforcement learning algorithm;

[0076] Figure 5 Collect experience process graphs for the interaction between Actor and Critic models and the environment. DETAILED DESCRIPTION

[0077] A reinforcement learning training method to improve the code structure perception ability of vulnerability repair models:

[0078] Step 1: Design a multi-angle structural similarity reward function. The comprehensive reward function is composed as follows: Figure 2 As shown:

[0079] 1.1) Calculate the positive reward function R + ;

[0080] In software engineering, evaluating the quality of patch repairs requires adherence to a rigorous set of criteria. The most important criterion is that they completely address vulnerabilities without introducing new ones. Automatically evaluating patch quality is a complex and laborious task, as it involves merging the patch into the appropriate source code area and reproducing the attack for a specific patch in a specific project and hardware environment. While merging the patch into the project can be automated, reproducing the attack in a specific project and hardware environment is extremely resource-intensive or even impossible. The most straightforward approach is manual verification by the project's programmers, which is impractical for large-scale deep learning. Deep learning evaluation methods can assess the quality of generated patches by examining factors such as whether the string matches the true label or whether the merged code structure is highly similar to the correct code. Therefore, when designing the reward function, this study focused on using a variety of NLP task metrics to develop a reward function that accurately evaluates patch quality across multiple dimensions.

[0081] In the evaluation of NLP models, a variety of metrics are widely used to measure model performance, including N-gram-based metrics, edit distance-based metrics, and embedding-based metrics. These metrics differ in their definitions, focus, and calculation methods, providing multi-faceted considerations for NLP model performance evaluation. Among N-gram-based metrics,

[0082] CodeBLEU is designed specifically for code. While incorporating BLEU, it also considers matching syntactic and semantic structures. Therefore, CodeBLEU can be used in the reward function. Among embedding-based metrics, MEANT2.0 and YISI-1 both use relatively simple similarity calculations. BertScore uses contextual embeddings for calculations, employing greedy matching and weighted importance to improve its effectiveness. However, the present invention makes appropriate modifications to BertScore to suit the tasks of this research.

[0083] Edit distance-based metrics quantify similarity by calculating the number of edit operations required to go from a candidate to a reference. The present invention does not use these traditional string edit distances but instead introduces a tree edit distance. The tree edit distance describes the minimum number of node deletions, insertions, and replacements required to convert from one tree to another. This is somewhat different from the metric used in CodeBLEU to measure the number of AST subtree matches. The present invention believes that for two ASTs, the smaller the edit distance from one to the other, the higher the similarity between the two ASTs. Therefore, the present invention incorporates the tree edit distance into the reward function calculation. In summary, the present invention takes into account CodeBLEU, T5Score, and tree edit distance when designing the positive reward function.

[0084] 1.1.1) CodeBLEU

[0085] CodeBLEU = α·BLEU + β·BLEU weight +γ·Match ast +δ·Match df (1)

[0086] Among them, α, β, γ, δ are weighted coefficients, BLEU is the standard BLEU, BLEU weight is weighted BLEU, Match ast is the AST subtree matching ratio, Match df The data flow matching ratio;

[0087] 1.1.2) T5Score: BertScore improves two common defects of N-gram-based indicators: one is that N-gram-based indicators generally cannot match sentence semantics; the other is that N-gram models cannot capture long-range dependencies. BertScore uses tag embeddings that contain contextual semantics to calculate similarity, because context embeddings are trained to effectively capture long-range dependencies and sequences. However, the present invention only uses the main method of BertScore and does not choose to use the Bert model. Instead, it uses the pre-trained SASTCodeT5 encoder model. Compared to the Bert model, it is trained on a large-scale code corpus and is retrained on error and vulnerability datasets. Its word embedding representation ability for code is much higher than that of the Bert model. The present invention calls this indicator T5Score, and its calculation method is divided into the following steps:

[0088] 1.1.2.1) Token representation: Assume that the synthetic repair code after generating the repair patch is represented as x = {x1, x2, ..., x k}, the reference code is Use SASTCodeT5 encoder to encode x and Generate context embedding x={x1,x2,…,x k}and

[0089] 1.1.2.2) Compute similarity: Compute the similarity between the context embeddings represented by the vectors, x i and The cosine similarity is By normalizing, the calculation is simplified to the inner product

[0090] 1.1.2.3) Calculate precision and recall: The complete T5Score uses greedy matching to match each token in x with The tokens in are matched to calculate the recall rate, and the Each token in is matched with a token in x to calculate the accuracy. Each token is matched to the most similar token in the other sentence. Then the BLEU in CodeBLEU is used. weight The weighted method w(x i ), that is, the keyword weight is set to 2.5 times the weight of other words, and the precision and recall rates are shown in the formulas:

[0091]

[0092] Among them, R T5 Recall rate, P T5 Indicates accuracy;

[0093] 1.1.2.4) Scaling: In order to scale the cosine similarity between -1 and 1 to between 0 and 1, calculate R in the previous step T5 、P T5 The average value of R T5 The average value is a, P t5 The mean value of is b, and after scaling it is:

[0094]

[0095] 1.1.2.5) Calculating T5Score

[0096]

[0097] 1.1.3) Tree Edit Distance (TED):

[0098] The Tree Edit Distance (TED) calculation uses a dynamic programming algorithm to calculate the minimum number of node deletions, insertions, and replacements required to move from one tree to another. This paper introduces TED into AST evaluation and customizes the distance calculation method and cost in the algorithm.

[0099] In TED, we define trees and forests, edit operations and edit scripts, and tree edit distance. For a tree on the alphabet V = {a, b, c, d, e}, it is expressed as Editing operations include deletion, replacement, and insertion, which are expressed as formulas 7, 8, and 9 respectively:

[0100]

[0101] Editing Scripts Defined as a series of editing operations δ1,…,δ T , applied to trees The above expression is: in The operation is The final tree edit distance is:

[0102]

[0103] When the editing operation fails due to unreasonable location, the cost c(δ i ,x,y)=0; when the edit operation includes a keyword node, c(δ i ,x,y)=2; normal editing operation, c(δ i ,x,y)=1, finally, the AST tree edit distance is expressed as formula 11:

[0104]

[0105] in Expressed as the total cost of editing the script:

[0106]

[0107] Finally, the positive reward function R + =CodeBLEU+ε·T5Score+TED AsT ,Since the score of T5Score is between 0-1, ε is used to compensate the proportion of T5Score in CodeBLEU, ε=50.

[0108] 1.2) Calculate the penalty term P KL ;

[0109] The penalty term is calculated as:

[0110] The penalty term P is used to penalize the situation when the Actor output distribution deviates too far from the Reference model. The penalty term is calculated using the KL divergence, which is an asymmetric measure of the difference between two probability distributions. The formula is as follows:

[0111]

[0112] Among them, B code is the vulnerable function, Rep is the generated repair patch, is the probability of each patch token generated by the Actor model, π Ref (Rep|B code ) Similarly. In the corresponding relationship between vulnerability repair and reinforcement learning, B code is the initial state. With the selection of strategy, the state changes to B code With the continuous combination of repair patches, the action space of the strategy is the vocabulary.

[0113] 1.3) Calculate the comprehensive reward function R φ , the comprehensive reward function formula is:

[0114] R φ =R + -P KL

[0115] Step 2: Create a Critic model. The Critic model is as follows: Figure 3 As shown, the MSE error is used as the loss function to update the model. The formula is:

[0116]

[0117] in, is the target state value of the Critic model, through the temporal difference (TD) error, Use the comprehensive reward function R at time t φ,t and the old state value v at time t+1 old (s t+1 ) is used for calculation, γ is the conversion factor, the value is 0.95, v new (s t ) is the new predicted state value at time t;

[0118] The loss function of the Critic model is:

[0119] The Critic model comes from the encoder part of the SASTCodeT5 model, and the input is the state s at time t t , by vulnerable code B codeIt is composed of the repair patch Rep and outputs the value v(s) of each token in the current state through the fully connected layer. For the Critic model, the MSE error is used as the loss function to update the model, as shown in Formula 4-15:

[0120]

[0121] in, is the target value of the Critic model. For the target value at time t, the time difference error is used to calculate the value v at time t+1 when collecting experience. old (s t+1 ) and the current reward value at time t to approximate the target value at time t, because at time t, the current true reward has been calculated. If the prediction is divided into two parts: the current prediction and the future prediction, when the true value of the current prediction is known, the current true value is used to replace the current prediction, reducing the error of the total prediction;

[0122]

[0123] Therefore, Equation 15 becomes:

[0124]

[0125] Step 3: Create an Actor model. The Actor model is a stacked model based on the transformer architecture. It uses the CodeT5 model proposed by Google. The formula is:

[0126]

[0127] Among them, φ represents the parameters of the Actor model, ∈ is a hyperparameter that controls the clipping range, usually taken as 0.15, and clip(a,b,c) means limiting a to the interval [b,c]; A t represents the advantage function at time t, min(·,·) is the minimum operation;

[0128] The loss function of the Actor model is:

[0129] Based on the policy gradient algorithm, the training goal of the Actor model is to maximize the expected value of the state. The state value is calculated by multiplying the probability distribution of the policy output at the current time t by the action state value at the current time:

[0130]

[0131] For the action state value, if calculated according to the above formula, no matter what action is taken, a positive value will be obtained. As a result, the Actor model is always updated in this direction. Therefore, a baseline standard is used as a measure for the action value function. Above this baseline, the value is positive; below this baseline, the value is negative. This ensures normal training of the model. The value function after adding the baseline is called the advantage function:

[0132] A t =Q(s t ,a t )-V(s t ) (19)

[0133] The loss function of the Actor model becomes:

[0134]

[0135] In PPO, after importance sampling, the on-policy strategy is changed to the off-policy strategy. Under on-policy, one trajectory is sampled each time. After the actor is updated, the parameters have changed and the sampled data cannot be used anymore, so new sampling is required. Under off-policy, another strategy is used to interact with the environment for sampling, and the data is used to update the actor's parameters. Thus, a batch of data is used to update the actor multiple times. The core of importance sampling is as follows:

[0136]

[0137] That is, the expectation of the function f(x) on the distribution of p(x) is transformed into The expectation on the distribution q(x) is transformed from Formula 20 to Formula 23 by combining Formula 22:

[0138]

[0139] The PPO-Clip method is used to further optimize the loss of Formula 23:

[0140]

[0141] in, ∈ is a hyperparameter that controls the clipping range, and clip(a,b,c) means limiting a to the interval [b,c].

[0142] Step 4: Using Pi φold Interact with the environment to collect experience and learn. The process of Actor and Critic models interacting with the environment to collect experience is as follows: Figure 5 As shown:

[0143] 4.1) Collection Process: The vulnerable code is patched by the actor in inference mode and merged with the source code through patch location. The reward value is then given by the comprehensive reward function in step 1. The vulnerable code and the patch are simply concatenated and input into the models in steps 2 and 3 and the reference model. The Actor model's output Actor logits, the Critic model's state value, and the Reference model's output Ref logits are obtained to form a complete trajectory experience.

[0144] 4.2) Learning process:

[0145] Using the collected trajectory experience, online reinforcement learning is performed, and 100 epochs are trained through the back propagation algorithm. The batch size of each epoch is 1. After training 100 epochs, the final Actor model obtained is the final repair model. The specific method is as follows Figure 4 shown.

Claims

1. A reinforcement learning training method for improving the code structure perception ability of a vulnerability repair model, characterized by: Step 1: Design a multi-angle structural similarity reward function: 1.1) Calculate the positive reward function R + ; 1.2) Calculate the penalty term P KL ; 1.3) Calculate the comprehensive reward function R φ , the comprehensive reward function formula is: R φ =R + -P KL Step 2: Create a critic model and use the MSE error as the loss function to update the model. The formula is: in, is the target state value of the Critic model, through the time difference TD error, Use the comprehensive reward function R at time t φ,t and the old state value v at time t+1 old (s t+1 ) is used for calculation, γ is the conversion factor, the value is 0.95, v new (s t ) is the new predicted state value at time t; Step 3: Create an Actor model. The Actor model is a stacked model based on the transformer architecture. It uses the CodeT5 model proposed by Google. The formula is: Among them, φ represents the parameters of the Actor model, ∈ is a hyperparameter that controls the clipping range, usually 0.15, and clip(a,b,c) means limiting a to the interval [b,c]; S t represents the advantage function at time t, min(·,·) is the minimum operation; Step 4: Using Pi φold Interact with the environment to gather experience and learn: 4.1) Collection Process: The vulnerable code is patched by the actor in inference mode and merged with the source code through patch location. The reward value is then given by the comprehensive reward function in step 1. The vulnerable code and the patch are simply concatenated and input into the models in steps 2 and 3 and the reference model. The Actor model's output Actor logits, the Critic model's state value, and the Reference model's output Ref logits are obtained to form a complete trajectory experience. 4.2) Learning process: Using the collected trajectory experience, we conduct online reinforcement learning and train 100 epochs through the back propagation algorithm. The batch size of each epoch is 1. After training for 100 epochs, the final Actor model obtained is the final repair model.

2. A reinforcement learning training method for improving the code structure perception ability of a vulnerability repair model according to claim 1, characterized in that: The positive reward function r in 1.1) + The calculation method is: 1.1.1) CodeBLEU CodeBLUE=α·BLUE+β·BLUE weight +γ·Match ast +δ·Match df (1) Among them, α, β, γ, δ are weighted coefficients, BLEU is the standard BLEU, BLEU weig is weighted BLEU, Match ast is the AST subtree matching ratio, Match df The data flow matching ratio; 1.1.2) T5Score 1.1.2.1) Token representation: Assume that the synthetic repair code after generating the repair patch is represented as x = {x1, x2, ..., x k }, the reference code is Use SASTCodeT5 encoder to encode x and Generate context embedding x={x1,x2,…,x k }and 1.1.2.2) Compute similarity: Compute the similarity between the context embeddings represented by the vectors, x i and The cosine similarity is By normalizing, the calculation is simplified to the inner product 1.1.2.3) Calculate precision and recall: The complete T5Score uses greedy matching to match each token in x with The tokens in are matched to calculate the recall rate, and the Each token in is matched with a token in x to calculate the accuracy. Each token is matched to the most similar token in the other sentence. Then the BLEU in CodeBLEU is used. weight The weighted method w(x i ), that is, the keyword weight is set to 2.5 times the weight of other words, and the precision and recall rates are shown in the formulas: Among them, R T5 Recall rate, P T5 Indicates accuracy; 1.1.2.4) Scaling: In order to scale the cosine similarity between -1 and 1 to between 0 and 1, calculate R in the previous step T5 、P T5 The average value of R T5 The average value is a, P T5 The mean value of is b, and after scaling it is: 1.1.2.5) Calculating T5Score 1.1.3) Tree Edit Distance TED The dynamic programming algorithm is used to calculate the minimum number of node deletions, insertions, and replacements required to go from one tree to another, and TED is introduced into the abstract syntax tree AST evaluation. In TED, we define trees and forests, edit operations and edit scripts, and tree edit distance. For a tree on the alphabet V = {a, b, c, d, e}, it is expressed as Editing operations include deletion, replacement, and insertion, which are expressed as formulas 7, 8, and 9 respectively: Editing Scripts Defined as a series of editing operations δ1,…,δ T , applied to trees The above expression is: in The operation is The final tree edit distance is: When the editing operation fails due to unreasonable location, the cost c(δ i ,x,y)=0; when the edit operation includes a keyword node, c(δ i ,x,y)=2; normal editing operation, c(δ i ,x,y)=1, finally, the AST tree edit distance is expressed as formula 11: in Expressed as the total cost of editing the script: Finally, the positive reward function R + =CodeBLEU+ε·T5Score+TED AST ,Since the score of T5Score is between 0-1, ε is used to compensate the proportion of T5Score in CodeBLEU, ε=50.

3. The reinforcement learning training method for improving the code structure perception ability of the vulnerability repair model according to claim 1 is characterized in that: The calculation method of the penalty term in 1.2) is: The penalty term P is used to penalize the situation when the Actor output distribution deviates too far from the Reference model. The penalty term is calculated using the KL divergence, which is an asymmetric measure of the difference between two probability distributions. The formula is as follows: Among them, B code is the vulnerable function, Rep is the generated repair patch, is the probability of each patch token generated by the Actor model, π Ref (Rep|B code ) Similarly. In the corresponding relationship between vulnerability repair and reinforcement learning, B code is the initial state. With the selection of strategy, the state changes to B code With the continuous combination of repair patches, the action space of the strategy is the vocabulary.

4. The reinforcement learning training method for improving the code structure perception ability of the vulnerability repair model according to claim 1 is characterized in that: In step 2), the loss function of the Critic model is: The Critic model comes from the encoder part of the SASTCodeT5 model, and the input is the state s at time t t , by vulnerable code B code It is composed of the repair patch Rep and outputs the value v(s) of each token in the current state through the fully connected layer. For the Critic model, the MSE error is used as the loss function to update the model, as shown in Formula 4-15: in, is the target value of the Critic model. For the target value at time t, the time difference error is used to calculate the value v at time t+1 when collecting experience. old (s t+1 ) and the current reward value at time t to approximate the target value at time t, because at time t, the current true reward has been calculated. If the prediction is divided into two parts: the current prediction and the future prediction, when the true value of the current prediction is known, the current true value is used to replace the current prediction, reducing the error of the total prediction; Therefore, Equation 15 becomes:

5. The reinforcement learning training method for improving the code structure perception ability of the vulnerability repair model according to claim 1 is characterized in that: In step 3), the loss function of the Actor model is: Based on the policy gradient algorithm, the training goal of the Actor model is to maximize the expected value of the state. The state value is calculated by multiplying the probability distribution of the policy output at the current time t by the action state value at the current time: For the action state value, if calculated according to the above formula, no matter what action is taken, a positive value will be obtained. As a result, the Actor model is always updated in this direction. Therefore, a baseline standard is used as a measure for the action value function. Above this baseline, the value is positive; below this baseline, the value is negative. This ensures normal training of the model. The value function after adding the baseline is called the advantage function: A t =Q(s t ,a t )-V(s t ) (19) The loss function of the Actor model becomes: In PPO, after importance sampling, the on-policy strategy is changed to the off-policy strategy. Under on-policy, one trajectory is sampled each time. After the actor is updated, the parameters have changed and the sampled data cannot be used anymore, so new sampling is required. Under off-policy, another strategy is used to interact with the environment for sampling, and the data is used to update the actor's parameters. Thus, a batch of data is used to update the actor multiple times. The core of importance sampling is as follows: That is, the expectation of the function f(x) on the distribution of p(x) is transformed into The expectation on the distribution q(x) is transformed from Formula 20 to Formula 23 by combining Formula 22: The PPO-Clip method is used to further optimize the loss of Formula 23: in, ∈ is a hyperparameter that controls the clipping range, and clip(a,b,c) means limiting a to the interval [b,c].