Vulnerability repairing method based on dynamic denoising and hybrid enhancement technology

Through dynamic denoising and hybrid enhancement technology, the vulnerability repair model is optimized, and the problems of data noise impact and noise label resistance in the existing technology are solved, achieving higher quality vulnerability repair and software development efficiency.

CN120105425APending Publication Date: 2025-06-06DALIAN MARITIME UNIVERSITY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510043711.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing automatic vulnerability repair methods cannot accurately learn the repair model due to the noise in the data set, and lack an effective resistance mechanism to noise labels.

Method used

Using a vulnerability repair method based on dynamic denoising and hybrid enhancement technology, the approximate distribution of clean data and noise data is derived through the expected maximization algorithm, the weight is calculated using the distribution-aware confidence function, mixed samples are generated and loss functions are reconstructed, and the vulnerability repair model is optimized to improve its resistance to noise.

Benefits of technology

It effectively reduces the impact of noise data during model training, improves the model's resistance to noise labels, improves the quality of the generated repair code, saves time and labor costs, and improves the efficiency and quality of software development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120105425A_ABST
    Figure CN120105425A_ABST
Patent Text Reader

Abstract

The invention provides a vulnerability repair method based on a dynamic denoising and hybrid enhancement technology, which comprises the following steps: S1, performing dynamic denoising and hybrid enhancement on an original vulnerability repair model to obtain an optimized vulnerability repair model; s11, deriving approximate distribution of clean data and noise data from the training loss of the original vulnerability repair model by using an expectation maximization algorithm; s12, calculating the weight of the clean data and the weight of the noise data by using a confidence coefficient function of distribution perception, and re-weighting the approximate distribution by using the weight of the clean data and the weight of the noise data; s13, generating a mixed sample by using the interpolation coefficient and reconstructing a loss function to obtain a total loss value; s14, repeatedly training the vulnerability repair model by using the value of the total loss until the value of the current total loss is not reduced in continuous n rounds or reaches the total number of rounds of training, and outputting the optimized vulnerability repair model; and S2, performing vulnerability repair by using the optimized vulnerability repair model. The software development efficiency and quality are improved, and software vulnerabilities are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of vulnerability repair technology, and in particular to a vulnerability repair method based on dynamic denoising and hybrid enhancement technology. Background Art

[0002] Vulnerability repair is the key to protecting software systems. Due to the popularity of computer software and the complexity of the development process, many software may have vulnerabilities. Attackers can use these unpatched vulnerabilities to carry out malicious activities, posing serious security risks to millions of users around the world. Vulnerability issues in software continue to increase, and the importance of vulnerability research and repair is increasing, requiring continuous investment of a lot of resources to meet the growing security challenges. In particular, according to statistics from 2022, it may take an average of 58 days to discover and fix a vulnerability. This shows that even if the number of vulnerabilities has decreased, their discovery and repair is still a time-consuming and complex process, and security analysts must spend a lot of energy tracking vulnerabilities and assessing the severity of vulnerabilities. With the development of technology and the increase in the complexity of software systems, it is expected that vulnerability repair will continue to be an important challenge, and effective vulnerability management strategies and security measures will be particularly important for reducing potential risks and maintaining system security.

[0003] Various deep learning-based automatic vulnerability repair methods have been proposed. Although these methods show great potential in automatic vulnerability repair, they still have some shortcomings in practical applications. First, the automatic vulnerability repair task relies on manually annotated datasets, which are usually noisy. Therefore, the vulnerability repair model cannot learn accurately, and this problem becomes more prominent as the dataset increases. (1) Label error Label errors occur during sample collection. Each sample is assigned a "positive" or "negative" label. However, given a large number of samples, label errors are inevitable and may lead to incorrect label assignment, that is, positive samples are labeled as negative and vice versa. (2) Representation noise When the code undergoes formal changes that do not affect the underlying semantics, representation noise is generated, resulting in misclassification, where positive and negative samples are incorrectly labeled. This noise has an adverse impact on the accuracy of learning-based models. Therefore, it is necessary to effectively identify such noise to improve the accuracy of inconsistency detection between annotations and source code. Second, existing automatic vulnerability repair methods lack an effective resistance mechanism to noisy labels during the repair process. Due to various factors, the model may mistakenly identify blank lines, comment lines, and code style changes as vulnerability repair code. Incorrect identification will hinder the vulnerability repair model from learning effective knowledge during the training process, causing the model to overfit to noisy labels. Summary of the invention

[0004] In view of this, the purpose of the present invention is to propose a vulnerability repair method based on dynamic denoising and hybrid enhancement technology to solve the technical problem that the existing automatic vulnerability repair method cannot accurately learn the repair model due to the noise in the data set.

[0005] The technical means adopted by the present invention are as follows:

[0006] A vulnerability repair method based on dynamic denoising and hybrid enhancement technology includes the following steps:

[0007] S1, dynamically denoising and hybrid enhancing the original vulnerability repair model to obtain an optimized vulnerability repair model;

[0008] S11. Use the expectation-maximization algorithm to derive the approximate distribution of clean and noisy data from the training loss of the original vulnerability repair model.

[0009] S12. Calculate the weights of clean data and noisy data using a distribution-aware confidence function, and re-weight the approximate distribution using the weights of the clean data and the noisy data to obtain a loss function;

[0010] S13, using the interpolation coefficient to generate mixed samples and reconstruct the loss function to obtain the value of the total loss;

[0011] S14, repeatedly training the vulnerability repair model using the total loss value until the current total loss value does not decrease for n consecutive rounds or reaches the total number of training rounds, and outputting the optimized vulnerability repair model;

[0012] S2. Use the optimized vulnerability repair model to repair the vulnerability: input the vulnerability code snippet to be repaired into the trained and optimized vulnerability repair model, and output the repair code.

[0013] Furthermore, in S1, the original vulnerability repair model is a transformer-based sequence-to-sequence generation model.

[0014] Furthermore, S11 specifically includes the following steps:

[0015] Create and update a dynamic distribution list to record the training loss during training;

[0016] The distribution-aware confidence function is used to reweight the training loss according to the distribution of the dynamic distribution list to obtain the approximate distribution of clean data and noisy data. The formula is as follows:

[0017] μ,v=EM(L)

[0018] Among them, μ is the approximate distribution of clean data, v is the approximate distribution of noisy data, and L is a dynamic distribution list that records the training loss of the approximate distribution.

[0019] Furthermore, calculating the weights of clean data and noise data in S12 specifically includes the following steps:

[0020] The training loss is redefined based on the distribution-aware confidence function of the approximate distribution of clean data and noisy data to obtain the loss function, which is as follows:

[0021]

[0022] Among them, C represents the weighted probability of vulnerability code data, P(l i |μ) represents the probability of given clean data, P(l i |v) represents the probability of given noise data, and α is a hyperparameter that compensates for the deviation of the expectation maximization algorithm;

[0023] For each data point, we first calculate the probability density of each data point in the clean data, and then weight the probability density in the noisy data. After obtaining a preliminary cleanliness measure, we perform probability normalization based on the probability of the data in the clean data distribution and the noisy data distribution, so that the confidence value is between 0 and 1, where 1 indicates that the data point is completely clean data and 0 indicates that the data point is completely noisy data. The confidence of different data is then calculated.

[0024] Furthermore, in S12, re-weighting the approximate distribution using the weights of the clean data and the noise data specifically includes the following steps:

[0025] Based on the weights of clean data and noisy data, the confidence level is used to readjust the training loss. The adjusted loss function is expressed as:

[0026]

[0027] Among them, τ (dsc) represents the reweighted training loss; represents the mean of the log-likelihood function of the label, p i (y i ) represents the model’s predicted probability for the label.

[0028] Further, S13 includes the following steps:

[0029] New virtual samples are generated by randomly extracting two code samples from the training data for interpolation; for a batch consisting of two random samples and their labels, new code samples are created by linear interpolation using the following formula:

[0030]

[0031]

[0032] Among them, λ is the interpolation parameter, λ~Βeta(α,α);

[0033] In order to take into account the accuracy of the original samples and the generated samples, a loss function is introduced, namely:

[0034]

[0035] The reconstructed loss function is a linear combination of the cross entropy loss and the supervision loss, and the formula is as follows:

[0036] L=L cross +λL sup

[0037] Among them, L cross is the cross entropy loss, which is used to measure the difference between the repair code predicted by the model and the actual repair code; L sup is the supervision loss, which is used to generate the vulnerability code snippets and their corresponding repair codes through the Mix Up technology; λ is the weight coefficient, which is used to balance the impact of the cross entropy loss and the supervision loss introduced by Mix Up on the model training.

[0038] The present invention also provides a storage medium, which includes a stored program, wherein when the program is run, any of the above-mentioned vulnerability repair methods based on dynamic denoising and hybrid enhancement technology is executed.

[0039] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes any one of the above-mentioned vulnerability repair methods based on dynamic denoising and hybrid enhancement technology through the computer program.

[0040] Compared with the prior art, the present invention has the following advantages:

[0041] The technical solution provided by the present invention performs automatic vulnerability repair based on dynamic denoising and hybrid enhancement technology. The method is designed for the field of automatic repair to reduce the impact of noise data in the model training process and then generate repair code. In the dynamic denoising part, the dynamic distribution list and distribution-aware consistency function are used to make the clean data obtain a greater weight in the training process. In the hybrid enhancement part, random interpolation is performed to expand the training distribution and joint label correction is performed to improve the model's resistance to noisy labels. Finally, it is applied to generate repair code. The main advantages are: (1) saving countless time and labor costs; (2) it is of great significance to improve code vulnerability repair; (3) improving software development efficiency and quality and reducing software vulnerabilities. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0043] Figure 1 The figure is a flow chart of the method of the present invention. DETAILED DESCRIPTION

[0044] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0045] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0046] The present invention uses vulnerable code collected from various sources to evaluate its effectiveness. The dataset includes vulnerability corpora written in C / C++ in Big-Vul and cveix. The Big-Vul dataset was obtained by crawling the CommonVulnerability and Exposure (CVE) database, which aggregates data from 348 open source projects on GitHub. From 2002 to 2019, Big-Vul contains a total of 3,754 code vulnerabilities. On the other hand, the cveix dataset is constructed in a similar way to Big-Vul, containing 5,365 vulnerabilities collected from 1,754 projects from 1999 to 2021. Using the two datasets preprocessed by Chen et al., after removing empty and duplicate samples, 5,417 samples spanning 2,095 different vulnerabilities were obtained.

[0047] like Figure 1 As shown, the present invention provides a vulnerability repair method based on dynamic denoising and hybrid enhancement technology, comprising the following steps:

[0048] S1, dynamically denoising and hybrid enhancing the original vulnerability repair model to obtain an optimized vulnerability repair model;

[0049] S11. Use the expectation-maximization algorithm to derive the approximate distribution of clean and noisy data from the training loss of the original vulnerability repair model.

[0050] For the two categories of clean data and noisy data, the number of clusters is selected as 2. First, the distribution parameters of the clean data distribution and the noisy data distribution are initialized respectively. Then the expectation step and the maximization step are iteratively performed: the expectation step is mainly to calculate the probability that each data point belongs to different distributions, and the maximization step is to update the distribution parameters according to the latest probability. When the absolute value of the distribution mean update is less than the threshold, it is considered that convergence has been achieved and the iteration is stopped.

[0051] In order to solve the annotation errors and characterization noise in the vulnerability patch annotations during the training process of the model, a dynamic distribution list is established and updated to record the training loss during the training process. In order to assign smaller weights to noisy data and larger weights to clean data, the present invention re-weights the training loss according to these distributions through a distribution-aware confidence function.

[0052] μ,v=EM(L)

[0053] Here, μ and v represent the approximate distributions of clean and noisy data, respectively, and L represents a dynamic distribution list that records the training loss.

[0054] S12. Calculate the weights of clean data and noisy data using a distribution-aware confidence function, and re-weight the approximate distribution using the weights of the clean data and the noisy data to obtain a loss function;

[0055] In order to assign smaller weights to noisy data and larger weights to clean data, the present invention redefines the training loss based on the distribution-aware confidence function of these two distributions.

[0056]

[0057] Where C represents the weighted probability (confidence function) of vulnerability code data, P(l i |μ) represents the probability of given clean data, P(l i|v) represents the probability of given noise data, α is a hyperparameter that compensates for the deviation of the expectation maximization algorithm. The role of this factor is to adjust the relative importance of the two probabilities (clean data and noise data). For each data point, first calculate its probability density in the clean data, and then weight the probability density in the noise data. After obtaining a preliminary cleanliness measure (that is, whether the data is more likely to belong to clean data or noise data), the probability is normalized according to the probability of the data in the clean data distribution and the noise data distribution. This ensures that the confidence value is between 0 and 1, where 1 means that the data point completely belongs to clean data, and 0 means that the data point is completely noise data. Then calculate the confidence of different data.

[0058] Finally, based on the weights of clean data and noise data obtained by the above method, the present invention uses confidence to readjust the training loss, thereby more accurately capturing the correspondence between vulnerability code and repair code and improving the vulnerability repair capability of the model. The adjusted training loss can be expressed as:

[0059]

[0060] Among them, τ (dsc) represents the reweighted training loss. Represents the mean of the log-likelihood function of the label, where p i (y i ) represents the model’s predicted probability for the label.

[0061] This loss function is used to calculate the prediction error of the model's prediction of the repair code. It calculates the loss by taking a weighted average of the log-likelihood of each sample. The weight is given by C(l i ) is determined, the higher the confidence, the greater the impact of the code sample on the model, which in turn affects the overall loss. In this way, not only the stability and efficiency of the model in processing noisy data are improved, but also the generalization ability of the model is enhanced, ensuring that the model can generate high-quality vulnerability patches, thereby improving the model's vulnerability repair efficiency.

[0062] S13, using the interpolation coefficient to generate mixed samples and reconstruct the loss function to obtain the value of the total loss;

[0063] In order to enhance the robustness and accuracy of the model in the task of generating repair code from vulnerability code datasets and reduce the excessive reliance on noise labels in the supervised learning stage, it is necessary to effectively expand the vulnerability repair code dataset to enhance the diversity and richness of data samples, so that the model can more easily capture the correspondence between codes.

[0064] Random interpolation: During the model training process, two code samples are randomly extracted from the training data for interpolation to generate new virtual samples. For a batch consisting of two random samples and their labels, the present invention uses the following formula to perform linear interpolation to create new code samples:

[0065]

[0066]

[0067] Among them, λ is the interpolation parameter, λ~Βeta(α,α). In order to take into account the accuracy of the original sample and the generated sample prediction, it is recommended to implement a robust sample loss correction method, namely:

[0068]

[0069] The complete loss function is a linear combination of the cross entropy loss and the supervision loss:

[0070] L=L cross +λL sup

[0071] Among them, L cross Cross-entropy loss, which measures the difference between the model’s predicted fix and the actual fix, L sup is the supervision loss, which takes into account the vulnerability code snippets generated by Mix Up technology and their corresponding repair codes, so that the model can accurately generate repair codes even in the presence of noisy labels. λ is a weight coefficient used to balance the impact of cross entropy loss and the supervision loss introduced by Mix Up on model training.

[0072] The hybrid enhancement technique provides a mechanism to combine clean and noisy samples and calculate a more representative loss to guide the training process. Even if two noisy samples are combined, the calculated loss is still useful because one of the noisy samples may contain the true label of the other sample. The mixture of samples and their labels can prevent overfitting on noisy samples, thereby improving the performance of the vulnerability code repair model.

[0073] S14, repeatedly training the vulnerability repair model using the total loss value until the difference between the current total loss value and the previous round total loss value is less than a given range, and outputting the optimized vulnerability repair model;

[0074] S141. Collect a code data set, which includes vulnerable codes and corresponding repair codes, and preprocess the data.

[0075] S142. Build a model, including an encoder and a decoder. The encoder is used to convert the input code token embedding and vulnerability mask into an embedding representation, where the vulnerability mask helps to emphasize the vulnerability embedding. The decoder is used to generate a repair code.

[0076] S143, input the preprocessed vulnerability code and vulnerability mask into the encoder to obtain an embedded representation of the input, and gradually generate a repair code corresponding to the vulnerability code through the decoder.

[0077] S144, calculate the cross entropy loss, and perform mixed data enhancement on the output and repair input of the model to obtain the mixed input, and then calculate the supervision loss. If the current number of training rounds is greater than 3, adjust the loss value according to the confidence list to achieve dynamic denoising.

[0078] S145. The model performs back propagation and gradient update.

[0079] S2. Use the optimized vulnerability repair model to repair the vulnerability: input the vulnerability code snippet to be repaired into the trained and optimized vulnerability repair model, and output the repair code.

[0080] Take the CWE-119 vulnerability dataset as an example. It is a code containing some potential vulnerabilities, including potential integer overflow errors. This code snippet is input into our optimized vulnerability repair model.

[0081] First, the encoder is used to encode the vulnerability code to obtain its embedded representation, and the mask mechanism is used to generate the corresponding mask for it. Finally, the decoder generates the corresponding repair code through vulnerability query.

[0082] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A vulnerability repair method based on dynamic denoising and hybrid enhancement technology, characterized in that: The steps include: S1, dynamically denoising and hybrid enhancing the original vulnerability repair model to obtain an optimized vulnerability repair model; S11. Use the expectation-maximization algorithm to derive the approximate distribution of clean and noisy data from the training loss of the original vulnerability repair model. S12. Calculate the weights of clean data and noisy data using a distribution-aware confidence function, and re-weight the approximate distribution using the weights of the clean data and the noisy data to obtain a loss function; S13, using the interpolation coefficient to generate mixed samples and reconstruct the loss function to obtain the value of the total loss; S14, repeatedly train the vulnerability repair model using the total loss value until the current total loss value does not decrease for n consecutive rounds or reaches the total number of training rounds, and output the optimized vulnerability repair model; S2. Use the optimized vulnerability repair model to repair the vulnerability: input the vulnerability code snippet to be repaired into the trained and optimized vulnerability repair model, and output the repair code.

2. The vulnerability repair method based on dynamic denoising and hybrid enhancement technology according to claim 1 is characterized in that: In S1, the original vulnerability repair model is a transformer-based encoder-decoder model.

3. The vulnerability repair method based on dynamic denoising and hybrid enhancement technology according to claim 1 is characterized in that: S11 specifically includes the following steps: Create and update a dynamic distribution list to record the training loss during training; The distribution-aware confidence function is used to reweight the training loss according to the distribution of the dynamic distribution list to obtain the approximate distribution of clean data and noisy data. The formula is as follows: μ,v=EM(L) Among them, μ is the approximate distribution of clean data, v is the approximate distribution of noisy data, and L is a dynamic distribution list that records the training loss of the approximate distribution.

4. The vulnerability repair method based on dynamic denoising and hybrid enhancement technology according to claim 1 is characterized in that: The calculation of the weights of clean data and noise data in S12 specifically includes the following steps: The training loss is redefined based on the distribution-aware confidence function of the approximate distribution of clean data and noisy data to obtain the loss function, which is as follows: Among them, C represents the weighted probability of vulnerability code data, P(l i |μ) represents the probability of given clean data, P(l i |v) represents the probability of given noise data, and α is a hyperparameter that compensates for the deviation of the expectation maximization algorithm; For each data point, we first calculate the probability density of each data point in the clean data, and then weight the probability density in the noisy data. After obtaining a preliminary cleanliness measure, we perform probability normalization based on the probability of the data in the clean data distribution and the noisy data distribution, so that the confidence value is between 0 and 1, where 1 indicates that the data point is completely clean data and 0 indicates that the data point is completely noisy data. The confidence of different data is then calculated.

5. The vulnerability repair method based on dynamic denoising and hybrid enhancement technology according to claim 1 is characterized in that: S12 uses the weights of clean data and noisy data to reweight the approximate distribution, which specifically includes the following steps: Based on the weights of clean data and noisy data, the confidence level is used to readjust the training loss. The adjusted loss function is expressed as: Among them, τ (dsc) represents the reweighted training loss; represents the mean of the log-likelihood function of the label, p i (y i ) represents the model’s predicted probability for the label.

6. The vulnerability repair method based on dynamic denoising and hybrid enhancement technology according to claim 1 is characterized in that: S13 includes the following steps: New virtual samples are generated by randomly extracting two code samples from the training data for interpolation; for a batch consisting of two random samples and their labels, new code samples are created by linear interpolation using the following formula: Among them, λ is the interpolation parameter, λ~Βeta(α,α); In order to take into account the accuracy of the original samples and the generated samples, a loss function is introduced, namely: The reconstructed loss function is a linear combination of the cross entropy loss and the supervision loss, and the formula is as follows: L=L cross +λL sup Among them, L cross is the cross entropy loss, which is used to measure the difference between the repair code predicted by the model and the actual repair code; L sup is the supervision loss, which is used to generate the vulnerability code snippets and their corresponding repair codes through the Mix Up technology; λ is the weight coefficient, which is used to balance the impact of the cross entropy loss and the supervision loss introduced by Mix Up on the model training.

7. A storage medium, characterized in that: The storage medium includes a stored program, wherein when the program is run, the vulnerability repair method based on dynamic denoising and hybrid enhancement technology described in any one of claims 1 to 6 is executed.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: The processor executes the vulnerability repair method based on dynamic denoising and hybrid enhancement technology as described in any one of claims 1 to 6 through the operation of the computer program.

Citation Information

Cited By

  • Semi-supervised submission level vulnerability classification method based on reinforcement learning enhancement

    CN122388765A

  • A semi-supervised submission-level vulnerability classification method based on reinforcement learning enhancement

    CN122388765B