Method for carrying out code bug repair detection by utilizing noise transfer
By merging the data set, extracting code changes characteristics using CodeBERT, and generating outputs through the Transformer model, setting confidence thresholds and correcting data using the transfer probability matrix, the problem of high noise in the existing vulnerability repair detection methods is solved, and the accuracy and efficiency of detection are improved.
Patent Information
- Application Number
- CN202510070426.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-30
AI Technical Summary
Existing vulnerability repair detection methods are difficult to distinguish vulnerability repair from other unrelated code changes when handling code changes, resulting in high noise and inaccurate detection results.
By merging existing vulnerability datasets and open source project datasets, extract code changes features using CodeBERT, and generate outputs through a pre-trained model based on Transformer. Set the confidence threshold to divide the code change characteristics into trustworthy and untrusted data, and use the transfer probability matrix to correct the prediction results of untrusted data.
Effectively reduce the interference of noise data on the model, improve the accuracy and efficiency of vulnerability repair detection, maintain good performance in complex scenarios, and reduce manpower and material consumption.
Smart Images

Figure CN120068085A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of code vulnerability repair, and more particularly, to a method for detecting code vulnerability repair by using noise transfer. Background Art
[0002] According to the requirements of the coordinated vulnerability disclosure model, vulnerabilities in software using third-party libraries should be "silently" repaired, that is, no information that may reveal the nature of the vulnerability should be disclosed. Past techniques classified commits by leveraging resources such as commit messages or issue reports. However, when processed according to the coordinated vulnerability disclosure method, these resources should not involve any security-related information during the repair process to avoid exposing its existence before the vulnerability is publicly disclosed. Therefore, it may not be practical to detect vulnerabilities and their repairs based on natural language resources such as commit messages or issue reports.
[0003] In this context, detecting code vulnerability repair using code changes becomes an inevitable choice. VulFixMiner was the first model to propose using deep learning techniques to extract code changes for code vulnerability repair detection. It uses CodeBERT to automatically represent code changes and extract features to detect vulnerability repair commits in mixed commits. Subsequently, CoLeFunDa proposed a new functional change data augmentation method, which uses supervised and unsupervised strategies to expand code changes at the function level and trains the function change encoder FCBERT through contrastive learning. Subsequently, FCBERT was further fine-tuned for three downstream tasks, including silent repair identification, corresponding Common Weakness Enumeration (CWE) classification, and exploitability rating.
[0004] There are several main problems in existing vulnerability repair detection methods. First, when dealing with code changes, existing methods are difficult to effectively distinguish vulnerability repairs from other unrelated code changes, such as code refactoring or the addition of new functions. Due to the existence of this mixed information, a large amount of noise is introduced into the model, reducing the detection accuracy. When a commit contains multiple types of changes simultaneously, existing models often cannot handle complex scenarios and are prone to confusing vulnerability repairs with other modifications. In addition, the scarcity of vulnerability repair code is also a key issue. Statistical data shows that vulnerability repairs account for only 0.35% of the entire commits. In this data-scarce situation, it is difficult for existing technologies to accurately locate true vulnerability repair commits from a large number of noisy commits, and the model is easily interfered by noise, resulting in unsatisfactory detection effects. Finally, existing methods usually rely on multiple information sources, such as commit messages and issue reports, which makes it consume a large amount of computing resources to process noisy commit data, resulting in high detection costs. These factors combined limit the efficiency and practical application effect of vulnerability repair detection. Summary of the Invention
[0005] In view of this, the object of the present invention is to propose a method for detecting code vulnerability repair by using noise transfer, so as to solve the technical problem that the existing vulnerability repair detection method has a large amount of noise, resulting in inaccurate detection results.
[0006] The technical means adopted by the present invention are as follows:
[0007] A method for detecting code vulnerability repair by using noise transfer, comprising the following steps:
[0008] S1. Merge the existing vulnerability data set and the data set collected from real open source projects to form a total data set;
[0009] S2. Use CodeBERT to extract and embed the code changes in the total data set, and obtain code change features through a feature extractor;
[0010] S3. Input the code change features into a pre-trained model based on Transformer, generate an output through the pre-trained model, and convert the output into a probability distribution for each category; the probability distribution indicates the prediction confidence of the model for each category. Set a confidence threshold, and divide the code change features into high-confidence credible data and low-confidence uncredible data according to the confidence threshold;
[0011] S4. Input the uncredible data into a classifier. After obtaining the output result of the uncredible data, jointly calculate the credible data and the uncredible data to obtain the transfer probability of the uncredible data;
[0012] S5. Construct a transfer probability matrix according to the obtained transfer probability, and perform an operation on the transfer probability matrix and the output result of the uncredible data to obtain the classification result of the uncredible data;
[0013] S6. Input the credible data into a classifier to obtain the classification result of the credible data.
[0014] Further, S2 specifically includes the following steps:
[0015] S21. Extract file-level code changes; extract the information of file-level changes into a separate code change document; for each document, further extract the deleted and added code lines, and split the deleted and added code lines into rem-code and add-code segments respectively;
[0016] S22. Process the deleted code and the added code lines; for each segment, first mark the code line as a series of tokens, and then further split the camel case style and snake case style tokens by using codeprep;
[0017] S23. Input token construction: To construct an input that meets the requirements of CodeBERT, the pre-trained tokenizer RobertaTokenizer of CodeBERT is used, and code snippets are concatenated with the delimiters [CLS], [SEP], and [EOS].
[0018] Further, S5 specifically includes the following steps:
[0019] For untrusted data, obtain the output through a classification model, and call the classification component Let and y be the set of possible labels, x be the feature vector of each data in the silver dataset, and the transition matrix C represents the transition probability from the true label y to the predicted label of Use the following formula to obtain the transition probability
[0020]
[0021] where the integral represents the weighted average of the values of x, and p ( x|y ) represents the probability distribution of x when the true label is y;
[0022] Use to replace to obtain the calculation formula for the transition probability, approximately calculate the integral of the transition probability, and obtain the transition probability formula as follows:
[0023]
[0024] where A i represents the set of all samples with the label y = i (i ∈ {0, 1}), that is
[0025] Further, S5 specifically includes the following steps:
[0026] Fill the obtained transition probability and the true label into the initialized 2x2 transition probability matrix:
[0027]
[0028] represents the probability that the predicted label is j when the true label is i.
[0029] Pass the untrusted data through the classifier to obtain the output S, and perform Then, it is trained with the Focal loss, a cross-entropy loss function with an adjustment factor, and finally the classification results of untrusted data are obtained.
[0030] The present invention also provides a storage medium, which includes a stored program. When the program runs, it executes any one of the above methods for detecting code vulnerability repair using noise transfer.
[0031] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor runs the computer program to execute any one of the above methods for detecting code vulnerability repair using noise transfer.
[0032] Compared with the prior art, the present invention has the following advantages:
[0033] The technical solution provided by the present invention divides data by calculating confidence, and corrects the prediction performance of untrusted data with trusted data, which can effectively reduce the interference of noisy data on the model. By reasonably using the present invention, the model can more accurately correct the wrong labels in the training process, thereby improving the overall performance. In addition, the present invention not only brings improvement in accuracy, but also enhances the performance of the model in dealing with complex scenarios. In a dataset containing a large amount of noise, this model can help the model avoid overfitting to wrong or noisy labels, ensuring that the model has better generalization performance when facing actual application data. Its main advantages are: (1) It can save a large amount of manpower and material resources in the field of vulnerability repair detection. (2) It can still obtain excellent detection results when the vulnerability repair submission information is very complex. (3) It can detect the security risks of third-party library vulnerabilities in application software in a timely manner and improve software security. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0035] Figure 1 It is a flowchart of the method of the present invention.
[0036] Figure 2 It is an example diagram of the input token constructed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0037] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0038] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0039] As Figure 1 shown, the present invention provides a method for detecting code vulnerability repair using noise transfer, including the following steps:
[0040] Merge the existing vulnerability datasets and collect datasets from real open-source projects to form a new dataset.
[0041] Extract embeddings for code changes using CodeBERT and obtain code change features through a feature extractor.
[0042] Input the code change features into a pre-trained model based on Transformer, generate an output through the pre-trained model, and use F.softmax(outs, dim = 1) to convert the output into a probability distribution for each category. These probabilities indicate the prediction confidence of the model for each category. Set the confidence threshold to 0.7, that is, if the prediction probability is greater than 70%, it is considered as trustworthy data, and if it is less than 70%, it is considered as untrustworthy data. In this way, the code change features can be classified into trustworthy data with high confidence and untrustworthy data with low confidence according to the confidence level.
[0043] Input untrusted data into a classifier, which consists of a linear layer that maps trusted and untrusted data to the same dimension, a ReLU activation function, a dropout layer, and a fully connected layer. After training for 60 epochs, convert the feature dimension to NUMBER_OF_LABELS through the fully connected layer (since the number of labels is 2, that is, obtain the prediction result of the untrusted data). After obtaining the output result of the untrusted data, jointly calculate the trusted and untrusted data to obtain the transfer probability of the untrusted data;
[0044] Construct a transition probability matrix based on the obtained transition probability, and perform operations on the transition probability matrix and the output result of the untrusted data to obtain the classification result of the untrusted data.
[0045] Input the trusted data into the classifier to obtain the classification result of the trusted data.
[0046] To evaluate the effectiveness of the model, collect data from two places and merge them to obtain a new dataset: (1) the existing vulnerability repair detection dataset (2) the dataset collected from real-world open-source projects.
[0047] For the existing vulnerability repair detection dataset, the VulFixminer dataset is used. The Java dataset contains 1,436 vulnerability-fixing commits and 839,682 non-vulnerability-fixing commits. At the same time, the Python dataset contains 885 vulnerability-fixing commits and 722,291 non-vulnerability-fixing commits. Then, VulFixMiner further filters the dataset by removing large commits that are unlikely to fix vulnerabilities. The removal led to 474,555 non-vulnerability-fixing commits and 1,353 vulnerability-fixing commits from 150 Java projects. For Python, the corresponding values are 357,696 non-vulnerability-fixing commits and 751 vulnerability-fixing commits from 106 projects. Finally, VulFixMiner enhances the dataset by marking more commits related to vulnerability repair. More specifically, the commit messages contain keywords related to vulnerabilities (i.e., "vuln", "CVE", and "NVD"). To ensure that the pattern is well-designed, VulFixMiner randomly selects a subset of the commits and manually verifies them. Therefore, in the Java dataset, they re-labeled 420 non-vulnerability-fixing commits in 123 projects. In the Python dataset, they re-labeled 501 non-vulnerability-fixing commits in 98 projects.
[0048] For the datasets of open-source projects, GitHub was used to collect commit links related to the total number of issues (471) and pull requests (508) from Java and Python projects. This resulted in 383 Java vulnerability fixes across 101 projects, corresponding to 186 CVEs. For Python, this resulted in 597 vulnerability fix commits across 141 projects, corresponding to 233 CVEs. Add this result to the list from the previous step and remove any duplicates.
[0049] In the data preprocessing stage, CodeBERT was selected as the pre-trained language model. First, code changes are treated as token sequences in a BoW. The code change preprocessor (see Figure 1 ) extracts code changes from commits and constructs a file-level input token sequence, which is then fed into CodeBERT for fine-tuning. The code change preprocessing is introduced as follows:
[0050] 1. Extract file-level code changes. First, the information of file-level changes is extracted into separate code change documents. For each document, the deleted and added code lines are further extracted and split into rem-code and add-code segments respectively.
[0051] 2. Process the deleted code and added code lines. For each segment, first, the code line is tokenized into a sequence of tokens, and then the camel case style (e.g., "addData") and snake case style (e.g., "add data") tokens are further split using codeprep (codeprep is the abbreviation of code preprocessing. Since CodeBERT is a pre-trained model designed specifically for processing code and natural language, when pre-training on code data, it is necessary to represent the code in a format that the model can understand. Therefore, the process of processing and transforming code data with codeprep is required to better input it into the model).
[0052] 3. Input token construction. To construct the same input representation in CodeBERT, the same tokenizer 12 used by CodeBERT to concatenate two segments was used. Three separator tokens (i.e., [CLS], [SEP], and [EOS]) are used for concatenation. [CLS] is a special classification token that is always the first token of the input. The final hidden state of CodeBERT corresponding to [CLS] is considered as the context embedding of the input tokens (i.e., the aggregated sequence embedding) and can be used for classification. [SEP] and [EOS] are used to separate the two segments and indicate the end of the input respectively. All inputs are padded or truncated to the same length (i.e., 256) by the CodeBERT tokenizer. The following figure shows an example of the constructed input tokens.
[0053] In the training phase of the data, it is divided into two steps: calculating the transition probability and correcting the data to obtain the prediction result.
[0054] 1. Calculate the transition probability matrix. First, for the untrusted data, the output is obtained through the classification model. For the convenience of describing the workflow, the classification component is called Let and y , be the set of possible labels, x be the feature vector of each data in the silver dataset, and the transition matrix C represents the transition probability from the true label y to the predicted label The transition probability To obtain the transition probability The following formula can be used:
[0055]
[0056] The integral here represents the weighted average of the values of x, where p( x|y ) represents the probability distribution when the true label is y is x Since the predicted label is obtained by the data feature x through the classifier That is to say The distribution of x only depends on y and has nothing to do with Using to replace we get the calculation formula of the transition probability. Next, we need to approximately calculate this integral to obtain the required transition probability. The formula for approximate calculation is as follows:
[0057]
[0058] In this formula, A i represents the set of all samples with the label y = i (i ∈ {0, 1}), that is According to the Monte Carlo method in statistics, the sample mean can be used to approximate this integral, and the value of the calculated integral is the value of the transition probability.
[0059] Fill the obtained transition probability and the true label into the initialized 2x2 transition probability matrix. That is, each element in the matrix Cij is stored in this corresponding way:
[0060]
[0061] Denotes the probability that the predicted label is j when the true label is i.
[0062] Estimate the probability transition matrix through the above process. In this process, the process of calculating the transition probability depends on and y are independent of each other. In the experiment, this premise is reasonable because The value of is obtained through the sample features x Through the classifier Predicted. In this process, it is not necessary to know the information of the true label y. Therefore, the entire process of calculating the transition probability is effective.
[0063] 2. Calibrate the data. After obtaining the transition probability matrix C, we reclassify the untrusted data through the classifier To obtain the output S, perform Then use the cross-entropy loss function Focalloss with a regularization factor to train for several rounds to make the AUC value of the prediction result (the AUC value is between 0 and 1. The closer the AUC is to 1, the better the classification performance of the model). Since in the industry, if the AUC value is greater than 0.7, the method can be considered available, so we train for 60 rounds to make the AUC value reach the best, and obtain the classification result of the untrusted data output at this time. For the trusted data, we directly obtain the result through the classifier Training.
[0064] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for detecting code vulnerability repair using noise transfer, characterized in that: The steps include: S1, merge the existing vulnerability datasets and the datasets collected from real open source projects to form the total dataset; S2. Use CodeBERT to extract and embed the code changes of the total data set, and obtain the code change features through the feature extractor; S3. Input the code change features into the Transformer-based pre-trained model, generate output through the pre-trained model, and convert the output into a probability distribution for each category; the probability distribution indicates the model's prediction confidence for each category, set a confidence threshold, and use the confidence threshold to divide the code change features into reliable data with high confidence and unreliable data with low confidence according to the confidence level; S4, inputting the untrusted data into the classifier, and after obtaining the output result of the untrusted data, jointly calculating the trusted data and the untrusted data to obtain the transition probability of the untrusted data; S5. construct a transition probability matrix according to the obtained transition probability, and perform operations on the transition probability matrix and the output result of the untrustworthy data to obtain the classification result of the untrustworthy data; S6. Input the trusted data into the classifier to obtain the classification result of the trusted data.
2. The method for detecting code vulnerability repair using noise transfer according to claim 1, characterized in that: S2 specifically includes the following steps: S21, extracting file-level code changes; extracting information of file-level changes into separate code change documents; for each document, further extracting deleted and added code lines, and splitting the deleted and added code lines into rem-code and add-code segments, respectively; S22, process deleted code and added code lines; for each snippet, first tokenize the code lines into a series of tokens, and then further split the camel case style and snake case style tokens by using codeprep; S23. Input token construction: To construct input that meets the requirements of CodeBERT, the RobertaTokenizer pre-trained by CodeBERT is used to connect code snippets through [CLS], [SEP], and [EOS] delimiters.
3. The method for detecting code vulnerability repair using noise transfer according to claim 1, characterized in that: S5 specifically includes the following steps: For untrusted data, the output is obtained through the classification model, and the classification component is called set up and y, is a possible label set, x is the feature vector of each data in the silver dataset, and the transfer matrix C is represented as the transition from the true label y to the predicted label The transition probability Use the following formula to get the transition probability The integral represents the weighted average of the values of x, p ( x|y ) Represents the probability distribution of x when the true label is y; use replace The calculation formula of the transition probability is obtained, and the integral of the transition probability is approximately calculated to obtain the following transition probability formula: Among them, A i represents the set of all samples with label y=i(i∈{0,1}), that is, 4. The method for detecting code vulnerability repair using noise transfer according to claim 1, characterized in that: S5 specifically includes the following steps: Fill the obtained transition probability and true label into the initialized 2x2 transition probability matrix: It represents the probability of predicting label j when the true label is i. Pass untrusted data through the classifier The result is output to S, and the new output is Then, the cross entropy loss function Focalloss with an adjustment factor is used for training. After several rounds of training, the AUC value of the prediction result reaches the specified value, and the classification result of the unreliable data output at this time is obtained.
5. A storage medium, characterized in that: The storage medium includes a stored program, wherein when the program is run, the method for code vulnerability repair detection using noise transfer described in any one of claims 1 to 4 is executed.
6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: The processor executes the method for detecting code vulnerability repair by using noise transfer as described in any one of claims 1 to 4 through the operation of the computer program.