Multi-granularity vulnerability repair detection method based on confidence learning

Through the multi-grained vulnerability repair detection method based on confidence learning, the problem of insufficient label noise and robustness in the existing technology is solved, and more accurate vulnerability repair detection is achieved, which improves the efficiency and quality of software development.

CN120372619APending Publication Date: 2025-07-25DALIAN MARITIME UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510277195.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

There are problems of insufficient label noise and robustness in the existing vulnerability repair detection methods, resulting in reduced misclassification and generalization capabilities.

Method used

Using a multi-grained detection method based on confidence learning, the object code is preprocessed, separated into repair commits and irrelevant commits, multi-layer information is extracted and encoded into numerical vectors, confidence matrix is calculated, and classifiers are optimized to identify and learn the characteristics of bug fix commits, and noise data are removed.

Benefits of technology

It improves the accuracy and robustness of vulnerability repair detection, saves time and labor costs, and improves software development efficiency and quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372619A_ABST
    Figure CN120372619A_ABST
Patent Text Reader

Abstract

The invention provides a detection method for multi-granularity vulnerability repair based on confidence learning, which comprises the following steps of: S1, performing data preprocessing on a target code, and changing and separating the target code into repair submission and irrelevant submission; s2, multi-layer information is extracted from repair submission, and submission information is coded into numerical vectors at the line, block, file and submission granularity levels; s3, extracting features from the numerical vector of each granularity level, calculating confidence, and obtaining a confidence matrix based on the confidence; and S4, identifying and learning the features of code change bug repair submission by using the confidence coefficient matrix. According to the invention, countless time and labor cost are saved; the method is of great significance in improving detection of code bug repair; the software development efficiency and quality are improved, and software vulnerabilities are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of software security, and in particular, to a detection method for multi-granularity vulnerability repair based on confidence learning. Background Art

[0002] In software development, there is an increasing reliance on third-party libraries, which makes users more vulnerable to security threats from these libraries. The focus on addressing the rising risk of security vulnerabilities in the software ecosystem lies in developing software composition tools that alert users to library vulnerabilities, but these tools face a delay in vulnerability exposure, leaving the system exposed to undetected threats. To identify security-related code changes before public disclosure, tools have been developed that automatically detect bug fix commits using resources such as commit messages or issue reports.

[0003] The detection of bug fixes is a coarse-grained classification method based on a classifier. Many methods have been proposed to address the security risks of third-party libraries, and despite the progress made, challenges still remain in bug fix detection. First, deep learning-based code analysis techniques rely on extensive data for optimal training, facing inherent challenges in the sample collection process. Massively labeling samples as "positive" or "negative" inevitably leads to label noise, resulting in misclassification. In addition, representational noise can also occur through code changes that do not alter the basic semantics. It can also occur by introducing new code features, leading to incorrect label classification. Second, code changes contain the necessary resources to identify bug fix commits, but only a very small portion of these commits are related to vulnerability fixes. Training on specific datasets may overfit to the characteristics of code changes unrelated to vulnerabilities, thereby reducing their ability to detect bug fix commits. This in turn leads to a reduction in robustness and generalization ability. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to propose a detection method for multi-granularity vulnerability repair based on confidence learning to solve the technical problem of label noise existing in the existing bug fix detection methods.

[0005] The technical means adopted by the present invention are as follows:

[0006] A detection method for multi-granularity vulnerability repair based on confidence learning, comprising the following steps:

[0007] S1. Perform data preprocessing on the target code to separate the target code changes into fix commits and irrelevant commits;

[0008] S2. Extract multi-layer information from the fix commits and encode the commit information into numerical vectors at several granularity levels of line, block, file, and commit;

[0009] S3. Extract features from the numerical vectors at each granularity level, calculate the confidence, and obtain a confidence matrix based on the confidence.

[0010] S4. Based on the classifier to be optimized, use the confidence matrix to identify and learn the features of the bug-fix submissions to be repaired, and obtain an optimized classifier.

[0011] S5. Input the bug-fix submissions into the optimized classifier to obtain the detection results.

[0012] Furthermore, S2 includes the following steps:

[0013] Decompose the bug-fix submission into four pairs of code snippets at the granularity levels, namely: submission, file, block, and line; encode the code snippets to be extracted into high-dimensional vectors; capture the specific features of the code changes at each granularity level by fine-tuning CodeBERT at each granularity level, and then use the fine-tuned model as a code embedding model to represent the code snippets; the code embedding model accepts two segments as input: one from natural language and one from a programming language, and the input format is:

[0014] [CLS]<NL>[SEP]<PL>[EOS]

[0015] Among them, [CLS] represents the start of the CodeBERT sequence, followed by natural language text; the [SEP] token separates the natural language text and the programming language source code, and [EOS] represents the end of the CodeBERT sequence.

[0016] Furthermore, S3 includes the following steps:

[0017] S31. Extract features at the line-level granularity;

[0018] Use a recursive neural network to extract features at the line-level granularity;

[0019] f line = BiLSTM([l1, l2,... l Tline )

[0020] Interpret the code change as a series of lines, each line represented by an embedding vector, i.e., [l1, l2,... l Tline , and use a bidirectional long short-term memory model, i.e., BiLSTM, as a feature extractor;

[0021] S32. Extract features at the block-level granularity;

[0022] Use a convolutional neural network to extract features at the block-level granularity; among them, the features of the i-th block-level embedding vector are represented by aggregating information from adjacent embedding vectors:

[0023] h i ' = Conv([h i ' -(w-1) / 2 ,...h i ' +(w-1) / 2 )

[0024] The max - pooling layer extracts key features from the input embeddings, generating the final representation:

[0025] f hunk = MaxPool([h1', h'2,...h' H )

[0026] S33. Extract features at the file - level granularity;

[0027] The file - level contains high - level relationships between all the codes in a single submission. A fully - connected neural network is used to capture the relationships between files in the submissions:

[0028]

[0029] S34. Extract features at the submission - level granularity;

[0030] A fully - connected neural network is used to capture the relationships between submissions:

[0031] f commit = FCN(x)

[0032] S35. For different types of data, two strategies of bimodal fusion and unimodal fusion are adopted;

[0033] For the bimodal representation, a single feature vector is obtained and directly input into a linear layer for feature fusion; for the unimodal representation, it is first concatenated into a vector, and then the vector is input into a linear layer for fusion; a neural network classifier is used to combine multi - level submission features to calculate the confidence value of code change samples for bug fixing.

[0034] Furthermore, S4 includes the following steps:

[0035] S41. Describe the joint distribution between noise labels and true labels. For each instance in the code change, calculate the probability that it is a bug - fixing submission and perform calibration; the representation of the confidence joint matrix is as follows:

[0036]

[0037] where the threshold t j for each class represents the expected confidence level for that specific class;

[0038]

[0039] Use a confidence joint matrix to estimate the noise in code changes, which is represented by the joint distribution matrix That is:

[0040]

[0041] S42. Use probability sorting to clean the jointly estimated, pruned, and sorted data, and estimate and clean incorrect labels;

[0042] S43. Input the code changes after denoising in S42 into the model for retraining, so as to identify and learn the characteristics of bug-fix commits in the code changes, and obtain an optimized classifier.

[0043] The present invention also provides a storage medium, which includes a stored program. When the program runs, it executes the detection method for multi-granularity vulnerability repair based on confidence learning as described in any one of the above.

[0044] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor runs the computer program to execute the detection method for multi-granularity vulnerability repair based on confidence learning as described in any one of the above.

[0045] Compared with the prior art, the present invention has the following advantages:

[0046] The technical solution provided by the present invention performs the detection of vulnerability repair through the detection method for multi-granularity vulnerability repair based on confidence learning. The detection of vulnerability repair is a coarse-grained classification method and is implemented based on a classifier. The present invention optimizes the performance of the classifier based on the original classifier. This method uses a code change confidence matrix to remove noisy data and achieve effort-aware adjustment, solving the problems of label and representation noise in code changes and the lack of robustness of the model. The present invention saves countless time and labor costs; the present invention is of great significance for improving the detection of code vulnerability repair; the present invention improves the development efficiency and quality of software and reduces software vulnerabilities. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0048] Figure 1 It is a flowchart of the method of the present invention. Detailed implementation manners

[0049] In order to enable those skilled in the art of the present technology to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0050] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0051] As Figure 1 shown, the present invention provides a detection method for multi-granularity vulnerability repair based on confidence learning, which includes the following steps:

[0052] S1. Perform data preprocessing on the target code, and separate the target code changes into repair commits and irrelevant commits;

[0053] The present invention uses the datasets proposed by VulFixMiner and SAP to evaluate its effectiveness. The datasets include bug fixes and non-bug fix commits from 150 Java and 106 Python open-source projects. The Java dataset contains 1,436 bug fix commits and 474,555 non-bug fix commits, and the Python dataset includes 885 bug fix commits and 357,696 non-bug fix commits.

[0054] S2. Extract multi-level information from the repair commits, and encode the commit information into numerical vectors at several granularity levels of line, block, file, and commit;

[0055] In terms of the structure of code submissions, a submission consists of a set of code changes applied to a set of files. Each file consists of multiple code blocks located in a specific area, and these code blocks are a series of code changes applied to lines. The method of the present invention follows the natural organizational structure of submissions and decomposes the submission into four pairs of code snippets at different granularity levels, namely: submission, file, block, and line. For example, at the line-level granularity, the code changes are divided into many lines that have been submitted.

[0056] Subsequently, the code snippets to be extracted are encoded into high-dimensional vectors. By fine-tuning CodeBERT at each granularity level to capture the specific features of code changes at each level, the fine-tuned model is then used as a code embedding model to represent the code snippets. By default, it accepts two segments as input: one from natural language and one from a programming language, and the input format is:

[0057] [CLS]<NL>[SEP]<PL>[EOS]

[0058] Among them, [CLS] represents the start of the CodeBERT sequence, followed by the natural language text. The [SEP] token separates the natural language text and the programming language source code, and [EOS] represents the end of the CodeBERT sequence. CodeBERT is pre-trained for different data patterns, bimodal data (natural language and source code pairs) and unimodal data (only source code). Therefore, it is necessary to pay attention to the added and deleted code in the input submission and consider the existence of context in two different ways.

[0059] S3. Extract features from the numerical vectors at each granularity level, calculate the confidence, and obtain a confidence matrix based on the confidence;

[0060] Since the features at the four granularity levels are different, corresponding different models are adopted. And the feature extractors at each granularity level follow a common structure.

[0061] S31. Line level; Considering that the lines in the code changes are read sequentially, a Recurrent Neural Network (RNN) is utilized, which is a standard model for processing sequential data. This is used to extract features at the line-level granularity.

[0062]

[0063] The code changes are interpreted as a series of lines, and each line is represented by an embedding vector, that is And a bidirectional long short-term memory model, namely BiLSTM, is used as the feature extractor;

[0064] S32. Block level; Different from lines, large chunks in a commit do not carry an order relationship, but there are still dependencies between adjacent chunks. Therefore, a convolutional neural network is used. Among them, the features of the i-th block-level embedding vector are represented by aggregating information from adjacent embedding vectors.

[0065] h i ' = Conv([h i ' -(w-1) / 2 ,...h i ' +(w-1) / 2 )

[0066] After that, the max pooling layer extracts key features from the input embeddings to produce the final representation. That is:

[0067] f hunk = MaxPool([h1', h'2,...h' H )

[0068] S33. File level; The file level contains the high-level relationships between all the codes in a commit. Therefore, a fully connected neural network is used to capture the relationships between files in a commit. That is:

[0069]

[0070] S34. Commit level; Similar to the feature extraction at the file level. That is:

[0071] f commit = FCN(x)

[0072] For different types of data, two strategies of bimodal fusion and unimodal fusion are adopted. For bimodal representation, a single feature vector is obtained and directly input into a linear layer for feature fusion. For unimodal representation, they are first concatenated into a vector, and then the vector is input into a linear layer for fusion. Finally, a neural network classifier is used to combine multi-level commit features to calculate the confidence value of the code change sample for bug fixing.

[0073] S4. Identify and learn the features of code change bug fix commits using the confidence matrix.

[0074] Enhance learning by identifying the representational noise and label errors in code changes, and adopt three processes for data denoising.

[0075] S41. Counting; Describes the joint distribution between the noise label and the true label. For each instance in the code change, calculate its possibility of being a bug fix commit and make a correction. The representation of the confidence joint matrix is as follows:

[0076]

[0077] Among them, the threshold t for each category j represents the expected confidence level for that specific category.

[0078]

[0079] The confidence joint matrix not only performs well in anomaly detection but also provides considerable flexibility in threshold selection. Using the confidence joint matrix to estimate the noise in code changes, it is represented by the joint distribution matrix That is:

[0080]

[0081] S42. Cleaning; cleaning the jointly estimated, pruned, and sorted data using probability sorting. Considering that identifying bug-fix commits is a binary classification task and actual bug-fix commits are rare, the task of the present invention is to estimate and clean incorrect labels.

[0082] S43. Retraining; using the above method, after filtering out the representational noise and label noise, inputting the denoised code changes into the model for retraining, so as to identify and learn the characteristics of bug-fix commits in the code changes.

[0083] S5. Input the bug-fix commit into the optimized classifier to obtain the detection result. The detection result refers to whether this fix commit is valid.

[0084] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A detection method for multi-granularity vulnerability repair based on confidence learning, characterized in that, It includes the following steps: S1. Perform data preprocessing on the target code, and separate the target code changes into fix commits and unrelated commits; S2. Extract multi-level information from the fix commits, and encode the commit information into numerical vectors at several granularity levels of line, block, file, and commit; S3. Extract features from the numerical vectors at each granularity level, calculate the confidence, and obtain a confidence matrix based on the confidence; S4. On the basis of the classifier to be optimized, use the confidence matrix to identify and learn the features of the fix commits of bug fix commits, and obtain an optimized classifier; S5. Input the commits for vulnerability repair into the optimized classifier to obtain the detection results.

2. The detection method for multi-granularity vulnerability repair based on confidence learning according to claim 1, wherein S2 includes the following steps: Decompose the fix commits into code fragments at four granularity levels, namely: commit, file, block, and line; encode the code fragments to be extracted into high-dimensional vectors; capture the specific features of code changes at each level by fine-tuning CodeBERT at each granularity level, and then use the fine-tuned model as a code embedding model to represent the code fragments; the code embedding model accepts two segments as input: one from natural language and one from programming language, and the input format is: [CLS]<NL>[SEP]<PL>[EOS] Among them, [CLS] represents the start of the CodeBERT sequence, followed by natural language text; the [SEP] token separates the natural language text and the programming language source code, and [EOS] represents the end of the CodeBERT sequence.

3. The detection method for multi-granularity vulnerability repair based on confidence learning according to claim 1, characterized in that, S3 includes the following steps: S31. Extract features at the line-level granularity; Use a recurrent neural network to extract features at the line-level granularity; f line = BiLSTM([l1, l2,... l Tline ) Interpret the code changes as a series of lines, each represented by an embedding vector, i.e., [l1, l2,... l Tline , and use a bidirectional long short-term memory model, i.e., BiLSTM, as a feature extractor; S32. Extract features at the block-level granularity; Use a convolutional neural network to extract features at the block-level granularity; among them, the features of the i-th block-level embedding vector are represented by aggregating information from adjacent embedding vectors: h i ' = Conv([h i ' -(w-1) / 2 ,...h i ' +(w-1) / 2 ) The max pooling layer extracts key features from the input embeddings to produce the final representation: f hunk = MaxPool([h1', h'2,... h' H ) S33. Extract features at the file-level granularity; The file-level contains the high-level relationships between all the codes in a commit, and a fully connected neural network is used to capture the relationships between files in the commits: S34. Extract features at the commit-level granularity; Use a fully connected neural network to capture the relationships between commits: f commit = FCN(x) S35. For different types of data, adopt two strategies of bimodal fusion and unimodal fusion; For the bimodal representation, obtain a single feature vector and directly input it into a linear layer for feature fusion; for the unimodal representation, first concatenate it into a vector, and then input the vector into a linear layer for fusion; use a neural network classifier to combine multi-level commit features to calculate the confidence value of the code change samples for bug repair.

4. The detection method for multi-granularity vulnerability repair based on confidence learning according to claim 1, wherein, S4 includes the following steps: S41. Describe the joint distribution between the noise labels and the true labels. For each instance in the code change, calculate its probability of being a bug fix commit and perform correction; the representation of the confidence joint matrix is as follows: where the threshold t for each category j represents the expected confidence level for that particular category; Use a confidence joint matrix to estimate the noise in code changes, represented by the joint distribution matrix That is: S42. Use probability sorting to clean the jointly estimated, pruned, and sorted data, and estimate and clean the incorrect labels; S43. Retrain the input model with the code changes after denoising in S42, so as to identify and learn the features of bug-fix commits in the code changes, and obtain an optimized classifier.

5. A storage medium, characterized in that, The storage medium includes a stored program, wherein when the program runs, it executes the detection method for multi-granularity vulnerability repair based on confidence learning according to any one of claims 1 to 4.

6. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor runs and executes the detection method for multi-granularity vulnerability repair based on confidence learning according to any one of claims 1 to 4 through the computer program.