Method for carrying out code vulnerability repair detection by utilizing multiple attention mechanisms

By adopting multiple attention mechanisms in code vulnerability repair detection, processing related and irrelevant features of patch content, calculating attention weights and weighting, the problem of insufficient detection capabilities of complex code repair scenarios in the existing technology is solved, and more accurate and robust vulnerability repair detection is achieved.

CN120068084APending Publication Date: 2025-05-30DALIAN MARITIME UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510070419.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing technology has poor detection capabilities for complex code repair scenarios in code vulnerability repair detection, and it is difficult to accurately capture the code modification content and its impact on the entire code structure. Especially when code modifications across files or modules are easily missed or misdetected.

Method used

Multiple attention mechanisms are adopted to process patch content related and irrelevant features through the self-attention component and the matching attention component, calculate attention weights and weight processing to improve the accuracy of feature extraction and performance ability in complex scenarios.

Benefits of technology

Effectively represent the mutual attention between the relevant and irrelevant patch content, improves the accuracy and robustness of vulnerability repair detection, enhances the model's performance ability in complex scenarios, and reduces the impact of noise.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068084A_ABST
    Figure CN120068084A_ABST
Patent Text Reader

Abstract

The invention provides a method for performing code bug repair detection by using a multiple attention mechanism. The method comprises the following steps of S1, constructing a bug repair data set; codeBERT is used for extracting and embedding code changes in the vulnerability repair data set, and the characteristics of the code changes are obtained through a characteristic extractor; s2, dividing patch content related features and patch content unrelated features by calculating the similarity between the features of code change; s3, obtaining features processed by the self-attention component; and S4, mapping the patch-independent and patch-related features processed by the self-attention component to the same latitude, inputting the patch-independent and patch-related features into a matched attention component for joint calculation to obtain an attention weight, and weighting through the attention weight to obtain a prediction result. According to the method, multiple attention mechanisms are introduced, the global dependency relationship of the code snippets is captured through a self-attention mechanism, and the difference before and after code modification is analyzed through a matching attention mechanism, so that the recognition precision of vulnerability repair is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of vulnerability repair, and in particular, to a method for detecting code vulnerability repair using a multi-attention mechanism. Background Art

[0002] Code vulnerability repair is an important part of software development and maintenance. The repair process generally includes vulnerability detection, analysis, solution formulation, repair implementation, test verification, deployment and online, etc. In the prior art, some methods have tried to use deep learning models to detect code vulnerability repair, but these methods have certain limitations in technical implementation. First, existing models based on a single attention mechanism, especially the self-attention mechanism, use self-attention to measure the correlation between different segments in the code, so as to model the global dependency relationship of the code. Through the self-attention mechanism, the model can focus on important segments in the code and improve the performance of vulnerability detection. However, these methods mainly focus on capturing the semantics and dependencies of the code, and do not fully consider the difference relationship before and after code changes, especially the distinction between patch-related content and irrelevant content.

[0003] In addition, existing methods have deficiencies in the processing ability of complex code. The prior art usually can only analyze the local features of the code and rely on the context information of the code sequence for vulnerability detection. Since these methods lack a more detailed code change analysis mechanism, they often cannot accurately capture the modified content (patch) of the code and its impact on the entire code structure. Especially when dealing with code modifications across files or modules, the prior art is prone to missing detections or false detections of the repaired parts. This limitation leads to confusion between patch-related and irrelevant code, thus affecting the accuracy of vulnerability repair detection.

[0004] Another deficiency of the prior art is the limited ability to handle complex code structures and cross-module dependency relationships. Although the self-attention mechanism can capture long dependencies in the code, when it comes to multi-module and multi-level dependencies, the local feature extraction mechanism of these methods cannot effectively express the global changes of the code. In addition, the current technology lacks in-depth exploration of specific change details involved in code repair, and only relies on the global modeling of the self-attention mechanism, making it difficult to effectively analyze the key content of code changes, resulting in deviations in the vulnerability repair detection results. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to propose a method for detecting code vulnerability repair using a multi-attention mechanism to solve the technical problem of the existing weak detection ability for complex code repair scenarios.

[0006] The technical means adopted by the present invention are as follows:

[0007] A method for detecting code vulnerability repair using a multi - attention mechanism, comprising the following steps:

[0008] S1. Construct a vulnerability repair dataset; use CodeBERT to extract embeddings for code changes in the vulnerability repair dataset, and obtain the features of the code changes through a feature extractor;

[0009] S2. By calculating the similarity between the features of the code changes, divide the features related to the patch content and the features unrelated to the patch content;

[0010] S3. Input the features related to the patch content and the features unrelated to the patch content into the self - attention component respectively to obtain the features processed by the self - attention component;

[0011] S4. Map the patch - unrelated and patch - related features processed by the self - attention component to the same dimension, input them into the matching attention component for joint calculation to obtain attention weights, and obtain the prediction result by weighting with the attention weights.

[0012] Further, S1 specifically includes the following steps:

[0013] S11. Extract file - level code changes; extract the change information of the file from before vulnerability repair to after vulnerability repair as an independent document, extract the deleted and added code lines, and organize them into rem - code and add - code segments respectively;

[0014] S12. Process the code lines; mark the deleted and added code lines as token sequences, and decompose the variable styles of camel - case naming and underscore naming;

[0015] S13. Construct input tokens; use the tokenization method of CodeBERT to connect the code segments through special tokens such as [CLS], [SEP] and [EOS] to form a fixed - length input sequence.

[0016] Further, the self - attention component is constructed based on Transformer, and the construction method is as follows:

[0017] Calculate the query matrix Q, key matrix K and value matrix V according to the input feature X, and the expressions are as follows:

[0018] Q = XW Q

[0019] K = XW K

[0020] V = XW V

[0021] Calculate the attention scores and normalize them using the softmax function, where d kAct on the normalized dot product result to prevent the value from being too large or too small:

[0022]

[0023] Use multi-head attention so that the self-attention component can focus on features at different positions and obtain information about the patch content from multiple perspectives. The multi-head attention divides the queries, keys, and values into multiple heads, then calculates the attention separately and concatenates the results together:

[0024] MultiHead(Q,K,V)=Concat(head 1 ,head 2 ,...,head h )W O

[0025] In Transformer, each encoder and decoder layer contains a feed-forward neural network that acts independently on each position and includes two fully-connected layers and a non-linear activation function; the output x of each attention head passes through the FFN:

[0026] FFN(x)=max(0,xW 1 +b 1 )W 2 +b 2

[0027] Apply residual connection and layer normalization to the output of the feed-forward neural network to obtain the feature y after being processed by the self-attention component:

[0028] y=LayerNorm(x+Sub-layer(x)).

[0029] Furthermore, S4 specifically includes the following steps:

[0030] Map the features processed by the self-attention component and the matching attention component and input them into the matching attention component. The matching attention component includes a linear dimension matching and feature transpose layer of the input features, an attention matrix calculation layer, a connection layer, and a pooling layer;

[0031] Map the patch content-related data D processed by the self-attention component to the dimension of the patch content-unrelated data F, expressed as:

[0032] D¢=DW+b

[0033] where D′ is the patch content-related data after mapping, W is the weight matrix, and b is the bias vector;

[0034] Calculate the similarity matrix S in the way of dot product. The definition of the element S ij in S is:

[0035]

[0036] Based on the similarity matrix, calculate the attention weights:

[0037]

[0038] Perform attention weight weighting on the data related to the patch content and the data unrelated to the patch content respectively:

[0039]

[0040] Perform a linear transformation on the feature dimension of the weighted data related to the patch content to transform it into the initial dimension:

[0041] attendtioned_D = linear(attentioned_D)

[0042] Make the weighted features pass through the ReLU activation function, the dropout layer and the fully connected layer to convert the dimension to the binary classification dimension and obtain the prediction result.

[0043] The present invention also provides a storage medium, the storage medium includes a stored program, wherein when the program runs, it executes any one of the above methods for code vulnerability repair detection using a multiple attention mechanism.

[0044] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and operable on the processor, and the processor runs through the computer program to execute any one of the above methods for code vulnerability repair detection using a multiple attention mechanism.

[0045] Compared with the prior art, the present invention has the following advantages:

[0046] The technical solution provided by the present invention divides the patch content into two types: relevant and irrelevant. By separately processing them with the self-attention component and then inputting them into the matching attention component for joint calculation to obtain attention weights, and using these attention weights for weighted processing, it can effectively represent the mutual attention between these two types and obtain more accurate prediction results. In addition, by dividing the patch content into relevant and irrelevant, the present invention can capture information that cannot be obtained by only using a single processing method of relevant or irrelevant patch content, enhancing the model's performance ability in dealing with complex scenarios. In a dataset containing a large amount of noise, the present invention can help the model obtain the important parts of the data, perform self-attention on them, and obtain important information. The main advantages of the present invention are as follows: (1) Dual attention mechanism: separately process relevant and irrelevant patch content data, and perform joint calculation through matching attention to improve the accuracy of feature extraction. (2) Capture implicit information: obtain additional information that cannot be obtained by a single processing method by dividing the patch content, enhancing the performance in complex scenarios. (3) Noise suppression: can focus on important data, filter noise, and improve the model's performance in a high-noise dataset. (4) Dynamic weighted processing: use attention weights to weight the data, improving the flexibility of feature fusion and prediction accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0048] Figure 1 It is a flowchart of the method of the present invention.

[0049] Figure 2 It is a diagram of the code change preprocessor of the present invention.

[0050] Figure 3 It is a diagram of the internal framework of the self-attention of the present invention.

[0051] Figure 4 It is a diagram of the internal framework of the matching attention of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0052] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative work shall fall within the scope of protection of the present invention.

[0053] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such data used can be interchanged under appropriate circumstances, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0054] As Figure 1 shown, the present invention provides a method for code vulnerability repair detection using a multi-attention mechanism, including the following steps:

[0055] S1. Construct a vulnerability repair dataset; extract embeddings for code changes in the vulnerability repair dataset using CodeBERT, and obtain the features of the code changes through a feature extractor;

[0056] The dataset of the present invention consists of commit information of Java and Python software projects. There are mainly three sources of data for the present invention:

[0057] · A manually managed Java vulnerability repair commit dataset of the SAPKB project. From this data source, 1,055 vulnerability repair links were obtained, spanning 183 Java OSS projects, corresponding to 615 CVEs.

[0058] · All CVEs related to Java and Python disclosed as of January 26, 2021 collected from the MitreCVE database. For Java, a total of 199 commits, 227 issues, and 155 pull requests were obtained, spanning 189 projects, corresponding to 340 CVEs. For Python, a total of 288 commits, 244 issues, and 353 pull requests were obtained, spanning 256 projects, corresponding to 444 CVEs.

[0059] · Use GitHub to collect commit links related to the total number of issues (471) and pull requests (508) from Java and Python projects. This led to 383 Java vulnerability fixes across 101 projects, corresponding to 186 CVEs. For Python, this led to 597 vulnerability fix commits across 141 projects, corresponding to 233 CVEs.

[0060] Merge these datasets and remove duplicate projects. Thus, the final Java dataset includes 1,436 vulnerability fixes and 839,682 non-vulnerability fix commits across 310 projects, corresponding to 839 CVEs. The Python dataset includes 885 vulnerability fixes and 722,291 non-vulnerability fix commits across 256 projects, corresponding to 444 CVEs.

[0061] In the data preprocessing stage, CodeBERT is adopted as the pre-trained language model, and code changes are represented as token sequences in the Bag of Words (BoW) model. The code change preprocessor (as Figure 2 shown) extracts file-level code changes from commit records, generates input token sequences, and inputs them into CodeBERT for fine-tuning. The preprocessing process includes the following steps:

[0062] S11. Extract file-level code changes: Extract the change information of the file as an independent document, extract the deleted and added code lines, and organize them into rem-code and add-code segments respectively.

[0063] S12. Process code lines: Mark the deleted and added code lines as token sequences and decompose the variable styles of camel case naming and underscore naming.

[0064] S13. Build input tokens: Using the tokenization method of CodeBERT, connect the code segments with special tokens [CLS], [SEP], and [EOS] to form a fixed-length input sequence for subsequent processing.

[0065] S2. Calculate the scores of tokens in these two code change features through Term-Frequency Inverse Document Frequency (TF-IDF). Then, we embed all the scores into a vector. Finally, we obtain two vectors and use cosine similarity to calculate the similarity between the two vectors. Based on the similarity between the features of code changes, divide the features related to patch content and the features unrelated to patch content;

[0066] S3. Input the features related to patch content and the features unrelated to patch content into the self-attention component respectively to obtain the features processed by the self-attention component;

[0067] During the training phase of the data, the data is processed through two components, namely the self-attention and the matching attention components, to obtain the prediction results.

[0068] 1. Self-attention component. First, in order to more comprehensively study the patch content-related and patch content-unrelated, the present invention integrates a self-attention component based on Transformer. This component solves the problem of ignoring the internal information of the patch content by capturing the feature information related to the patch content and the feature information unrelated to the patch content respectively. The internal structure of the self-attention component is as Figure 3 shown.

[0069] First, the query matrix Q, the key matrix K, and the value matrix V are calculated according to the input feature X, and the expressions are as follows:

[0070] Q = XW Q

[0071] K = XW K

[0072] V = XW V

[0073] Then, the attention scores are calculated and normalized using the softmax function, where d k acts on the normalized dot product result to prevent the values from being too large or too small:

[0074]

[0075] Next, multi-head attention is used to enable the component to focus on features at different positions and obtain information about the patch content from multiple perspectives. In this process, the multi-head attention divides the query, key, and value into multiple heads, then calculates the attention separately and concatenates the results together.

[0076] MultiHead(Q, K, V) = Concat(head 1 , head 2 ,..., head h )W O

[0077] In Transformer, each encoder and decoder layer contains a feed-forward neural network FeedForwardNeural Network (FFN). This feed-forward neural network acts independently on each position and mainly includes two fully connected layers and a non-linear activation function. The output x of each attention head passes through the FFN.

[0078] FFN(x) = max(0, xW 1 + b 1)W 2 +b 2

[0079] Finally, residual connections and layer normalization are applied to the output of the feed-forward neural network.

[0080] y = LayerNorm(x + Sub-layer(x))

[0081] S4. Map the patch-independent and patch-related features processed by the self-attention component to the same dimension, input them into the matching attention component for joint calculation to obtain attention weights, and obtain the prediction result by weighting with the attention weights.

[0082] Through these steps, the classification model of Vul-Tracker can make full use of the information inside the patch content, so as to more accurately judge the patch content through the patch content, improving the overall performance and robustness of the model.

[0083] 2. Matching attention component. Map the features processed by the self-attention component and the matching attention component, and input them into the matching attention component. The internal structure of the matching attention component is as Figure 4 shown

[0084] When performing matching attention, the two features to be input need to have the same dimension. Therefore, a linear transformation needs to be performed first. In this process, the patch content-related data D processed by the self-attention component is mapped to the dimension of the patch content-independent data F. It can be expressed as:

[0085] D′ = DW + b

[0086] where D′ is the patch content-related data after mapping, W is the weight matrix, and b is the bias vector

[0087] After obtaining D′ and F with the same dimension, calculate the similarity matrix S in a dot product manner. The element S in S ij is defined as:

[0088]

[0089] Based on the obtained similarity matrix, the attention weights can be calculated:

[0090]

[0091] After obtaining the attention weights, the patch content-related data and the patch content-independent data can be weighted by the attention weights respectively:

[0092]

[0093] Finally, perform a linear transformation on the feature dimensions of the weighted patch content-related data to transform them into the initial dimensions.

[0094] attendtioned_D = linear(attentioned_D)

[0095] Through the joint calculation of patch content-related and patch content-unrelated, the attention weights between them are obtained. Using these attention weights for weighting processing can effectively represent the mutual attention between these two categories. The experimental results of the present invention show that this mutual attention mechanism is effective, and it reduces the impact of the repetition between patch content-related and patch content-unrelated data on the classification model. Through mutual attention, information that cannot be obtained by only using a single processing method of patch content-related or patch content-unrelated can be captured.

[0096] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for code vulnerability repair detection using a multi-attention mechanism, characterized in that: The steps include: S1. Build a vulnerability repair dataset; use CodeBERT to extract embeddings for code changes in the vulnerability repair dataset, and obtain features of code changes through feature extractors; S2. Divide the patch content-related features and the patch content-irrelevant features by calculating the similarity between the features of the code changes; S3, inputting patch content-related features and patch content-irrelevant features into the self-attention component respectively to obtain features processed by the self-attention component; S4. Map the patch-independent and patch-dependent features processed by the self-attention component to the same dimension, input them into the matching attention component for joint calculation, obtain the attention weight, and obtain the prediction result by weighting with the attention weight.

2. The method for code vulnerability repair detection using a multiple attention mechanism according to claim 1, characterized in that: S1 specifically includes the following steps: S11. Extract file-level code changes: extract the change information of the file from before the vulnerability is fixed to after the vulnerability is fixed as an independent document, extract the deleted and added code lines, and organize them into rem-code and add-code fragments respectively; S12, process code lines; mark deleted and added code lines as token sequences, and decompose variable styles of camel case naming and underscore naming; S13, construct input token; Using CodeBERT’s tagging method, the code snippets are connected through special tokens [CLS], [SEP], and [EOS] to form a fixed-length input sequence.

3. The method for code vulnerability repair detection using a multiple attention mechanism according to claim 1, characterized in that: The self-attention component is built based on Transformer, and the construction method is as follows: The query matrix Q, key matrix K and value matrix V are calculated based on the input feature X. The expressions are as follows: Q=XW Q K=XW K V=XW V Calculate the attention score and normalize it using the softmax function, where d k Acts on the normalized dot product result to prevent the value from being too large or too small: Using multi-head attention, the self-attention component can focus on features at different positions and obtain information about the patch content from multiple angles. Multi-head attention divides the query, key, and value into multiple heads, and then calculates the attention separately and splices the results together: MultiHead(Q,K,V)=Concat(head1,head2,...,head h )W O In Transformer, each encoder and decoder layer contains a feedforward neural network, which acts independently on each position and includes two fully connected layers and a nonlinear activation function; the output x of each attention head passes through the FFN: FFN(x)=max(0,xW1+b1)W2+b2 Apply residual connections and layer normalization to the output of the feedforward neural network to obtain the feature y after processing by the self-attention component: y=LayerNorm(x+Sub-layer(x)).

4. The method for code vulnerability repair detection using a multiple attention mechanism according to claim 1, characterized in that: S4 specifically includes the following steps: Mapping the features processed by the self-attention component and the matching attention component and inputting them into the matching attention component, wherein the matching attention component includes a linear dimension matching and feature transposition layer of the input features, an attention matrix calculation layer, a connection layer, and a pooling layer; Map the patch content-dependent data D processed by the self-attention component to the dimension of the patch content-independent data F, expressed as: D¢=DW+b Where D′ is the mapped patch content related data, W is the weight matrix, and b is the bias vector; Use the dot product method to calculate the similarity matrix S, the element S in S ij is defined as: Based on the similarity matrix, calculate the attention weight: The attention weights are weighted for the patch content related data and the patch content irrelevant data respectively: The feature dimension of the weighted patch content-related data is linearly transformed to the initial dimension: attentioned_D=linear(attentioned_D) The weighted features are passed through the ReLU activation function, the dropout layer, and the fully connected layer to transform the dimension into a binary classification dimension to obtain the prediction result.

5. A storage medium, characterized in that: The storage medium includes a stored program, wherein when the program is run, the method for code vulnerability repair detection using a multiple attention mechanism as described in any one of claims 1 to 4 is executed.

6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: The processor executes the method for code vulnerability repair detection using a multiple attention mechanism as described in any one of claims 1 to 4 through the computer program.