Code detection method, electronic device and computer readable medium

By preprocessing and embedding representation of the detection code, the modeling of the semantic vector sequence of code is enhanced by using relative position coding, the problem of inaccurate semantic representation of code detection in the prior art is solved, and the accuracy and generalization ability of code detection are improved.

CN119622750BActive Publication Date: 2025-05-13XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510156853.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-05-13
Estimated Expiration
2045-02-13

AI Technical Summary

Technical Problem

In the prior art, the code detection method based on neural network model is inaccurate in semantic representation in complex code scenarios, resulting in missed detection or false positives, affecting the accuracy of code detection.

Method used

By preprocessing the detection code, the code word block index sequence is obtained, and the code word block index is embedded to represent it, and the relative position encoding is obtained. Based on this relative position encoding, the embedding representation of the code word block index sequence is modeled after embedding enhancement to obtain the code semantic vector sequence. Then, vulnerability detection is performed on the code semantic vector sequence to obtain the code detection results.

Benefits of technology

By introducing relative position coding, it can better capture the relative position information of code word blocks, enhance the acquisition of local structural information of code semantic vector sequences, generate more accurate code semantic vector representations, improve the accuracy and generalization ability of code detection, and avoid global position deviations caused by relying on absolute position information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119622750B_ABST
    Figure CN119622750B_ABST
Patent Text Reader

Abstract

The present invention discloses a code detection method, an electronic device and a computer-readable medium. When performing code detection, the code to be detected is first preprocessed to obtain a code word block index sequence; then, the code word block index sequence is embedded to represent the code word block index sequence, and the relative position encoding of each word block in the code word block index sequence is obtained, and based on the relative position encoding, the embedded representation of the code word block index sequence is embedded and enhanced and then modeled to obtain a code semantic vector sequence; then, the code semantic vector sequence is tested for vulnerabilities to obtain code detection results. The present invention introduces relative position encoding during code detection, which can better capture the relative position information of each word block in the code word block index sequence; and based on the relative position encoding, the embedded representation of the code word block index sequence is embedded and enhanced, generating a vector representation of code semantics with rich semantic information and more accurate, which greatly improves the accuracy of code detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a code detection method, an electronic device and a computer-readable medium. Background Art

[0002] In the process of software design, implementation and maintenance, some errors or defects are inevitable. These errors or defects are likely to be maliciously attacked, thus affecting the integrity, availability and confidentiality of the software, and further threatening the privacy and data security of users; therefore, code detection technology has become a necessary means to maintain software security.

[0003] In recent years, deep learning, as a powerful machine learning technology, has been widely used in the field of software security. Among the related technologies, there is a code detection method based on a neural network model and an attention mechanism. The neural network model can learn the features in the vulnerability code dataset, and the introduction of the attention mechanism enables the neural network model to focus on the key information of the features and ignore or weaken the non-critical information.

[0004] However, the semantic representation accuracy of the vulnerability code dataset by the above method is still limited, especially when the code logic is complex and the dependency is high; this inaccurate semantic representation may lead to missed detections or false positives during code detection, affecting the accuracy of code detection. Summary of the invention

[0005] In order to solve the above problems existing in the prior art, the present invention provides a code detection method, an electronic device and a computer readable medium. The technical problem to be solved by the present invention is achieved by the following technical solutions:

[0006] According to a first aspect of an embodiment of the present invention, a code detection method is provided, the method comprising:

[0007] Preprocess the code to be detected to obtain a code word block index sequence;

[0008] The code word block index sequence is embedded to represent the code word block index sequence, a relative position code of each word block in the code word block index sequence is obtained, and based on the relative position code, the embedding representation of the code word block index sequence is embedded and enhanced and then modeled to obtain a code semantic vector sequence;

[0009] Vulnerability detection is performed on the code semantic vector sequence to obtain a code detection result.

[0010] Optionally, the embedding representation of the code word block index sequence to obtain the relative position encoding of each word block in the code word block index sequence includes:

[0011] By using a target language model, the code block index sequence is embedded and represented to obtain a relative position code of each block in the code block index sequence;

[0012] The method of performing embedding enhancement and modeling on the embedding representation of the code word block index sequence based on the relative position encoding to obtain a code semantic vector sequence comprises:

[0013] Using the attention weight calculation formula , calculating the attention score of each word block in the code word block index sequence in the target language model to obtain a first attention weight;

[0014] Based on the first attention weight, embedding enhancement is performed on the embedding representation of the code word block index sequence to model the code semantic vector sequence;

[0015] In the attention weight calculation formula, is a query matrix, which represents the information that any word block in the code word block index sequence wants to obtain from other words blocks; is a key matrix used to determine the correlation between any word block in the code word block index sequence and other word blocks, Represents the key matrix The transposed matrix of represents the dimension of the key matrix; is a relative position matrix, used to represent the relative position distance of the word blocks in the code word block index sequence, and the relative position matrix is ​​constructed based on the relative position encoding; represents the first attention weight.

[0016] Optionally, performing vulnerability detection on the code semantic vector sequence to obtain a code detection result includes:

[0017] The bidirectional gated recurrent unit is used to perform function vulnerability detection on the code semantic vector sequence to obtain preliminary code detection results.

[0018] Optionally, after performing function vulnerability detection on the code semantic vector sequence by using a bidirectional gated loop unit to obtain a preliminary code detection result, the method further includes:

[0019] When the preliminary code detection result indicates that there is a vulnerability, obtaining an attention score of each semantic vector in the bidirectional gated recurrent unit in a code semantic vector sequence corresponding to the function where the vulnerability is located, and obtaining a second attention weight;

[0020] The code semantic vector sequence corresponding to the function where the vulnerability is located is split according to the line break character to obtain a semantic vector line list;

[0021] Based on the semantic vector row list, weighted summing the first attention weight and the second attention weight by row is performed to obtain a row attention score of the code semantic vector sequence corresponding to the function where the vulnerability is located;

[0022] The row attention scores are sorted in descending order, and the final code detection result is obtained according to the sorting result.

[0023] Optionally, the embedding representation of the code word block index sequence to obtain the relative position encoding of each word block in the code word block index sequence includes:

[0024] The code block index sequence is embedded and represented through the pre-trained CodeT5+ language model to obtain the relative position encoding of each block in the code block index sequence.

[0025] Optionally, the CodeT5+ language model is trained in the following manner:

[0026] Obtaining a training code sample, wherein the training code sample includes: a data set containing row-level vulnerabilities and row labels corresponding to row data in the data set;

[0027] The CodeT5+ language model is trained using the training code samples.

[0028] Optionally, the using the training code sample to train the CodeT5+ language model includes:

[0029] Performing word segmentation processing on the training code sample to obtain a sample word block vector;

[0030] Performing index mapping on the sample chunk vectors according to the vocabulary to obtain a sample chunk index sequence;

[0031] The sample chunk index sequence is embedded to obtain a relative position code of each chunk in the sample chunk index sequence, and the embedded representation of the sample chunk index sequence is embedded and enhanced based on the relative position code to perform modeling to obtain a sample semantic vector sequence;

[0032] Using the target loss function, calculating the loss value of the sample semantic vector sequence;

[0033] The CodeT5+ language model is trained according to the loss value.

[0034] Optionally, performing word segmentation processing on the training code sample to obtain a sample word block vector includes:

[0035] The training code sample is segmented by using a subword tokenization method to obtain a sample word block vector.

[0036] According to a second aspect of an embodiment of the present invention, there is provided an electronic device, the device comprising:

[0037] one or more processors;

[0038] A computer readable medium configured to store one or more computer programs;

[0039] When the one or more computer programs are executed by the one or more processors, the one or more processors implement the code detection method as described in any one of the first aspects.

[0040] According to a third aspect of an embodiment of the present invention, a computer-readable medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the code detection method as described in any one of the first aspects is implemented.

[0041] Compared with the prior art, the present invention has the following beneficial effects:

[0042] The code detection method provided by the embodiment of the present invention, when performing code detection, first pre-processes the code to be detected to obtain a code word block index sequence; then obtains the relative position encoding of each word block in the code word block index sequence, and based on the relative position encoding, performs embedding enhancement and modeling on the embedding representation of the code word block index sequence to obtain a code semantic vector sequence; then performs vulnerability detection on the code semantic vector sequence to obtain a code detection result. The present invention introduces relative position encoding during code detection, which can better capture the relative position information of each word block in the code word block index sequence, thereby enhancing the acquisition of local structural information in the code semantic vector sequence; and based on the relative position encoding, performs dynamic word embedding, that is, embedding enhancement, on the embedding representation of the code word block index sequence, thereby generating a vector representation of code semantics with rich semantic information and more accurate, which can fully capture the dependency and logical structure in the code; compared with the prior art, the present invention has stronger generalization ability in complex code scenarios, and avoids the global position deviation caused by relying on absolute position information in the prior art, greatly improving the accuracy of code detection.

[0043] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 A flowchart of a code detection method provided by an embodiment of the present invention;

[0045] Figure 2 Another step flow chart of the code detection method provided by an embodiment of the present invention;

[0046] Figure 3 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0047] The present invention is further described in detail below with reference to specific embodiments, but the embodiments of the present invention are not limited thereto.

[0048] Embodiment 1

[0049] Reference Figure 1 , shows a step flow chart of a code detection method provided according to embodiment 1 of the present invention.

[0050] The code detection method of the embodiment of the present invention comprises the following steps:

[0051] Step 101: pre-process the code to be detected to obtain a code word block index sequence.

[0052] In an embodiment of the present invention, the code to be detected can be represented by a vector. First, the code to be detected is segmented to obtain a code word block vector. In this step, the subword tokenization word segmentation method can be used to implement the word segmentation. This word segmentation method can effectively avoid the OOV (out-of-vocabulary, unregistered words) problem, that is, when encountering a word that is not in the vocabulary, that is, an unregistered word, the traditional word-based word segmentation method will regard it as a whole unknown symbol, and subword tokenization can decompose it into known subwords, thereby reducing the occurrence of unknown words. Then, for the obtained code word block vector, it is indexed and mapped according to the vocabulary to obtain a code word block index sequence.

[0053] Step 102: embed the code word block index sequence to obtain the relative position code of each word block in the code word block index sequence, and based on the relative position code, perform embedding enhancement and modeling on the embedding representation of the code word block index sequence to obtain a code semantic vector sequence.

[0054] Specifically, the target language model can be used to embed the code block index sequence to obtain the relative position encoding of each block in the code block index sequence; then, based on the relative position encoding, the embedding representation of the code block index sequence is modeled after embedding enhancement to obtain a code semantic vector sequence. The target language model can be a pre-trained CodeT5+ language model, but is not limited thereto.

[0055] Introducing relative position encoding in this step can better capture the relative position information of each word block in the code word block index sequence, thereby obtaining a code semantic vector sequence with rich semantic information.

[0056] Step 103: Perform vulnerability detection on the code semantic vector sequence to obtain code detection results.

[0057] Specifically, the bidirectional gated loop unit can be used to perform function vulnerability detection on the code semantic vector sequence to obtain the code detection result, but it is certainly not limited to this.

[0058] In the embodiment of the present invention, when the detection is passed, that is, when there is no vulnerability in the code, the code detection result can be output as text information or icon information indicating that the detection is passed. When the detection fails, the code detection result can be output as text information or icon information indicating that the detection fails, and the code vulnerability location information is output. The specific output method of the code detection result is not limited by the present invention.

[0059] The code detection method provided by the embodiment of the present invention introduces relative position coding during code detection, which can better capture the relative position information of each word block in the code semantic vector sequence, thereby enhancing the acquisition of local structural information in the sequence; and based on the relative position coding, the embedding representation of the code word block index sequence is dynamically embedded, that is, embedding enhancement, to generate a vector representation of the code semantics with rich semantic information and more accurate, which can comprehensively capture the dependencies and logical structures in the code; compared with the prior art, the embodiment of the present invention has stronger generalization ability in complex code scenarios, and avoids the global position deviation caused by reliance on absolute position information in the prior art, thereby greatly improving the accuracy of code detection.

[0060] Embodiment 2

[0061] The code detection method implemented by the present invention will be described in more detail below. Figure 2 As shown, Figure 2 Another flowchart of the code detection method provided in the embodiment of the present application may include the following steps:

[0062] Step 201: perform word segmentation on the code to be detected to obtain a code word block vector.

[0063] Here, the subword tokenization method can be used to segment the code to be detected, thereby effectively avoiding the OOV problem.

[0064] Step 202: perform index mapping on the code word block vectors according to the vocabulary to obtain a code word block index sequence.

[0065] Step 203: embed the code word block index sequence through the pre-trained CodeT5+ language model to obtain the relative position code of each word block in the code word block index sequence.

[0066] Specifically, the code word block index sequence is input into the pre-trained CodeT5+ language model so that it embeds the code word block index sequence and outputs the relative position encoding of each word block in the code word block index sequence.

[0067] In the embodiment of the present invention, the CodeT5+ language model can be trained in the following manner:

[0068] A training code sample is obtained, where the training code sample includes: a data set containing row-level vulnerabilities, and a row label corresponding to the row data in the data set, where the row label is used to mark whether there is a vulnerability in the row; then, the CodeT5+ language model is trained using the obtained training code sample.

[0069] In an embodiment of the present invention, the training code sample can use the Big-Vul dataset, which is a large dataset containing C / C++ vulnerabilities, including row-level vulnerabilities, and the row-level vulnerabilities include vulnerability row data and their corresponding row labels. The dataset can be randomly divided into a training set, an evaluation set, and a test set in a ratio of 8:1:1, and then the training set is used as a training code sample to train the CodeT5+ language model.

[0070] Specifically, when the CodeT5+ language model is trained using the above-mentioned training code samples, the training code samples are first segmented to obtain sample word block vectors. For example, the subword tokenization method can be used to segment the training code samples; then, the sample word block vectors are indexed and mapped according to the vocabulary to obtain a sample word block index sequence; then, the sample word block index sequence is embedded to represent the sample word block index sequence, the relative position encoding of each word block in the sample word block index sequence is obtained, and based on the relative position encoding, the embedded representation of the sample word block index sequence is embedded and enhanced and then modeled to obtain a sample semantic vector sequence; then, the target loss function is used to calculate the loss value of the sample semantic vector sequence; and the CodeT5+ language model is trained according to the loss value.

[0071] The above-mentioned calculation method of the sample semantic vector sequence is similar to the method of obtaining the code semantic vector sequence, which will not be repeated here. The target loss function can be a loss function commonly used for the loss calculation of the CodeT5+ language model, and the embodiment of the present invention is not limited. The model parameters of the CodeT5+ language model are adjusted according to the loss value calculated by the target loss function until the training is completed; finally, the evaluation set can be used to evaluate the performance of the CodeT5+ language model under different hyperparameter settings, and the optimal hyperparameter combination can be selected as the final model parameters of the trained CodeT5+ language model.

[0072] In an embodiment of the present invention, through training with large-scale data, the CodeT5+ language model can automatically extract potential vulnerability features in the code, significantly improving the accuracy of detection.

[0073] Step 204: Based on the relative position coding, the embedding representation of the code word block index sequence is modeled after embedding enhancement to obtain a code semantic vector sequence.

[0074] Specifically, the attention weight calculation formula is used to calculate the attention score of each word block in the code word block index sequence in the pre-trained CodeT5+ language model to obtain the first attention weight; then based on the first attention weight, the embedding representation of the code word block index sequence is modeled after embedding enhancement to obtain the code semantic vector sequence.

[0075] Among them, the attention weight calculation formula is:

[0076] ;

[0077] In the attention weight calculation formula, is the query matrix, which represents the information that any word block in the code word block index sequence wants to obtain from other words blocks; is a key matrix, which is used to determine the correlation between any word block and other word blocks in the code word block index sequence. Represents the key matrix The transposed matrix of represents the dimension of the key matrix, is a relative position matrix, which is used to represent the relative position distance of the word blocks in the code word block index sequence. The relative position matrix is ​​constructed based on the relative position encoding; represents the first attention weight.

[0078] It can be understood that in step 203, the code word block index sequence is embedded to obtain a sequence with preliminary embedding representation. In step 204, the first attention weight of each word block is obtained by the attention weight calculation formula, and then the word block information in the sequence with preliminary embedding representation is updated according to the first attention weight, that is, the embedding representation of the code word block index sequence is embedded and enhanced. Subsequently, the enhanced embedding representation is context modeled, and finally a code semantic vector sequence containing rich contextual semantic information can be generated.

[0079] The present invention combines relative position coding to enhance the embedding representation of the code word block index sequence to capture the relative relationship between the code word blocks. Subsequently, the enhanced embedding representation can be context modeled to ultimately generate a code semantic vector sequence containing rich context semantic information.

[0080] In an embodiment of the present invention, by introducing relative position encoding, the embedding representation of the code word block index sequence is enhanced to capture the relative relationship between code word blocks, which can more accurately capture the relative position information in the code word block index sequence, thereby enhancing the target language model's acquisition of local structural information in the sequence, avoiding the global position deviation caused by relying solely on absolute position information in related technologies, and improving the performance and flexibility of the model in long sequence processing. Moreover, compared with the traditional static word embedding method, word embedding through the CodeT5+ language model can dynamically capture the context information of the input sequence, that is, the code word block index sequence, and comprehensively capture the dependencies and logical structures in the code. Therefore, it can adjust the value of the embedding vector according to the specific context of the code, so that the vector representation of the generated code semantics is more accurate. Compared with the prior art, the context-aware embedding method is particularly effective in detecting code vulnerabilities, can better identify potential security risks in the code, improve the accuracy and efficiency of code detection, and has stronger generalization capabilities in complex code scenarios.

[0081] Step 205: Use a bidirectional gated recurrent unit to perform function vulnerability detection on the code semantic vector sequence to obtain preliminary code detection results.

[0082] Among them, the bidirectional gated recurrent unit is also called the BiGRU model. This BiGRU model can also be pre-trained. The code semantic vector sequence with rich semantic information output by the trained CodeT5+ language model is used as a sample and input into the BiGRU model for one-step feature learning. The BiGRU model finally performs function-level code vulnerability detection and classification through a linear classification layer, and calculates the loss (Loss) based on the classification results, so as to adjust the learning parameters of the BiGRU model for the sample through backpropagation, thereby achieving a higher classification accuracy. Loss can be calculated using the cross entropy loss function, and its calculation formula is as follows:

[0083]

[0084] in, is the number of samples; is the true value of the label, that is, the real result; It is the probability distribution predicted by the BiGRU model, which is also the output result of the BiGRU model.

[0085] In the embodiment of the present invention, this step 205 is to perform function-level vulnerability detection on the code to be detected. The indicators for evaluating the detection effect include Percision, Recall and F1-Score. F1-Score is the harmonic mean of precision and recall. It comprehensively considers precision and recall and can more comprehensively evaluate the performance of the BiGRU model. When both precision and recall are high, F1-Score will also be high. Conversely, if one of the indicators is low, F1-Score will also decrease accordingly). Its calculation formula is as follows:

[0086] ;

[0087] ;

[0088] ;

[0089] Among them, TP is a true positive example, which means that the BiGRU model predicts that it is a positive class and the actual situation is also a positive class; FP is a false positive example, which means that the BiGRU model predicts that it is a positive class, but the actual situation is a negative class; FN is a false negative example, that is, the BiGRU model predicts that it is a negative class, but the actual situation is a positive class.

[0090] Since the CodeT5+ language model in the embodiment of the present invention covers the code generation and completion tasks during the pre-training process, a more accurate vector representation of the code semantics is obtained after word embedding through the CodeT5+ language model, so that the accuracy of code vulnerability prediction by the BiGRU model is higher. See Table 1. Table 1 shows the experimental results under different embedding methods. After word embedding through the CodeT5+ language model of the present invention, the BiGRU model is input, and the final test indicators are higher than those of CodeBERT and T5 in the related technologies as a whole.

[0091] Table 1 Experimental results of different embedding methods

[0092]

[0093] Then, based on the above three detection indicators, the preliminary code detection results are obtained. If the preliminary code detection result is that there is a vulnerability, it means that there is a vulnerability in a function of the detected code. The position of the vulnerable function can be directly located, that is, the position of the function to which the code with the vulnerability belongs. This is also the granularity of locating code vulnerabilities in related technologies. However, the specific position of the vulnerable code cannot be located. If this detection ends, manual screening will be required later.

[0094] The following method steps of the embodiment of the present invention are proposed to solve the problem that the location of code vulnerabilities in the related art is not accurate enough.

[0095] It can be understood that when the preliminary code detection result indicates that there is no vulnerability, the detection ends; when the preliminary code detection result indicates that there is a vulnerability, the following line-level code vulnerability detection is performed:

[0096] Step 206: When the preliminary code detection result indicates that there is a vulnerability, the attention score of each semantic vector in the code semantic vector sequence corresponding to the function where the vulnerability is located in the bidirectional gated recurrent unit is obtained to obtain a second attention weight.

[0097] The calculation method of the attention score in this step can refer to step 204 and will not be repeated here.

[0098] Step 207: split the code semantic vector sequence corresponding to the function where the vulnerability is located according to the line break to obtain a semantic vector line list.

[0099] Step 208: Based on the semantic vector row list, perform weighted summation on the first attention weight and the second attention weight by row to obtain a row attention score of the code semantic vector sequence corresponding to the function where the vulnerability is located.

[0100] In an embodiment of the present invention, the code semantic vector sequence corresponding to the function where the vulnerability is located is segmented according to line breaks to obtain a semantic vector row list, which includes row labels and their corresponding code semantic vectors. According to the row labels, the first attention weight and the second attention weight corresponding to each row of the code semantic vector are weighted and summed to obtain a row attention score for each row of the code semantic vector in the vulnerable function.

[0101] Step 209: sort the row attention scores obtained above in descending order, and obtain the final code detection result according to the sorting result.

[0102] Specifically, after sorting the row attention scores obtained above in descending order, a set of detection indicators are obtained according to the sorting results, and then the final code detection results are obtained according to these detection indicators.

[0103] Specifically, in the embodiment of the present invention, the detection index can be set according to the actual situation. For example, in this step, the detection index can include Top-10-Accuracy, Effort@20%Recall and Recall@1%LOC.

[0104] Among them, Top-10-Accuracy indicates the percentage of vulnerable functions with row-level vulnerabilities that are ranked in the top 10 rows; Effort@20%Recall indicates the total number of rows that need to be scanned when 20% of the vulnerable rows are scanned out among all vulnerable functions. This indicator can evaluate the cost of row positioning; Recall@1%LOC indicates the proportion of vulnerable rows scanned out among the top 1% of the total number of rows of all vulnerable functions.

[0105] Then, based on these detection indicators, the code lines that contribute most significantly to the function vulnerability prediction results can be selected to obtain the final code detection results: such as the code data and line labels of the line.

[0106] The code detection method of the embodiment of the present invention optimizes the representation effect of the code semantic vector in complex scenarios by introducing the CodeT5+ language model. CodeT5+ covers code generation and completion tasks in the pre-training process, and can more comprehensively capture the semantic information and contextual relationships of the code. The vector representation generated by it is more accurate and can better reflect the potential dependencies, detailed features and logical structures in the code. Combined with the BiGRU model with an attention mechanism, more fine-grained code vulnerability positioning can be performed. In the line-level code vulnerability detection, that is, the fine-grained vulnerability detection process, based on the comprehensive calculation and sorting of the attention weight scores of the CodeT5+ language model and the attention weight scores of the BiGRU model, the code detection granularity can be refined to the code line, thereby accurately locating the location of the code vulnerability. The embodiment of the present invention accurately locates vulnerabilities from function-level vulnerabilities to line-level vulnerabilities, reduces subsequent manual retrieval work, significantly reduces manual screening costs, and improves detection efficiency.

[0107] In addition, among many recurrent neural networks, BiGRU is better at handling long sequence problems than other models, and maintains the stability of the model by alleviating gradient vanishing and gradient exploding phenomena. Moreover, BiGRU can process code sequences in both forward and backward directions. This two-way processing method helps to fully capture the features in the code and further improve the effect of code detection.

[0108] The method provided in the embodiment of the present invention can be applied to electronic devices. Specifically, the electronic device can be: a desktop computer, a portable computer, an intelligent mobile terminal, a server, etc. This is not limited here, and any electronic device that can implement the present invention belongs to the protection scope of the present invention.

[0109] Embodiment 3

[0110] The embodiment of the present invention further provides an electronic device, such as Figure 3As shown, it includes a processor 301, a communication interface 302, a memory 303 and a communication bus 304, wherein the processor 301, the communication interface 302, and the memory 303 communicate with each other through the communication bus 304;

[0111] Memory 303, for storing computer program 305, the computer program is Figure 3 In short, it is referred to as program;

[0112] The processor 301, when used to execute the computer program 305 stored in the memory 303, implements the following steps: preprocessing the code to be detected to obtain a code word block index sequence; embedding the code word block index sequence to obtain the relative position code of each word block in the code word block index sequence, and based on the relative position code, embedding enhances the embedding representation of the code word block index sequence and models it to obtain a code semantic vector sequence; performing vulnerability detection on the code semantic vector sequence to obtain a code detection result.

[0113] The communication bus mentioned in the above electronic device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0114] The communication interface is used for communication between the above electronic device and other devices.

[0115] The memory may include a random access memory (RAM) or a non-volatile memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.

[0116] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0117] An embodiment of the present invention further provides a computer-readable medium on which a computer program is stored. When the computer program is executed by a processor, the method steps of any of the above-mentioned code detection methods are implemented.

[0118] As for the electronic device / computer-readable medium embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0119] It should be noted that the electronic device and computer-readable medium of the embodiments of the present invention are respectively an electronic device and a computer-readable medium to which the above-mentioned code detection method is applied. All embodiments of the above-mentioned code detection method are applicable to the electronic device and computer-readable medium, and can achieve the same or similar beneficial effects.

[0120] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.

[0121] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine different embodiments or examples described in this specification.

[0122] Although the present application is described herein in conjunction with various embodiments, in the process of implementing the claimed application, those skilled in the art may understand and implement other variations of the disclosed embodiments by viewing the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "one" or "an" does not exclude multiple situations. A single processor or other unit may implement several functions listed in a claim. Certain measures are recorded in different dependent claims, but this does not mean that these measures cannot be combined to produce good results.

[0123] The above contents are further detailed descriptions of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, several simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as falling within the protection scope of the present invention.

Claims

1. A code detection method, characterized in that: The method comprises: Preprocess the code to be detected to obtain a code word block index sequence; The code block index sequence is embedded by using the pre-trained CodeT5+ language model to obtain the relative position encoding of each block in the code block index sequence. Based on the relative position encoding, the embedding representation of the code block index sequence is modeled after embedding enhancement to obtain a code semantic vector sequence. Performing vulnerability detection on the code semantic vector sequence to obtain a code detection result; The method of performing embedding enhancement and modeling on the embedding representation of the code word block index sequence based on the relative position encoding to obtain a code semantic vector sequence comprises: Using the attention weight calculation formula , calculating the attention score of each word block in the code word block index sequence in the pre-trained CodeT5+ language model to obtain a first attention weight; Based on the first attention weight, embedding enhancement is performed on the embedding representation of the code word block index sequence to model the code semantic vector sequence; In the attention weight calculation formula, is a query matrix, which represents the information that any word block in the code word block index sequence wants to obtain from other words blocks; is a key matrix used to determine the correlation between any word block in the code word block index sequence and other word blocks, Represents the key matrix The transposed matrix of represents the dimension of the key matrix; is a relative position matrix, used to represent the relative position distance of the word blocks in the code word block index sequence, and the relative position matrix is ​​constructed based on the relative position encoding; represents the first attention weight.

2. The method according to claim 1, characterized in that The performing vulnerability detection on the code semantic vector sequence to obtain a code detection result includes: The bidirectional gated recurrent unit is used to perform function vulnerability detection on the code semantic vector sequence to obtain preliminary code detection results.

3. The method according to claim 2, characterized in that After performing function vulnerability detection on the code semantic vector sequence by using a bidirectional gated loop unit to obtain a preliminary code detection result, the method further includes: When the preliminary code detection result indicates that there is a vulnerability, obtaining an attention score of each semantic vector in the bidirectional gated recurrent unit in a code semantic vector sequence corresponding to the function where the vulnerability is located, and obtaining a second attention weight; The code semantic vector sequence corresponding to the function where the vulnerability is located is split according to the line break character to obtain a semantic vector line list; Based on the semantic vector row list, weighted summing the first attention weight and the second attention weight by row is performed to obtain a row attention score of the code semantic vector sequence corresponding to the function where the vulnerability is located; The row attention scores are sorted in descending order, and the final code detection result is obtained according to the sorting result.

4. The method according to claim 1, characterized in that The CodeT5+ language model is trained in the following way: Obtaining a training code sample, wherein the training code sample includes: a data set containing row-level vulnerabilities and row labels corresponding to row data in the data set; The CodeT5+ language model is trained using the training code samples.

5. The method according to claim 4, characterized in that The using the training code sample to train the CodeT5+ language model includes: Performing word segmentation processing on the training code sample to obtain a sample word block vector; Performing index mapping on the sample chunk vectors according to the vocabulary to obtain a sample chunk index sequence; The sample chunk index sequence is embedded to obtain a relative position code of each chunk in the sample chunk index sequence, and the embedded representation of the sample chunk index sequence is embedded and enhanced based on the relative position code to perform modeling to obtain a sample semantic vector sequence; Using the target loss function, calculating the loss value of the sample semantic vector sequence; The CodeT5+ language model is trained according to the loss value.

6. The method according to claim 5, characterized in that The word segmentation processing of the training code sample to obtain a sample word block vector includes: The training code sample is segmented by using a subword tokenization method to obtain a sample word block vector.

7. An electronic device, characterized in that: The device comprises: one or more processors; A computer readable medium configured to store one or more computer programs, When the one or more computer programs are executed by the one or more processors, the one or more processors implement the code detection method according to any one of claims 1 to 6.

8. A computer readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the code detection method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • C source code vulnerability detection method based on Bert model and BiLSTM

    CN113420296A

  • Vulnerability detection method based on code data stream enhanced large model

    CN118246029A