A method, device, medium, and equipment for detecting code vulnerabilities

By identifying and deleting comment information, adjusting and combining code snippets, and building an evaluation set to adjust the vulnerability detection model, solving the problems of low vulnerability detection efficiency and high false alarm rate in the existing technology, improving detection accuracy and repair effect.

CN119939608BActive Publication Date: 2025-06-20ZHEJIANG LAB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510423675.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-06-20
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

The prior art is inefficient in code vulnerability detection, has a high false positive rate, and large models are easily affected by annotation information during semantic analysis, resulting in distortion of detection results.

Method used

By obtaining the sample code set, identifying and deleting annotation information, adjusting the desensitization code to generate enhancement code, combining code snippets to form composite code, and building evaluation sets through these composite codes to adjust the vulnerability detection model.

Benefits of technology

It improves the identification accuracy and code repair effect of the vulnerability detection model, reduces the false positive rate, and makes the model output consistent with its actual vulnerability recognition capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939608B_ABST
    Figure CN119939608B_ABST
Patent Text Reader

Abstract

The present application discloses a code vulnerability detection method, device, medium and equipment. After obtaining a sample code set, it can identify and delete the annotation information in each sample code included in the sample code set to obtain each desensitized code. Then, according to a preset adjustment strategy, each desensitized code is adjusted to obtain at least one enhanced code with the same code identifier as the desensitized code. Subsequently, by combining each code segment belonging to the same code identifier, multiple composite codes corresponding to one code identifier can be obtained. Furthermore, through these composite codes, a test set is constructed to adjust the vulnerability detection model through the test set, and code vulnerability detection is performed through the adjusted vulnerability detection model. Thus, in subsequent actual applications, the recognition accuracy of the vulnerability detection model can be significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of computer technology and artificial intelligence, and particularly to a method, apparatus, medium, and device for code vulnerability detection. Background Art

[0002] As the complexity of software systems continues to increase, detecting vulnerabilities in software has become an important means to ensure the normal operation and security of software. Traditional vulnerability detection methods mainly rely on manual analysis and static code analysis tools, but these methods are inefficient when dealing with software with large-scale code and often have a high false positive rate.

[0003] With the rapid development of artificial intelligence technology in recent years, using large models for vulnerability detection has become a new idea. By learning a large number of software samples with vulnerabilities, large models can understand the semantics of long code and automatically identify potential vulnerabilities, and through semantic analysis, the false positive rate can be effectively reduced, and an interpretable analysis report can be given.

[0004] However, there are also certain problems with using large models to perform vulnerability detection. For example, since the software code in the sample set contains a large number of comment statements and declarations of classes, functions, variables, etc. with explanatory meanings, this easily leaks the semantics of the software code to the large model, causing the large model to be able to know whether there are vulnerabilities in the code only through these comments without actually identifying the code itself, resulting in distorted vulnerability detection results. Moreover, the length of some code fragments in the current sample set is insufficient and lacks a complete context, thus significantly affecting the accurate positioning of vulnerabilities by the large model.

[0005] In addition, current large models are limited to judging whether there are vulnerabilities in software and the types of vulnerabilities, and cannot effectively repair the vulnerabilities existing in the software. Moreover, current large models may detect vulnerabilities only through simple feature matching methods rather than through semantic understanding methods, resulting in insufficient generalization ability. Summary of the Invention

[0006] Embodiments of this application provide a method, apparatus, medium, and device for code vulnerability detection to partially solve the above problems existing in the prior art.

[0007] This application adopts the following technical solutions:

[0008] Embodiments of this application provide a method for code vulnerability detection, including:

[0009] Obtain a sample code set, where each sample code is included in the sample code set. For each sample code, the sample code corresponds to a code identifier, and each sample code includes various code segments. The code categories of the various code segments include: vulnerability code, the repair code corresponding to the vulnerability code, and the context code of the vulnerability code;

[0010] Identify and delete the comment information in each sample code to obtain a desensitized code corresponding to each sample code;

[0011] For each desensitized code, adjust the desensitized code according to a preset adjustment strategy to obtain at least one enhanced code corresponding to the desensitized code, where the at least one enhanced code has the same code identifier as the desensitized code;

[0012] For each code identifier, combine the code segments belonging to the code identifier to obtain multiple composite codes corresponding to the code identifier, and save them in a preset database. The code categories of the code segments included in one composite code are all different;

[0013] According to the composite codes corresponding to each code identifier saved in the preset database, construct an evaluation set, and through the evaluation set, adjust a preset vulnerability detection model to perform code vulnerability detection through the adjusted vulnerability detection model.

[0014] Optionally, for each desensitized code, adjust the desensitized code according to a preset adjustment strategy to obtain at least one enhanced code corresponding to the desensitized code, specifically including:

[0015] For each desensitized code, adjust the control flow of the desensitized code to obtain at least one enhanced code corresponding to the desensitized code. The control flow is used to represent the execution order and manner of the code statements in the desensitized code.

[0016] Optionally, for each desensitized code, adjust the desensitized code according to a preset adjustment strategy to obtain at least one enhanced code corresponding to the desensitized code, specifically including:

[0017] For each desensitized code, add redundant code to the desensitized code to obtain at least one enhanced code corresponding to the desensitized code.

[0018] Optionally, according to the composite codes corresponding to each code identifier saved in the preset database, construct an evaluation set, specifically including:

[0019] For each code identifier, select one composite code from the multiple composite codes corresponding to this code identifier as the question stem code, select at least some of the composite codes from the other composite codes corresponding to this code identifier, and select at least one composite code from the composite codes corresponding to other code identifiers;

[0020] According to the question stem code, at least some of the composite codes selected from the other composite codes corresponding to this code identifier, at least one composite code selected from the composite codes corresponding to other code identifiers, and a preset prompt word template, construct a first evaluation sample, and based on the first evaluation sample, construct an evaluation set, where the first evaluation sample is used to enable the vulnerability detection model to identify codes that are semantically irrelevant to the question stem code.

[0021] Optionally, construct an evaluation set according to the composite codes corresponding to each code identifier saved in the preset database, specifically including:

[0022] For each code identifier, select one composite code from the multiple composite codes corresponding to this code identifier, and use the vulnerability code contained in this composite code as the question stem code. According to this code identifier, determine the repair code corresponding to the question stem code, obtain at least some of the vulnerability codes contained in the other composite codes corresponding to this code identifier, and obtain the code fragments contained in the composite codes corresponding to other code identifiers;

[0023] According to the question stem code, the repair code corresponding to the question stem code, at least some of the vulnerability codes obtained from the other composite codes corresponding to this code identifier, the code fragments contained in the composite codes corresponding to other code identifiers, and a preset prompt word template, construct a second evaluation sample, and based on the second evaluation sample, construct an evaluation set, where the second evaluation sample is used to enable the vulnerability detection model to select the repair code for repairing the question stem code.

[0024] Optionally, for each composite code, the preset database also saves the line numbers of the code corresponding to the vulnerability code contained in this composite code;

[0025] According to the composite codes corresponding to each code identifier saved in the preset database, construct an evaluation set, specifically including:

[0026] For each code identifier, select the vulnerability code contained in one composite code from the multiple composite codes corresponding to this code identifier as the question stem code, read out the line number of the code corresponding to the question stem code from the preset database, and randomly generate several line numbers for the question stem code;

[0027] Generate a third evaluation sample according to the stem code, the line number corresponding to the stem code read from the preset database, several randomly generated line numbers for the stem code, and a preset prompt word template, so as to construct an evaluation set according to the third evaluation sample. The third evaluation sample is used to enable the vulnerability detection model to determine the line number of the stem code.

[0028] Optionally, for each composite code, the preset database also stores the vulnerability type information corresponding to the vulnerability code included in the composite code.

[0029] Construct an evaluation set according to the composite codes corresponding to each code identifier stored in the preset database, specifically including:

[0030] For each code identifier, select the vulnerability code included in one composite code corresponding to the code identifier as the stem code, read the vulnerability type information corresponding to the stem code from the preset database, and randomly obtain the vulnerability type information of several other vulnerability types.

[0031] Generate a fourth evaluation sample according to the stem code, the vulnerability type information corresponding to the stem code read from the preset database, the randomly obtained vulnerability type information of several other vulnerability types, and a preset prompt word template, so as to construct an evaluation set according to the fourth evaluation sample. The fourth evaluation sample is used to enable the vulnerability detection model to identify the vulnerability type corresponding to the stem code.

[0032] An embodiment of the present application provides a code vulnerability detection device, including:

[0033] An acquisition module, configured to acquire a sample code set, where each sample code is included in the sample code set. For each sample code, the sample code corresponds to a code identifier, and each code is included in the sample code. The code categories of the code segments include: vulnerability code, the repair code corresponding to the vulnerability code, and the context code of the vulnerability code.

[0034] An identification module, configured to identify and delete the annotation information in each sample code to obtain a desensitized code corresponding to each sample code.

[0035] An adjustment module, configured to adjust each desensitized code according to a preset adjustment strategy to obtain at least one enhanced code corresponding to the desensitized code, where the at least one enhanced code has the same code identifier as the desensitized code.

[0036] Combination module, which is used to combine each code segment belonging to a code identifier for each code identifier, obtain multiple composite codes corresponding to the code identifier, and save them in a preset database. Each code segment included in a composite code has a different code category;

[0037] Construction module, which is used to construct an evaluation set according to the composite codes corresponding to each code identifier saved in the preset database, and adjust a preset vulnerability detection model through the evaluation set, so as to perform code vulnerability detection through the adjusted vulnerability detection model.

[0038] An embodiment of the present application provides a computer-readable storage medium. The storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned code vulnerability detection method is implemented.

[0039] An embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the code vulnerability detection method is implemented.

[0040] At least one of the above technical solutions adopted in the embodiments of the present application can achieve the following beneficial effects:

[0041] For the code vulnerability detection method provided by the embodiment of the present application, after obtaining a sample code set, the annotation information in each sample code included in the sample code set can be identified and deleted to obtain each desensitized code. Then, according to a preset adjustment strategy, each desensitized code is adjusted to obtain at least one enhanced code with the same code identifier as the desensitized code. Subsequently, by combining each code segment belonging to the same code identifier, multiple composite codes corresponding to one code identifier can be obtained. Furthermore, through these composite codes, an evaluation set is constructed to adjust the vulnerability detection model through the evaluation set, and code vulnerability detection is performed through the adjusted vulnerability detection model.

[0042] It can be seen from the above method that since the annotation information in each sample code included in the sample code set can be deleted after obtaining the sample code set, the evaluation set constructed subsequently will not disclose the software code semantics to the vulnerability detection model, so that the vulnerability detection model can output results that conform to its actual vulnerability identification ability. Moreover, by adjusting the desensitized code, multiple composite codes can be obtained. Through these composite codes, a variety of evaluation samples can be obtained, so that the ability of the vulnerability detection model to identify code vulnerabilities can be enhanced from multiple angles. Furthermore, in subsequent actual applications, the identification accuracy of the vulnerability detection model can be significantly improved. Description of the Drawings

[0043] The accompanying drawings described herein are used to provide a further understanding of this specification and form a part of this application. The illustrative embodiments of this application and their descriptions are used to explain this specification and do not unduly limit this application. In the drawings:

[0044] Figure 1 It is a schematic flowchart of a code vulnerability detection method provided by an embodiment of this application;

[0045] Figure 2 It is a schematic diagram of a storage structure of sample code provided by this specification;

[0046] Figure 3 It is a schematic diagram of a storage structure of a code segment provided by this specification;

[0047] Figure 4 It is a schematic diagram of a storage structure adopted for storing each code segment in the desensitized code provided by this specification;

[0048] Figure 5 It is a schematic diagram of a storage structure of the enhanced code segment obtained after adjustment provided by this specification;

[0049] Figure 6 It is a schematic diagram of a storage structure of a composite code provided by this specification;

[0050] Figure 7 It is a schematic diagram of a storage structure of a test set provided by this specification;

[0051] Figure 8 It is a schematic diagram of a device for code vulnerability detection provided by an embodiment of this application;

[0052] Figure 9 It is provided by an embodiment of this application corresponding to Figure 1 a schematic diagram of the structure of an electronic device. Detailed implementation manners

[0053] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this specification will be clearly and completely described below in conjunction with specific embodiments of this application and the corresponding accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.

[0054] The following will describe in detail the technical solutions provided by each embodiment of this application in conjunction with the accompanying drawings.

[0055] Figure 1The flowchart of a code vulnerability detection method provided by an embodiment of this application includes the following steps:

[0056] S101: Obtain a sample code set.

[0057] In this specification, in order to enable the vulnerability detection model to accurately identify and repair the vulnerabilities contained in software code in actual applications, therefore, before using the vulnerability detection model, it is necessary to adjust it through the constructed evaluation set, so that the vulnerability detection model after being adjusted by the evaluation set has strong software vulnerability identification and repair capabilities.

[0058] Before that, it is necessary to first construct an evaluation set for adjusting the vulnerability detection model. Therefore, the code vulnerability detection method provided in this specification can actually be divided into a data set construction stage, a model adjustment stage, and a vulnerability detection stage. The focus of this specification is mainly concentrated on the data set construction stage because the evaluation set constructed by the method provided in this specification contains a variety of evaluation samples. These evaluation samples can adjust the vulnerability detection model from multiple angles, enabling the vulnerability detection model to learn "knowledge" about software code vulnerability detection and repair from different angles, so as to exert better code vulnerability detection and repair capabilities in actual applications and provide more convenience for users.

[0059] In the process of constructing the evaluation set, it is necessary to first obtain a sample code set. Among them, the sample code set contains multiple sample codes. For any one sample code, the sample code can contain various code segments. Among them, the code categories of the code segments can include: vulnerable code (i.e., code with vulnerabilities), the repaired code corresponding to the vulnerable code (i.e., code without vulnerabilities after repairing the vulnerable code), and the context code of the vulnerable code (i.e., code used to connect the vulnerable code).

[0060] The above sample code set can be obtained from an existing code set. For the obtained sample code set, the vulnerable codes and repaired codes in each sample code contained therein are marked in advance, and they can be stored in a preset database in a certain manner, such as Figure 2 shown.

[0061] Figure 2 The schematic diagram of a sample code storage structure provided by this specification.

[0062] From Figure 2 it can be seen that for a sample code, the sample code corresponds to a unique code identifier, and the code identifier can be obtained by calculating the hash value. The sample code can contain vulnerable code, repaired code, upper context, and lower context.

[0063] Among them, the upper context can be understood as the code connected before the vulnerability code or the repair code, while the lower context can be understood as the code connected after the vulnerability code or the repair code.

[0064] In Figure 2 the storage structure, the line numbers of the vulnerability code in the sample code and the line numbers of the repair code in the sample code are also recorded. In addition, Figure 2 the programming language corresponding to the sample code is also recorded in the storage structure. The reason for recording the programming language is that when obtaining the composite code in the subsequent process, the operation methods for the codes of different programming languages are different.

[0065] Figure 2 The vulnerability types corresponding to the vulnerability codes included in the sample code are also recorded in the storage structure, such as out-of-bounds writing (which means that during the execution of the code program, the position where data is attempted to be written into the memory exceeds the memory area allocated to it), out-of-bounds reading (which means that the code program attempts to access data outside the memory area allocated to it during execution), improper input validation (the code program fails to fully check whether the data provided by the user or an external system meets the expected format, type, range, etc. requirements), etc. This is helpful for the subsequent vulnerability detection model to identify the vulnerability types.

[0066] In this specification, the vulnerability detection model for adjustment and for vulnerability detection can be a large language model. In actual applications, the code to be detected can be input into the vulnerability detection model together with the prompt statement. The vulnerability detection model can identify the user's intention through the received prompt statement, thereby completing the vulnerability detection task.

[0067] In addition, the execution subject of the code vulnerability detection method provided in this specification can be a terminal device such as a desktop computer or a laptop computer, a server, or a client installed in the terminal device. For the convenience of description, below, only the case where the server is the execution subject will be taken as an example to describe in detail the code vulnerability detection method provided in this specification.

[0068] After the server obtains the above sample code set, it can perform a series of processes on each sample code included in the sample code set. Specifically, the server can uniformly convert each sample code included in the sample code set into a specified encoding format, thereby obtaining the converted sample code set. The reason for converting to a unified encoding format is to avoid exceptions during the parsing process. Considering factors such as universality and extensiveness, each sample code in the sample code set can be uniformly converted into the UTF-8 encoding format.

[0069] Subsequently, the server can perform code location on each sample code in the sample code set to separately extract vulnerability code, repair code, and context code therefrom. Among them, for each sample code in the sample code set, the server can locate the interval of the vulnerability code in the sample code according to the line number of the vulnerability code, so as to extract the vulnerability code located in this interval. Similarly, the server can locate the interval of the repair code in the sample code according to the line number of the repair code, so as to extract the repair code located in this interval.

[0070] For the context code, the server can determine the upper context segment and the lower context segment according to the line number of the vulnerability code or the line number of the repair code. Specifically, the following formulas can be used to determine the line number of the upper context segment and the line number of the lower context segment respectively:

[0071] The line number of the upper context segment is: [0, z1 = min(x1, y1)]

[0072] The line number of the lower context segment is: [z2 = max(x2, y2), -1], where -1 represents the last line.

[0073] After obtaining the above line numbers, the server can extract the upper context segment PC = S[0, z1], and extract the lower context segment as SC = S[z2, -1]. Among them, PC represents the upper context segment, and SC represents the lower context segment.

[0074] The server can extract the vulnerability code, repair code, and context code from the sample code respectively in the above manner. After that, the server can store these code snippets in a preset database according to a certain storage structure, such as Figure 3 shown.

[0075] Figure 3 is a schematic diagram of a storage structure of a code snippet provided in this specification.

[0076] For any code snippet, the server can calculate the hash value according to the code information contained in the code snippet to obtain the identifier of the code snippet. And Figure 3 the code identifier of the sample code to which the code snippet belongs is also recorded in the storage structure shown. Figure 3 The code part in Figure 3 is used to record the code snippet itself. And

[0077] S102: Identify and delete the comment information in each sample code to obtain the desensitized code corresponding to each sample code.

[0078] Since the current code contains a large amount of annotation information, this annotation information will disclose the semantics of the software code, causing the vulnerability detection model to no longer perform semantic analysis on the code itself, but rather identify vulnerabilities solely based on the annotation information, resulting in distorted detection results.

[0079] Therefore, in this specification, after obtaining the above sample code set, the server can identify the annotation information contained in each sample code in the sample code set to obtain the desensitized code corresponding to each sample code. Among them, the server can use a preset regular expression to identify and delete the annotation information in the sample code.

[0080] For example, for C / C++, comments are divided into single-line comments (starting with / / ) and multi-line comments (starting with / * and ending with * / ). Among them, the regular expression for deleting C / C++ single-line comments is:

[0081] re.sub(r' / / .*', '', code)

[0082] In this regular expression, the start symbol for matching comments is / / ; and the symbol.* means matching all characters after / / in that line.

[0083] The regular expression for deleting C / C++ multi-line comments is:

[0084] re.sub(r' / \*.*?\* / ', '', code, flags=re.DOTALL)

[0085] In this regular expression, / \\*: this part matches the start marker / * of the comment;.*? matches any character (non-greedy mode) until * / is encountered; \* / matches the end * / of the comment; flags=re.DOTALL makes. able to match newline characters, thus supporting multi-line comments across lines.

[0086] For Python, comments are divided into single-line comments (starting with #) and multi-line comments (starting and ending with single quotes or double quotes, including multi-line strings).

[0087] Among them, the regular expression for deleting Python multi-line comments is:

[0088] re.compile(

[0089] r"""

[0090] (\"\"\"[^\"\"\"]*\"\"\" | # Match multi-line double-quoted strings

[0091] \'\'\'[^\'\'\']*\'\'\' | # Matches multi-line single-quoted strings

[0092] \"[^\"]*\" | # Matches single-line double-quoted strings

[0093] \'[^\']*\' ) # Matches single-line single-quoted strings

[0094] """, re.VERBOSE )

[0096] Among them, the regular expression for deleting single-line Python comments is:

[0097] re.compile(r"#.*$", re.MULTILINE)

[0098] In this regular expression, re.MULTILINE modifies the behavior of ^ and $ so that they match the beginning and end of each line respectively, rather than the beginning and end of the entire string.

[0099] And in order to further enhance the vulnerability detection ability and vulnerability repair ability of the vulnerability detection model, in this specification, in addition to being able to delete the comments in the above sample code, the server can also modify some identifiers in the sample code on the basis of the sample code with comments deleted.

[0100] Among them, the so-called partial identifiers here can be understood as information that does not change the original semantic features of the code. In this specification, the so-called identifiers can refer to, for example, class names, variable names, and function names.

[0101] Therefore, after the server identifies and deletes the comment information in each sample code, it obtains the initial desensitized code corresponding to each sample code. Then, the server can identify the non-critical identifiers included in the initial desensitized code corresponding to each sample code and modify these non-critical identifiers, thereby obtaining the desensitized code corresponding to each sample code. Among them, the non-critical identifiers mentioned here include the class names of non-critical classes, the variable names of non-critical variables, and the function names of non-built-in functions. Modifying the non-critical identifiers does not change the original semantic features of the code.

[0102] Therefore, when the server extracts identifiers such as class names, variable names, and function names from the above initial desensitized code, it will detect whether these identifiers are critical identifiers. The so-called critical identifiers are those that will change the original semantic features of the code after being modified. Therefore, the modification of some identifiers in the initial desensitized code in this specification only plays a certain role in obfuscation, rather than completely changing the original semantic features of the initial desensitized code.

[0103] Further, for different programming languages, the server can use different regular expressions to modify the non-critical identifiers in the above initial desensitization code.

[0104] For example, C / C++ belongs to statically typed languages and strongly typed languages. At compile time, the data type of a variable must be determined, and its rules are relatively clear. Class names, function names, and variable names can all be regarded as a type of variable.

[0105] Therefore, the regular expressions used to match C / C++ class names, function names, and variable names are as follows:

[0106] var_regex = re.compile(r"\b(int|double|float|char|bool|void|short|long|size_t|class)\s+([a-zA-Z_][a-zA-Z0-9_]*)\b").

[0107] The above expression matches common variable types (such as int, double, float, etc.) and captures the variable name.

[0108] Among them, \b: This is a marker for a word boundary. It does not match any actual characters but represents a position - a position where the characters before and after it are not part of the same word. Specifically, \b is usually used to ensure that the match is for the entire word rather than just a part of a word. For example, when matching variable names, using \b can prevent misinterpreting a part of a variable name as another variable name.

[0109] (int|double|float|char|bool|void|short|long|size_t|class): This is a defined capture group that contains multiple keywords separated by the pipe character |. The listed ones are common data types and the keyword class in C / C++. The vertical bar | in the regular expression means "or", so this capture group will match any of the listed data types or the class keyword.

[0110] \s+: This represents one or more whitespace characters. Whitespace characters include spaces, tabs, etc. The plus sign + here means that the previous element (here, whitespace characters) appears at least once.

[0111] ([a-zA-Z_][a-zA-Z0-9_]*): This is another capturing group, used to match variable or class names. The first character must be a letter (uppercase or lowercase) or an underscore, which is represented by [a-zA-Z_]. The following part can contain zero or more letters, numbers, or underscores, which is represented by [a-zA-Z0-9_]*. The asterisk * means that the previous element (in this case, a set of letters, numbers, or underscores) can appear zero or more times.

[0112] For another example, Python is a dynamic language that determines data types at runtime. No type declaration is required before using a variable. For example, in Python, if variable a=1, then the type of a is integer. If a="abc", the type of a is string. Therefore, it is relatively responsible when implementing obfuscation.

[0113] To this end, the server can use a variety of regular expressions to match class, function, and variable names. In Python, when defining classes and functions, the iconic keywords class and def are prepended, so the C / C++ matching mode can be used; variable declarations are more obscure, and given the characteristics of the Python programming language, this manual treats the left side of the assignment statement as a variable.

[0114] Therefore, the regular expression used to match Python classes and functions is as follows:

[0115] class_func_regex = re.compile(r"\b(class|def)\s+([a-zA-Z_][a-zA-Z0-9_]*)")

[0116] Among them, \b: This is a word boundary marker, which ensures that the match is an independent word rather than a part of a word. For example, when matching class or def, using \b can avoid mistaking these keywords for parts of other words.

[0117] (class|def): This is a capture group, which contains two options separated by a vertical bar |, indicating an "or" relationship. Here it means matching one of the two keywords class or def. Class is used to define a class, while def is used to define a function.

[0118] \s+: This represents one or more whitespace characters. Whitespace characters include space, tab, etc. The plus sign + here means that the preceding element (in this case, whitespace characters) occurs at least once.

[0119] ([a-zA-Z_][a-zA-Z0-9_]*): This is another capture group, used to match class names or function names. The specific rules are as follows:

[0120] [a-zA-Z_]: It means that the first character must be a letter (either uppercase or lowercase) or an underscore.

[0121] [a-zA-Z0-9_]*: It means that the following characters can contain zero or more letters, digits, or underscores. The asterisk * means that the previous element (here, the set of letters, digits, or underscores) can appear zero or more times.

[0122] In particular, Python supports multi-value returns. Therefore, the regular expression needs to match all the results on the left side of the assignment statement. The regular expression for matching multi-variable assignment is as follows:

[0123] multi_assign_regex = re.compile(r"\b([a-zA-Z_][a-zA-Z0-9_]*)\s*,\s*([a-zA-Z_][a-zA-Z0-9_]*)(?:\s*,\s*([a-zA-Z_][a-zA-Z0-9_]*))*\s*=\s*([^=\n]+)")

[0124] For the explanations of each symbol in this regular expression, refer to the symbol explanations of the above regular expression, which will not be elaborated here in detail.

[0125] In addition, since the server can extract the vulnerable code, repair code, and context code included in the sample code and store them according to the storage structure as Figure 3 shown, then when the server performs the modification operation of the above non-critical identifiers, it needs to first merge the vulnerable code, repair code, and context code corresponding to the same code identifier before it can perform the modification of the non-critical identifiers. This is mainly to ensure that the names of each code snippet are consistent during obfuscation.

[0126] On this basis, the server can store each code snippet included in the desensitized code in a preset database, as Figure 4 shown.

[0127] Figure 4 It is a schematic diagram of the storage structure adopted for storing each code snippet in the desensitized code provided in this specification.

[0128] From Figure 4It can be seen that a code snippet refers to a code snippet after modification of non-critical identifiers (i.e., desensitization). The server can calculate the hash value of the actual code snippet after modification of non-critical identifiers, so as to obtain the code snippet identifier of the code snippet, and the code identifier refers to the identifier information of the sample code to which the code snippet belongs. The code category is used to indicate whether the code snippet is vulnerable code, repaired code, upper context paragraph or lower context paragraph. The code is the code snippet itself after desensitization.

[0129] S103: For each desensitized code, according to a preset adjustment strategy, adjust the desensitized code to obtain at least one enhanced code corresponding to the desensitized code.

[0130] In order to further enhance the vulnerability recognition ability and repair ability of the vulnerability detection model, after obtaining the above-mentioned desensitized codes, the server can adjust each desensitized code according to a preset adjustment strategy, so as to obtain each enhanced code.

[0131] Among them, this specification mainly introduces two adjustment strategies. One is to adjust the control flow in the code. The so-called control flow is used to represent the execution order and manner of code statements in the code. Then, adjusting the control flow means changing the execution order and manner of code statements in the above desensitized code. And since for a desensitized code, its control flow may involve multiple ones, after the server adjusts the control flow in the desensitized code, at least one enhanced code can be obtained.

[0132] In this specification, the server can use a large language model to implement the adjustment of the control flow in the above desensitized code. For example, the server can adopt the following prompt word template:

[0133] <lang>{programming_language}< / lang> \n <code>{code_snippet}< / code> ,\n <code>The code in the label is <lang>Written in the programming language, it is required to adjust the control flow in the code (if any) or replace it with an equivalent expression without changing the semantics and affecting the original logic execution, such as: changing the conditional branch if(x>0) to if(!(x<=0)), replacing a+b with b+a or replacing ++i in the loop with i+=1. Note that only the modified code is output.

[0134] {programming_language} and {code_snippet} are placeholders, which are filled in when generating the large language model prompt.

[0135] The server can use the above-mentioned prompt word template to input the desensitized code into the large language model. The large language model will perform intent recognition based on the content of the prompt word template, and then adjust the control flow of the obtained desensitized code to obtain at least one enhanced code and output it.

[0136] In addition, the server may also adjust the desensitized code by means of code expansion. The so-called code expansion is to add or subtract redundant codes (ie, increase redundant noise) in the desensitized code, thereby obtaining at least one enhanced code.

[0137] The server can also use a large language model to complete the code expansion process. For example, the server can use the following prompt word template:

[0138] <lang>{programming_language}< / lang> \n <code>{code_snippet}< / code> ,\n <code>The code in the label is <lang>Written in the programming language, it is required that without changing the semantics and without affecting the original execution logic, in <code>Add 10 - 20 lines of redundant code. Note that only the modified code is to be output.

[0139] Regarding the prompt words in the prompt word template, they have been explained in the above content and will not be elaborated in detail here.

[0140] In this specification, the server can obtain at least one enhanced code by solely adopting the method of adjusting the control flow, or can also obtain at least one enhanced code by solely adopting the method of code inflation. Of course, these two methods can be used simultaneously, that is, not only adjusting the control flow of the desensitized code but also performing code inflation on the desensitized code. As for whether to perform control flow adjustment first or code inflation first, it can be determined according to actual requirements.

[0141] In addition, in practical applications, there can be other adjustment methods to obtain at least one enhanced code, which will not be elaborated in detail here.

[0142] After the server obtains the enhanced code corresponding to each desensitized code through the above method, it can store the obtained enhanced code. Since the server can store each code fragment separately in a preset database in advance, the enhanced code can also be obtained for each code fragment. That is, for a code fragment, the server can use the above method to obtain at least one enhanced code corresponding to that code fragment.

[0143] After obtaining the enhanced code for these code fragments, the server can store these enhanced codes according to a certain data storage structure, such as Figure 5 shown.

[0144] Figure 5 It is a schematic diagram of the storage structure of the enhanced code fragment obtained after adjustment provided in this specification.

[0145] In Figure 5 the code fragment identifier refers to the identifier of the code fragment corresponding to the enhanced code fragment, which can be obtained by the server calculating the hash value. The code identifier refers to the identifier information of the sample code to which the enhanced code fragment belongs. The code refers to the enhanced code fragment itself. The code category is used to indicate whether the enhanced code fragment specifically corresponds to a vulnerability code, a repair code, or a context code (that is, whether it is an enhanced code fragment obtained by adjusting a vulnerability code, a repair code, or a context code).

[0146] Figure 5 The adjustment strategy in it is used to record the adjustment strategy adopted to obtain the enhanced code snippet. Among them, CFP means that the enhanced code snippet is obtained only by adopting the adjustment strategy of adjusting the control flow; Bloat means that the enhanced code snippet is obtained only by adopting the adjustment strategy of code bloat; All means that both the adjustment strategy of adjusting the control flow and the adjustment strategy of code bloat are adopted; None means that no adjustment strategy is adopted.

[0147] Therefore, in actual applications, some of the code in the enhanced code snippet finally obtained by the server only undergoes desensitization processing but is not adjusted by any adjustment strategy.

[0148] S104: For each code identifier, combine the code snippets belonging to the code identifier to obtain multiple composite codes corresponding to the code identifier, and save them in a preset database.

[0149] After obtaining the above enhanced codes, the server can combine the code snippets in each enhanced code to obtain multiple composite codes.

[0150] During the combination process, the server is not limited to only combining the code snippets in the enhanced codes. It can also combine the code snippets included in the enhanced codes and the original sample codes (as mentioned above, the desensitized codes finally obtained by the server also include codes that only undergo desensitization processing but are not adjusted by the adjustment strategy. Therefore, the enhanced codes here actually include codes that only undergo desensitization processing and codes that are adjusted by the adjustment strategy) to obtain multiple composite codes.

[0151] In the specific combination process, the server can actually combine a code snippet in a code (including enhanced code and sample code) with one or more code snippets included in other codes to obtain a composite code. Among them, during the entire combination process, it is necessary to ensure that the code snippets included in the codes corresponding to the same code identifier are combined. Moreover, it is necessary to ensure that the code categories of the code snippets included in a composite code are all different. That is to say, there will not be two or more vulnerability codes, repair codes, or context codes in a composite code.

[0152] Furthermore, during the process of combining code snippets, the server can combine with the vulnerability code as the core and combine with the repair code as the core. Among them, combining with the vulnerability code as the core means combining the adjusted vulnerability code with the context code (including adjusted and unadjusted context codes), and combining the unadjusted vulnerability code with the context code (also including adjusted and unadjusted context codes).

[0153] Combining with the vulnerability code as the core can result in the following 16 cases:

[0154] CFP(PC / SC)+None(B), CFP(PC / SC)+CFP(B), CFP(PC / SC)+Bloat(B),

[0155] CFP(PC / SC)+All(B);

[0156] Bloat(PC / SC)+None(B), Bloat(PC / SC)+CFP(B), Bloat(PC / SC)+Bloat(B),

[0157] Bloat(PC / SC)+All(B);

[0158] All(PC / SC)+None(B), All(PC / SC)+CFP(B), All(PC / SC)+Bloat(B), All(PC / SC)+All(B);

[0159] None(PC / SC)+None(B), None(PC / SC)+CFP(B), None(PC / SC)+Bloat(B),

[0160] None(PC / SC)+All(B).

[0161] Among them, CFP means only adopting the adjustment strategy of adjusting the control flow to adjust the code segment, Bloat means only adopting the adjustment strategy of code bloat to adjust the code segment; All means adopting both the adjustment strategy of adjusting the control flow and the adjustment strategy of code bloat; None means not adopting any adjustment strategy.

[0162] PC represents the upper context segment, SC represents the lower context segment, and B represents the vulnerability code.

[0163] Taking the combination of CFP(PC / SC)+None(B) as an example, it means the composite code obtained by combining the upper and lower context segments adjusted by the adjustment strategy of adjusting the control flow with the vulnerability code that has not been adjusted by any adjustment strategy. The other combination forms will not be elaborated in detail here.

[0164] Similarly, combining with the repair code as the core on the server can also result in the following 16 cases:

[0165] CFP(PC / SC)+None(G), CFP(PC / SC)+CFP(G), CFP(PC / SC)+Bloat(G),

[0166] CFP(PC / SC)+All(G);

[0167] Bloat(PC / SC)+None(G), Bloat(PC / SC)+CFP(G), Bloat(PC / SC)+Bloat(G),

[0168] Bloat(PC / SC)+All(G);

[0169] All(PC / SC)+None(G), All(PC / SC)+CFP(G), All(PC / SC)+Bloat(G), All(PC / SC)+All(G);

[0170] None(PC / SC)+None(G), None(PC / SC)+CFP(G), None(PC / SC)+Bloat(G),

[0171] None(PC / SC)+All(G).

[0172] Among them, PC, SC, CFP, Bloat, All, and None have all been explained and will not be elaborated further, while G represents the repair code.

[0173] Taking the combination of Bloat(PC / SC)+None(G) as an example, it represents a composite code obtained by combining the upper and lower segments of the context with code bloat and the repair code without any adjustment.

[0174] It should be noted that in order to improve the vulnerability recognition ability of the vulnerability detection model, the server can further count the line number range of the vulnerable code in the composite code. In this way, during the subsequent adjustment process of the vulnerability detection model, the vulnerability detection model can identify the line numbers of the vulnerable code contained in the input code and adjust the vulnerability detection model according to the actual line numbers of the vulnerable code.

[0175] For this reason, in this specification, the server can count the line number range of the vulnerable code. Specifically, the following function can be used for counting:

[0176] Start line of vulnerable code: len(PC)

[0177] End line of vulnerable code: len(PC) + len(Bad)

[0178] Among them, the len() function is used to count the line numbers.

[0179] In addition, the server can also mark line numbers for the composite code. Among them, according to different programming languages, C / C++ adds comments at the end of each line of code: / / Line x; for Python, comments are added at the end of each line: # Line x.

[0180] Furthermore, the server can store the obtained composite code according to a certain storage structure, as Figure 6 shown.

[0181] Figure 6 is a schematic diagram of a storage structure of a composite code provided in this specification.

[0182] In Figure 6 , the composite code identifier, code identifier, and the description of the code in Figure 6 have all been explained. For the code type in Figure 6 , from the above composite code, it can be seen that if it is combined with the vulnerability code as the core, then the composite code contains the vulnerability code; if it is combined with the repair code as the core, then the composite code does not contain the repair code.

[0183] Therefore, the code type is used to indicate whether the composite code is combined with the vulnerability code as the core or the repair code as the core. When the code type is Bad, it means that the composite code contains the vulnerability code; when the code type is Good, it means that the composite code does not contain the vulnerability code.

[0184] Correspondingly, for the vulnerability type in Figure 6 , the server can query the specific vulnerability type (such as the out-of-bounds write, out-of-bounds read, improper input validation, etc. mentioned above) from the original sample code set according to the code identifier. If the composite code does not contain the vulnerability code, then the server will not be able to query the corresponding vulnerability type from the original sample code set through the code identifier. At this time, the vulnerability type of the composite code is None.

[0185] Furthermore, the same is true for the vulnerability line number in Figure 6 . If the composite code contains the vulnerability code, the line number where the vulnerability code is located is recorded here; if the composite code does not contain the vulnerability code, the line number where the vulnerability code is located is not recorded here.

[0186] Figure 6 The "list of corresponding identifiers of the referenced code fragments" in Figure 6 means that since the composite code contains enhanced code fragments, the identifiers corresponding to the enhanced code fragments can be recorded in the storage structure shown in Figure 6 The storage structure in

[0187] Figure 6 The enhancement type in

[0188] It should be noted that the reason for recording various identifiers in the storage structure shown in Figure 6 is mainly to indicate the sources of the various code segments (including the enhanced code segments) included in the composite code, so as to facilitate data management and subsequent data maintenance.

[0189] S105: Construct an evaluation set according to the composite code corresponding to each code identifier saved in the preset database, and adjust the preset vulnerability detection model through the evaluation set, so as to perform code vulnerability detection through the adjusted vulnerability detection model.

[0190] The server can save the above composite code in a preset database, and then, according to these composite codes, construct an evaluation set for adjusting the vulnerability detection model.

[0191] Among them, the evaluation set constructed in this specification is mainly to improve the vulnerability recognition and repair capabilities of the vulnerability detection model. Then, the construction of the evaluation set can be achieved from the following four perspectives, that is, by generating four different types of evaluation samples to construct the evaluation set. These four different types of evaluation samples can enhance the capabilities of the vulnerability detection model from different perspectives.

[0192] The first evaluation sample:

[0193] In this specification, the server can select one composite code as the question stem code from the multiple composite codes corresponding to each code identifier, select at least part of the composite codes from the other composite codes corresponding to the code identifier, and select at least one composite code from the composite codes corresponding to other code identifiers.

[0194] Then, the server can construct the first evaluation sample according to the above question stem code, at least part of the composite codes selected from the other composite codes corresponding to the code identifier, at least one composite code selected from the composite codes corresponding to other code identifiers, and the preset prompt word template, and then, in the subsequent process, construct an evaluation set through the first evaluation sample.

[0195] As can be seen from the above process of constructing the first evaluation sample, the server subsequently uses the first evaluation sample to adjust the vulnerability detection model, mainly to enable the vulnerability detection model to identify code that is semantically unrelated to the code in the question stem. For the first evaluation sample, at least some of the composite codes selected from the other composite codes corresponding to the code identifier belong to the code that is semantically related to the code in the question stem, while at least one of the composite codes selected from the composite codes corresponding to other code identifiers belongs to the code that is semantically unrelated to the code in the question stem.

[0196] In practical applications, the server can generate the first evaluation sample through the following instance of the prompt template:

[0197] <code>{Q}< / code> <options> <option>i. {S[i]}< / option> for i in range(N)]​< / options> , you are a professional code analysis assistant who can understand and explain complex code logic. Now there is a relatively long piece of code <code>, please select from <options>Find the option with the least semantic relevance. Note that only output the correct option you think meets the requirements of the question. Do not output anything other than the option.

[0198] where {Q}, {S[i]} are placeholders, <option>i. {S[i]}< / option> for i in range(N)] means generating N options.

[0199] The server can generate a large number of first evaluation samples by repeatedly using the above examples.

[0200] Second evaluation sample:

[0201] For each code identifier, the server can select a composite code from multiple composite codes corresponding to the code identifier, and use the vulnerability code included in the composite code as the question stem code. Then, according to the code identifier, the server can determine the repair code corresponding to the question stem code, and obtain at least part of the vulnerability codes included in other composite codes corresponding to the code identifier, as well as obtain code segments included in composite codes corresponding to other code identifiers.

[0202] After that, the server can construct a second evaluation sample based on the question stem code, the repair code corresponding to the question stem code, at least part of the vulnerability codes obtained from other composite codes corresponding to the code identifier, the code segments included in composite codes corresponding to other code identifiers, and a preset prompt word template, so as to construct an evaluation set based on the second evaluation sample in the subsequent process.

[0203] From the process of constructing the second evaluation sample above, it can be seen that the server uses the second evaluation sample to adjust the vulnerability detection model later, mainly to enable the vulnerability detection model to identify the repair code for repairing the question stem code from numerous codes. For the second evaluation sample, the repair code corresponding to the question stem code is the correct option for repairing the question stem code, and the others are wrong options. In this way, after using the second evaluation sample to adjust the vulnerability detection model, the vulnerability detection model will have the ability to repair codes, and then can repair the identified vulnerability codes in actual applications.

[0204] In actual applications, the server can generate a second evaluation sample through the following instance of the prompt word template:

[0205] <code>{Q}< / code> <options> <option>i. {G[i]}< / option> for i in range(N)]​< / options> , you are a professional code analysis assistant who can understand and explain complex code logic. Now there is a long piece of code with vulnerabilities <code>, please select from <options>Find the option to repair the vulnerability code. Note that only output the correct options you think meet the requirements of the question. Do not output anything other than the options.

[0206] Each prompt word in the above prompt word template has been explained in the previous examples and will not be elaborated here. The server can reuse this prompt word template to obtain a large number of second evaluation samples.

[0207] It should be noted that in order to further increase the perplexity of the vulnerability detection model, it can be required that the adjustment strategy adopted for the context code in the correct option (i.e., the repair code corresponding to the question stem code) and the adjustment strategy adopted for the context code in the vulnerability code with the same code identifier as the question stem code are consistent with the adjustment strategy adopted by the question stem code. For example, assume that the adjustment strategy adopted in the question stem code is CFP(PC / SC). Then, the repair code and the composite code to which the vulnerability code with the same code identifier as the question stem code belongs also need to adopt CFP(PC / SC).

[0208] In addition, in order to further increase the perplexity, it can be required that the adjustment strategy adopted by the vulnerability code with the same code identifier as the question stem code is different from that of the question stem code. For example, assume that the question stem code adopts the adjustment strategy of CFP(PC / SC). Then, for the composite code to which the vulnerability code with the same code identifier as the question stem code belongs, only CFP(PC / SC)+None(B), CFP(PC / SC)+Bloat(B), and CFP(PC / SC)+All(B) can be selected.

[0209] The third evaluation sample:

[0210] The server can select the vulnerability code contained in a composite code corresponding to each code identifier as the question stem code from multiple composite codes corresponding to the code identifier, read out the line number of the code corresponding to the question stem code from the preset database, and randomly generate several line numbers for the question stem code. Then, the server can generate the third evaluation sample according to the question stem code, the line number of the code corresponding to the question stem code read out from the preset database, the several randomly generated line numbers for the question stem code, and the preset prompt word template, so as to construct an evaluation set according to the third evaluation sample in the subsequent process.

[0211] As can be seen from the above third evaluation sample, the server subsequently uses the third evaluation sample to adjust the vulnerability detection model, mainly to enable the vulnerability detection model to accurately locate the position of the vulnerability code in the code, that is, to accurately determine the line number of the code line where the vulnerability code is located. For the third evaluation sample, the line number of the code corresponding to the question stem code read out from the preset database is the correct option, and the randomly generated line numbers are the wrong options.

[0212] Among them, it is necessary to ensure that the line numbers corresponding to the vulnerability codes read from the preset database do not intersect with the line numbers of several randomly generated code lines in the line number range, and the codes where the starting and ending line numbers are located are not empty.

[0213] In practical applications, the server can generate the third evaluation sample through the following instance of the prompt word template:

[0214] <code>{Q}< / code> <options> <option>i. {L[i]}< / option> for i in range(N)]​< / options> , you are a professional code analysis assistant who can understand and explain complex code logic. Now there is a long piece of code with vulnerabilities <code>, please select from <options>Option to find the line number where the vulnerability lies. Note that only output the correct option you think meets the question. It is prohibited to output anything other than the option.

[0215] Each prompt word in the above prompt word template has been explained in the previous examples and will not be elaborated here. The server can reuse this prompt word template to obtain a large number of third evaluation samples.

[0216] Fourth evaluation sample:

[0217] For each code identifier, the server can select the vulnerability code included in a composite code from the multiple composite codes corresponding to the code identifier as the question stem code, read the vulnerability type information corresponding to the question stem code from the preset database, and randomly obtain the vulnerability type information of several other vulnerability types. Then, the server can generate a fourth evaluation sample based on the question stem code, the vulnerability type information corresponding to the question stem code read from the preset database, the randomly obtained vulnerability type information of several other vulnerability types, and the preset prompt word template, so as to construct an evaluation set based on the generated fourth evaluation sample in the subsequent process.

[0218] As can be seen from the above fourth evaluation sample, the server uses the fourth evaluation sample to adjust the vulnerability detection model later, mainly to enable the vulnerability detection model to identify the vulnerability type corresponding to the question stem code. For the fourth evaluation sample, the vulnerability type information corresponding to the question stem code read from the preset database is the correct option, while the randomly obtained vulnerability type information of several other vulnerability types belongs to the misaligned option.

[0219] In practical applications, the server can generate a fourth evaluation sample through the following instance of the prompt word template:

[0220] <code>{Q}< / code> <options> <option>i. {T[i]}< / option> for i in range(N)]​< / options> , you are a professional code analysis assistant who can understand and explain complex code logic. Now there is a long piece of code with vulnerabilities <code>, please select from <options>Select the option corresponding to the vulnerability type. Note that only output the correct option that you think meets the requirements of the question. It is prohibited to output anything other than the options.

[0221] Each prompt word in the above prompt word template has been explained in the previous examples and will not be elaborated here. The server can reuse this prompt word template to obtain a large number of fourth evaluation samples.

[0222] From the four evaluation samples generated by the above server, it can be seen that the first evaluation sample is used to improve the vulnerability code recognition ability of the vulnerability detection model, the second evaluation sample is used to improve the code repair ability of the vulnerability detection model, the third evaluation sample is used to improve the vulnerability code location ability of the vulnerability detection model, and the fourth evaluation sample is used to improve the vulnerability type recognition ability of the vulnerability detection model.

[0223] Through the above evaluation samples, an evaluation set can be constructed. By using the constructed evaluation set, the ability of the vulnerability detection model can be enhanced from different perspectives, providing a basis for accurate vulnerability identification and repair of the vulnerability detection model in actual applications.

[0224] Furthermore, in this specification, after the server constructs the above evaluation set, it can store the evaluation samples in the evaluation set, specifically through the Figure 7 shown storage structure to achieve.

[0225] Figure 7 It is a schematic diagram of the storage structure of an evaluation set provided in this specification.

[0226] In Figure 7 the evaluation sample identifier can be obtained by calculating the hash value, and the evaluation sample type is used to mark whether the evaluation sample in an evaluation set belongs to the first evaluation sample, the second evaluation sample, the third evaluation sample, or the fourth evaluation sample.

[0227] The code identifier is used to represent the code identifier corresponding to the stem code in the evaluation sample, and the evaluation sample text refers to the text obtained after filling in the text according to the preset prompt word template. Finally, Figure 7 the correct answer options corresponding to the evaluation samples are also recorded in the shown storage structure.

[0228] It should be noted that for the convenience of retrieval and management, the database storing various data in this specification can be a relational database.

[0229] The server can adjust the vulnerability detection model through the evaluation set constructed in the above manner. Among them, the adjustment methods adopted can be existing conventional model adjustment methods, and this specification does not limit the specific adjustment methods. Once the adjustment is completed, the server can deploy the vulnerability detection model to execute the code vulnerability detection task through this vulnerability detection model.

[0230] As can be seen from the above method, since the comment information in each sample code included in the sample code set can be deleted after obtaining the sample code set, the evaluation set constructed subsequently will not leak the software code semantics to the vulnerability detection model, so that the vulnerability detection model can output results that conform to its actual vulnerability identification ability.

[0231] Moreover, by adjusting and modifying the control flow and non-critical identifiers of the sample code, the final obtained evaluation samples can provide a certain degree of recognition difficulty for the vulnerability detection model. Therefore, after the model adjustment stage, the ability of the vulnerability detection model can be significantly enhanced.

[0232] Furthermore, by adjusting the desensitized code, multiple composite codes can be obtained. Through these composite codes, a variety of evaluation samples can be obtained, so that the vulnerability detection model can be enhanced from multiple angles in terms of code vulnerability identification, location ability, vulnerability type analysis ability, and code repair ability. Therefore, in subsequent actual applications, the recognition accuracy of the vulnerability detection model and the code repair effect can be significantly improved.

[0233] The above is a code vulnerability detection method provided by one or more embodiments of the present application. Based on the same idea, the embodiments of the present application also provide a corresponding code vulnerability detection device, as Figure 8 shown.

[0234] Figure 8 It is a schematic diagram of a code vulnerability detection device provided by an embodiment of the present application, specifically including:

[0235] An acquisition module 801, configured to acquire a sample code set. Each sample code is included in the sample code set. For each sample code, the sample code corresponds to a code identifier, and each code segment is included in the sample code. The code categories of the code segments include: vulnerability code, the repair code corresponding to the vulnerability code, and the context code of the vulnerability code;

[0236] An identification module 802, configured to identify and delete the comment information in each sample code to obtain a desensitized code corresponding to each sample code;

[0237] An adjustment module 803 is configured to adjust each desensitized code according to a preset adjustment strategy to obtain at least one enhanced code corresponding to the desensitized code, where the at least one enhanced code has the same code identifier as the desensitized code;

[0238] A combination module 804 is configured to combine each code segment belonging to a code identifier to obtain multiple composite codes corresponding to the code identifier, and store them in a preset database. Each code segment included in a composite code has a different code category;

[0239] A construction module 805 is configured to construct a test set according to the composite codes corresponding to each code identifier stored in the preset database, and adjust a preset vulnerability detection model through the test set, so as to perform code vulnerability detection through the adjusted vulnerability detection model.

[0240] Optionally, the adjustment module 803 is specifically configured to, for each desensitized code, adjust the control flow of the desensitized code to obtain at least one enhanced code corresponding to the desensitized code, where the control flow is used to represent the execution order and manner of code statements in the desensitized code.

[0241] Optionally, the adjustment module 803 is specifically configured to, for each desensitized code, add redundant code to the desensitized code to obtain at least one enhanced code corresponding to the desensitized code.

[0242] Optionally, the construction module 805 is specifically configured to, for each code identifier, select a composite code as a question stem code from the multiple composite codes corresponding to the code identifier, select at least some composite codes from the other composite codes corresponding to the code identifier, and select at least one composite code from the composite codes corresponding to other code identifiers; construct a first test sample according to the question stem code, the at least some composite codes selected from the other composite codes corresponding to the code identifier, the at least one composite code selected from the composite codes corresponding to other code identifiers, and a preset prompt word template, so as to construct a test set according to the first test sample, and the first test sample is used to enable the vulnerability detection model to identify codes that are semantically irrelevant to the question stem code.

[0243] Optionally, the building module 805 is specifically configured to, for each code identifier, select one composite code from multiple composite codes corresponding to the code identifier, and use the vulnerability code included in the composite code as the question stem code. According to the code identifier, determine the repair code corresponding to the question stem code, obtain at least some of the vulnerability codes included in other composite codes corresponding to the code identifier, and obtain the code snippets included in the composite codes corresponding to other code identifiers; according to the question stem code, the repair code corresponding to the question stem code, at least some of the vulnerability codes obtained from other composite codes corresponding to the code identifier, the code snippets included in the composite codes corresponding to other code identifiers, and a preset prompt word template, construct a second evaluation sample, so as to construct an evaluation set according to the second evaluation sample. The second evaluation sample is used to enable the vulnerability detection model to select a repair code for repairing the question stem code.

[0244] Optionally, for each composite code, the preset database also stores the line numbers of the code corresponding to the vulnerability code included in the composite code;

[0245] The building module 805 is specifically configured to, for each code identifier, select the vulnerability code included in one composite code from multiple composite codes corresponding to the code identifier as the question stem code, read the line numbers of the code corresponding to the question stem code from the preset database, and randomly generate several line numbers for the question stem code; according to the question stem code, the line numbers of the code corresponding to the question stem code read from the preset database, several line numbers randomly generated for the question stem code, and a preset prompt word template, generate a third evaluation sample, so as to construct an evaluation set according to the third evaluation sample. The third evaluation sample is used to enable the vulnerability detection model to determine the line numbers of the code of the question stem code.

[0246] Optionally, for each composite code, the preset database also stores the vulnerability type information corresponding to the vulnerability code included in the composite code;

[0247] The building module 805 is specifically configured to, for each code identifier, select the vulnerability code included in one composite code from multiple composite codes corresponding to the code identifier as the question stem code, read the vulnerability type information corresponding to the question stem code from the preset database, and randomly obtain the vulnerability type information of several other vulnerability types; according to the question stem code, the vulnerability type information corresponding to the question stem code read from the preset database, the vulnerability type information of several other vulnerability types randomly obtained, and a preset prompt word template, generate a fourth evaluation sample, so as to construct an evaluation set according to the fourth evaluation sample. The fourth evaluation sample is used to enable the vulnerability detection model to identify the vulnerability type corresponding to the question stem code.

[0248] The embodiments of the present application further provide a computer-readable storage medium storing a computer program, which can be used to execute the Figure 1 code vulnerability detection method provided above.

[0249] The embodiments of the present application also provide Figure 9 a schematic structural diagram of the electronic device shown in the figure. As Figure 9 shown, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the Figure 1 code vulnerability detection method described above.

[0250] Of course, in addition to the software implementation, the present application does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, and can also be hardware or logic devices.

[0251] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to circuit structures such as diodes, transistors, switches, etc.) or software improvements (improvements to method flows). However, with the development of technology, many method flow improvements today can be regarded as direct improvements to hardware circuit structures. Almost all designers obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement to a method flow cannot be implemented using a hardware entity module. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is an integrated circuit whose logical function is determined by the user programming the device. Designers can program a digital system "integrated" onto a single PLD by themselves, without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compiler used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a Hardware Description Language (HDL), and there is not just one type of HDL, but many types, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones currently are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that by simply performing a little logical programming on the method flow using the above-mentioned several hardware description languages and programming it into an integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.

[0252] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to implement the same functions of the controller in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or structures within the hardware component.

[0253] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0254] For the convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0255] Those skilled in the art should understand that the embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, this specification can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.

[0256] This specification is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the specification. It should be understood that each flow and / or block in the flowchart and / or block diagram, and combinations of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0257] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0258] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are performed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0259] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0260] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). Memory is an example of computer-readable media.

[0261] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0262] It should also be noted that the term "includes", "comprising" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of further restrictions, an element defined by the sentence "comprising a ..." does not exclude the existence of other identical elements in the process, method, commodity or device including the element.

[0263] It should be understood by those skilled in the art that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0264] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0265] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For the relevant parts, reference can be made to the corresponding descriptions in the method embodiment.

[0266] The above description is only for the embodiments of this specification and is not intended to limit this specification. For those skilled in the art, various modifications and changes can be made to this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included within the scope of the claims of this specification.< / options> < / code> < / options> < / code> < / options> < / code> < / options> < / code> < / code> < / lang> < / code> < / lang> < / code>

Claims

1. A code vulnerability detection method, characterized in that: include: Acquire a sample code set, wherein the sample code set includes various sample codes, and for each sample code, the sample code corresponds to a code identifier, and the sample code includes various code fragments, and the code category of each code fragment includes: vulnerability code, repair code corresponding to the vulnerability code, and context code of the vulnerability code; Identify and delete the annotation information in each sample code to obtain a desensitized code corresponding to each sample code; For each desensitization code, the desensitization code is adjusted according to a preset adjustment strategy to obtain at least one enhanced code corresponding to the desensitization code, wherein the at least one enhanced code is the same as the code identifier of the desensitization code; For each code identifier, the code fragments belonging to the code identifier are combined to obtain multiple composite codes corresponding to the code identifier, and the composite codes are stored in a preset database, and the code categories of the code fragments contained in a composite code are all different; According to the composite code corresponding to each code identifier stored in the preset database, an evaluation set is constructed, and the preset vulnerability detection model is adjusted through the evaluation set to perform code vulnerability detection through the adjusted vulnerability detection model.

2. The method according to claim 1, characterized in that For each desensitized code, the desensitized code is adjusted according to a preset adjustment strategy to obtain at least one enhanced code corresponding to the desensitized code, specifically including: For each desensitized code, the control flow of the desensitized code is adjusted to obtain at least one enhanced code corresponding to the desensitized code, wherein the control flow is used to represent the execution order and method of the code statements in the desensitized code.

3. The method according to claim 1, characterized in that For each desensitized code, the desensitized code is adjusted according to a preset adjustment strategy to obtain at least one enhanced code corresponding to the desensitized code, specifically including: For each desensitized code, a redundant code is added to the desensitized code to obtain at least one enhanced code corresponding to the desensitized code.

4. The method according to claim 1, characterized in that According to the composite code corresponding to each code identifier stored in the preset database, an evaluation set is constructed, specifically including: For each code identifier, select one compound code from multiple compound codes corresponding to the code identifier as the question stem code, select at least part of the compound codes from other compound codes corresponding to the code identifier, and select at least one compound code from compound codes corresponding to other code identifiers; A first evaluation sample is constructed based on the question stem code, at least part of the compound code selected from other compound codes corresponding to the code identifier, at least one compound code selected from compound codes corresponding to other code identifiers, and a preset prompt word template, and an evaluation set is constructed based on the first evaluation sample, wherein the first evaluation sample is used to enable the vulnerability detection model to identify codes that are semantically irrelevant to the question stem code.

5. The method according to claim 1, characterized in that According to the composite code corresponding to each code identifier stored in the preset database, an evaluation set is constructed, specifically including: For each code identifier, select a composite code from multiple composite codes corresponding to the code identifier, and use the vulnerability code contained in the composite code as the question stem code, determine the repair code corresponding to the question stem code according to the code identifier, obtain at least part of the vulnerability code contained in other composite codes corresponding to the code identifier, and obtain code snippets contained in composite codes corresponding to other code identifiers; A second evaluation sample is constructed based on the question stem code, the repair code corresponding to the question stem code, at least part of the vulnerability code obtained from other composite codes corresponding to the code identifier, code snippets contained in the composite code corresponding to other code identifiers, and a preset prompt word template, and an evaluation set is constructed based on the second evaluation sample, and the second evaluation sample is used to enable the vulnerability detection model to select the repair code for repairing the question stem code.

6. The method according to claim 1, characterized in that For each composite code, the preset database also stores the code line number corresponding to the vulnerability code contained in the composite code; According to the composite code corresponding to each code identifier stored in the preset database, an evaluation set is constructed, specifically including: For each code identifier, a vulnerability code contained in a composite code is selected from multiple composite codes corresponding to the code identifier as the question stem code, the code line number corresponding to the question stem code is read from the preset database, and a number of code line numbers for the question stem code are randomly generated; A third evaluation sample is generated based on the question stem code, the code line number corresponding to the question stem code read from the preset database, a number of code line numbers randomly generated for the question stem code, and a preset prompt word template, so as to construct an evaluation set based on the third evaluation sample, and the third evaluation sample is used to enable the vulnerability detection model to determine the code line number of the question stem code.

7. The method according to claim 1, characterized in that For each composite code, the preset database also stores vulnerability type information corresponding to the vulnerability code contained in the composite code; According to the composite code corresponding to each code identifier stored in the preset database, an evaluation set is constructed, specifically including: For each code identifier, a vulnerability code contained in a composite code is selected from multiple composite codes corresponding to the code identifier as a question stem code, vulnerability type information corresponding to the question stem code is read from the preset database, and vulnerability type information of several other vulnerability types is randomly obtained; A fourth evaluation sample is generated based on the question code, the vulnerability type information corresponding to the question code read from the preset database, the vulnerability type information of several other vulnerability types obtained randomly, and a preset prompt word template, so as to construct an evaluation set based on the fourth evaluation sample, and the fourth evaluation sample is used to enable the vulnerability detection model to identify the vulnerability type corresponding to the question code.

8. A code vulnerability detection device, characterized in that: include: An acquisition module is used to acquire a sample code set, wherein the sample code set includes various sample codes, and for each sample code, the sample code corresponds to a code identifier, and the sample code includes various code fragments, and the code category of each code fragment includes: a vulnerability code, a repair code corresponding to the vulnerability code, and a context code of the vulnerability code; An identification module, used to identify and delete the annotation information in each sample code to obtain a desensitized code corresponding to each sample code; An adjustment module, configured to adjust each desensitization code according to a preset adjustment strategy to obtain at least one enhanced code corresponding to the desensitization code, wherein the at least one enhanced code is the same as the code identifier of the desensitization code; A combination module is used to combine the code fragments belonging to each code identifier to obtain multiple composite codes corresponding to the code identifier, and store them in a preset database. The code categories of the code fragments contained in a composite code are different. The construction module is used to construct an evaluation set according to the composite code corresponding to each code identifier stored in the preset database, and adjust the preset vulnerability detection model through the evaluation set to perform code vulnerability detection through the adjusted vulnerability detection model.

9. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method described in any one of claims 1 to 7 is implemented.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Source code detection method and device, computer equipment and storage medium

    CN116257856A

  • Code vulnerability remediation

    US20210124830A1