Code vulnerability detection method and device, medium and equipment

By identifying and deleting comment information, adjusting and combining code snippets, building evaluation sets, and adjusting vulnerability detection models, the problems of low code vulnerability detection efficiency and high false alarm rate in the existing technology are solved, and higher recognition accuracy and code repair effects are achieved.

CN119939608AActive Publication Date: 2025-05-06ZHEJIANG LAB
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510423675.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-05-06
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

The existing technology is inefficient in code vulnerability detection, has a high false alarm rate, and large models rely on comment information and code snippet lengths, which affects the accuracy of vulnerability identification.

Method used

By obtaining sample code sets, identifying and deleting annotation information, adjusting desensitization codes, combining code snippets, building evaluation sets, and adjusting vulnerability detection models to improve the accuracy and effectiveness of vulnerability detection.

Benefits of technology

It significantly improves the identification accuracy and code repair effect of the vulnerability detection model, reduces the false positive rate, and enhances the multi-angle recognition ability of the model for code vulnerabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939608A_ABST
    Figure CN119939608A_ABST
Patent Text Reader

Abstract

The invention discloses a code vulnerability detection method and device, a medium and equipment, after a sample code set is obtained, annotation information in each sample code contained in the sample code set can be identified and deleted to obtain each desensitization code, and then each desensitization code is adjusted according to a preset adjustment strategy, so that the code vulnerability detection efficiency is improved. The method comprises the steps of obtaining at least one enhanced code which is the same as a code identifier of a desensitization code, subsequently combining code snippets belonging to the same code identifier to obtain a plurality of composite codes corresponding to one code identifier, and then constructing an evaluation set through the composite codes, so as to adjust a vulnerability detection model through the evaluation set, thereby improving the vulnerability detection accuracy. And code vulnerability detection is performed through the adjusted vulnerability detection model, so that the recognition accuracy of the vulnerability detection model can be remarkably improved in subsequent practical application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the fields of computer technology and artificial intelligence, and in particular to a code vulnerability detection method, device, medium and equipment. Background Art

[0002] As the complexity of software systems continues to increase, vulnerability detection has become an important means to ensure the normal operation and security of software. Traditional vulnerability detection methods mainly rely on manual analysis and static code analysis tools, but these methods are inefficient when facing software with large-scale code, and the false positive rate is often high.

[0003] With the rapid development of artificial intelligence technology in recent years, vulnerability detection through large models has become a new idea. By learning a large number of software samples with vulnerabilities, large models can understand the semantics of long codes and automatically identify potential vulnerabilities. They can also effectively reduce the false positive rate through semantic analysis and provide explainable analysis reports.

[0004] However, there are also certain problems in using large models to perform vulnerability detection. For example, since the software code in the sample set contains a large number of comment statements and declarations of classes, functions, variables, etc. with descriptive meanings, it is easy for the large model to leak the semantics of the software code, so that the large model can know whether the code has vulnerabilities only through these comments without identifying the code itself, resulting in distorted vulnerability detection results. In addition, some code snippets in the current sample set are not long enough and lack complete context, which significantly affects the large model's accurate positioning of vulnerabilities.

[0005] In addition, the current large models are limited to determining whether there are vulnerabilities in the software and the types of vulnerabilities, and cannot effectively repair the vulnerabilities in the software. Moreover, the current large models may only detect vulnerabilities through simple feature matching, rather than identifying vulnerabilities through semantic understanding, resulting in insufficient generalization capabilities. Summary of the invention

[0006] The embodiments of the present application provide a code vulnerability detection method, device, medium and equipment to partially solve the above-mentioned problems existing in the prior art.

[0007] This application adopts the following technical solutions: The present application provides a method for detecting code vulnerabilities, including: Acquire a sample code set, wherein the sample code set includes various sample codes, and for each sample code, the sample code corresponds to a code identifier, and the sample code includes various code fragments, and the code category of each code fragment includes: vulnerability code, repair code corresponding to the vulnerability code, and context code of the vulnerability code; Identify and delete the annotation information in each sample code to obtain a desensitized code corresponding to each sample code; For each desensitization code, the desensitization code is adjusted according to a preset adjustment strategy to obtain at least one enhanced code corresponding to the desensitization code, wherein the at least one enhanced code is the same as the code identifier of the desensitization code; For each code identifier, the code fragments belonging to the code identifier are combined to obtain multiple composite codes corresponding to the code identifier, and the composite codes are stored in a preset database. The code categories of the code fragments contained in a composite code are all different. An evaluation set is constructed according to the composite code corresponding to each code identifier stored in the preset database, and a preset vulnerability detection model is adjusted through the evaluation set to perform code vulnerability detection through the adjusted vulnerability detection model.

[0008] Optionally, for each desensitized code, adjusting the desensitized code according to a preset adjustment strategy to obtain at least one enhanced code corresponding to the desensitized code specifically includes: For each desensitized code, the control flow of the desensitized code is adjusted to obtain at least one enhanced code corresponding to the desensitized code, wherein the control flow is used to represent the execution order and method of the code statements in the desensitized code.

[0009] Optionally, for each desensitized code, adjusting the desensitized code according to a preset adjustment strategy to obtain at least one enhanced code corresponding to the desensitized code specifically includes: For each desensitized code, a redundant code is added to the desensitized code to obtain at least one enhanced code corresponding to the desensitized code.

[0010] Optionally, constructing an evaluation set according to the composite code corresponding to each code identifier stored in the preset database specifically includes: For each code identifier, select one compound code from multiple compound codes corresponding to the code identifier as the question stem code, select at least some compound codes from other compound codes corresponding to the code identifier, and select at least one compound code from compound codes corresponding to other code identifiers; A first evaluation sample is constructed based on the question stem code, at least part of the compound code selected from other compound codes corresponding to the code identifier, at least one compound code selected from compound codes corresponding to other code identifiers, and a preset prompt word template, and an evaluation set is constructed based on the first evaluation sample, wherein the first evaluation sample is used to enable the vulnerability detection model to identify codes that are semantically irrelevant to the question stem code.

[0011] Optionally, constructing an evaluation set according to the composite code corresponding to each code identifier stored in the preset database specifically includes: For each code identifier, select a composite code from multiple composite codes corresponding to the code identifier, and use the vulnerability code contained in the composite code as the question stem code, determine the repair code corresponding to the question stem code according to the code identifier, obtain at least part of the vulnerability code contained in other composite codes corresponding to the code identifier, and obtain code snippets contained in composite codes corresponding to other code identifiers; A second evaluation sample is constructed based on the question stem code, the repair code corresponding to the question stem code, at least part of the vulnerability code obtained from other composite codes corresponding to the code identifier, code snippets contained in the composite code corresponding to other code identifiers, and a preset prompt word template, and an evaluation set is constructed based on the second evaluation sample, and the second evaluation sample is used to enable the vulnerability detection model to select the repair code for repairing the question stem code.

[0012] Optionally, for each composite code, the preset database also stores the code line number corresponding to the vulnerability code contained in the composite code; According to the composite code corresponding to each code identifier stored in the preset database, an evaluation set is constructed, specifically including: For each code identifier, a vulnerability code contained in a composite code is selected from multiple composite codes corresponding to the code identifier as the question stem code, the code line number corresponding to the question stem code is read from the preset database, and a number of code line numbers for the question stem code are randomly generated; A third evaluation sample is generated based on the question stem code, the code line number corresponding to the question stem code read from the preset database, a number of code line numbers randomly generated for the question stem code, and a preset prompt word template, so as to construct an evaluation set based on the third evaluation sample, and the third evaluation sample is used to enable the vulnerability detection model to determine the code line number of the question stem code.

[0013] Optionally, for each composite code, the preset database also stores vulnerability type information corresponding to the vulnerability code contained in the composite code; According to the composite code corresponding to each code identifier stored in the preset database, an evaluation set is constructed, specifically including: For each code identifier, a vulnerability code contained in a composite code is selected from multiple composite codes corresponding to the code identifier as a question stem code, vulnerability type information corresponding to the question stem code is read from the preset database, and vulnerability type information of several other vulnerability types is randomly obtained; A fourth evaluation sample is generated based on the question code, the vulnerability type information corresponding to the question code read from the preset database, the vulnerability type information of several other vulnerability types obtained randomly, and a preset prompt word template, so as to construct an evaluation set based on the fourth evaluation sample, and the fourth evaluation sample is used to enable the vulnerability detection model to identify the vulnerability type corresponding to the question code.

[0014] The present application provides a code vulnerability detection device, including: An acquisition module is used to acquire a sample code set, wherein the sample code set includes various sample codes, and for each sample code, the sample code corresponds to a code identifier, and the sample code includes various codes, and the code category of each code fragment includes: a vulnerability code, a repair code corresponding to the vulnerability code, and a context code of the vulnerability code; An identification module, used to identify and delete the annotation information in each sample code to obtain a desensitized code corresponding to each sample code; An adjustment module, configured to adjust each desensitization code according to a preset adjustment strategy to obtain at least one enhanced code corresponding to the desensitization code, wherein the at least one enhanced code is the same as the code identifier of the desensitization code; A combination module is used to combine the code fragments belonging to each code identifier to obtain multiple composite codes corresponding to the code identifier, and store them in a preset database. The code categories of the code fragments contained in a composite code are different. The construction module is used to construct an evaluation set according to the composite code corresponding to each code identifier stored in the preset database, and adjust the preset vulnerability detection model through the evaluation set to perform code vulnerability detection through the adjusted vulnerability detection model.

[0015] An embodiment of the present application provides a computer-readable storage medium, wherein the storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned code vulnerability detection method is implemented.

[0016] An embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements a code vulnerability detection method when executing the program.

[0017] At least one of the above technical solutions adopted in the embodiments of the present application can achieve the following beneficial effects: The code vulnerability detection method provided in the embodiment of the present application can, after obtaining a sample code set, identify and delete the comment information in each sample code contained in the sample code set to obtain each desensitized code, and then, according to a preset adjustment strategy, adjust each desensitized code to obtain at least one enhanced code with the same code identifier as the desensitized code, and subsequently, by combining the code fragments belonging to the same code identifier, multiple composite codes corresponding to a code identifier can be obtained, and then an evaluation set is constructed through these composite codes, so that the vulnerability detection model can be adjusted through the evaluation set, and code vulnerability detection is performed through the adjusted vulnerability detection model.

[0018] From the above method, it can be seen that after obtaining the sample code set, the annotation information in each sample code contained in the sample code set can be deleted, so that the evaluation set constructed subsequently will not leak the software code semantics to the vulnerability detection model, so that the vulnerability detection model can output results that are consistent with its actual vulnerability identification capabilities. Moreover, by adjusting the desensitized code, multiple composite codes can be obtained, and through these composite codes, a variety of evaluation samples can be obtained, so that the vulnerability detection model's ability to identify code vulnerabilities can be enhanced from multiple angles, and then in subsequent practical applications, the recognition accuracy of the vulnerability detection model can be significantly improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The drawings described herein are used to provide a further understanding of the present specification and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present specification and do not constitute an improper limitation on the present application. In the drawings: Figure 1 A flowchart of a code vulnerability detection method provided in an embodiment of the present application; Figure 2 A schematic diagram of a sample code storage structure provided in this specification; Figure 3 A schematic diagram of a storage structure of a code snippet provided in this specification; Figure 4 A schematic diagram of a storage structure used to store each code fragment in the desensitized code provided in this specification; Figure 5 A schematic diagram of the storage structure of the enhanced code snippet obtained after adjustment provided in this specification; Figure 6 A schematic diagram of a storage structure of a composite code provided in this specification; Figure 7 A schematic diagram of the storage structure of a test set provided in this specification; Figure 8A schematic diagram of a device for detecting code vulnerabilities provided in an embodiment of the present application; Fig. 9 A method corresponding to the embodiment of the present application is provided Figure 1 Schematic diagram of the structure of an electronic device. DETAILED DESCRIPTION

[0020] In order to make the purpose, technical solutions and advantages of this application clearer, the technical solutions of this specification will be clearly and completely described below in combination with the specific embodiments of this application and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of this specification, not all of them. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0021] The technical solutions provided by various embodiments of the present application are described in detail below in conjunction with the accompanying drawings.

[0022] Figure 1 A flowchart of a code vulnerability detection method provided in an embodiment of the present application includes the following steps: S101: Obtain a sample code set.

[0023] In this specification, in order to enable the vulnerability detection model to accurately identify and repair vulnerabilities contained in software codes in actual applications, it is necessary to adjust the vulnerability detection model through the constructed evaluation set before using it, so that the vulnerability detection model after the adjustment of the evaluation set has powerful software vulnerability identification and repair capabilities.

[0024] Prior to this, it is necessary to first construct an evaluation set for adjusting the vulnerability detection model. Therefore, the code vulnerability detection method provided in this manual can actually be divided into a data set construction stage, a model adjustment stage, and a vulnerability detection stage. The focus of this manual is mainly on the data set construction stage, because the evaluation set constructed by the method provided in this manual contains a variety of evaluation samples. These evaluation samples can adjust the vulnerability detection model from multiple angles, so that the vulnerability detection model can learn about software code vulnerability detection and repair "knowledge" from different angles, thereby exerting better code vulnerability detection and repair capabilities in actual applications, and providing more convenience for users.

[0025] In the process of building an evaluation set, it is necessary to first obtain a sample code set, wherein the sample code set contains multiple sample codes, and for any sample code, the sample code may contain various code snippets, wherein the code category of each code snippet may include: vulnerability code (i.e., code with vulnerabilities), repair code corresponding to the vulnerability code (i.e., code without vulnerabilities after the vulnerability code is repaired), and context code of the vulnerability code (i.e., code used to connect to the vulnerability code).

[0026] The above sample code set can be obtained from an existing code set. For the obtained sample code set, the vulnerability code and the repair code in each sample code contained therein are marked in advance, and can be stored in a preset database in a certain way, such as Figure 2 shown.

[0027] Figure 2 A schematic diagram of a sample code storage structure provided in this specification.

[0028] from Figure 2 As can be seen from the figure, for a sample code, the sample code corresponds to a unique code identifier, which can be obtained by calculating a hash value. The sample code can include vulnerability code, repair code, context upper segment and context lower segment.

[0029] The upper context segment can be understood as the code connected before the vulnerability code or the repair code, and the lower context segment can be understood as the code connected after the vulnerability code or the repair code.

[0030] exist Figure 2 The storage structure in also records the line number of the vulnerability code in the sample code and the line number of the repair code in the sample code. Figure 2 The storage structure in also records the programming language corresponding to the sample code. The reason why the programming language needs to be recorded is that when the composite code is obtained in the subsequent process, the operation methods used for codes in different programming languages ​​are different.

[0031] Figure 2 The storage structure in the sample code also records the vulnerability types corresponding to the vulnerability codes contained in the sample code, such as out-of-bounds write (the code program attempts to write data in the memory at a location beyond the memory area allocated to it during execution), out-of-bounds read (the code program attempts to access data outside its allocated memory area during execution), improper input validation (the code program fails to fully check whether the data provided by the user or external system meets the expected format, type, range, etc.), etc. This helps the subsequent vulnerability detection model to identify vulnerability types.

[0032] In this specification, the vulnerability detection model that needs to be adjusted and used for vulnerability detection can be a large language model. In actual applications, the code that needs to be detected can be input into the vulnerability detection model together with the prompt statement. The vulnerability detection model can recognize the user's intention through the received prompt statement, thereby completing the vulnerability detection task.

[0033] In addition, the execution subject of the code vulnerability detection method provided in this specification can be a terminal device such as a desktop computer, a laptop computer, etc., which can be a server or a client installed in the terminal device. For the sake of convenience, the following only takes the server as the execution subject as an example to explain in detail the code vulnerability detection method provided in this specification.

[0034] After obtaining the above sample code set, the server can perform a series of processing on each sample code contained in the sample code set. Specifically, the server can convert each sample code contained in the sample code set into a specified encoding format, thereby obtaining a converted sample code set. The reason for converting into a unified encoding format is to avoid exceptions during the parsing process. In order to consider factors such as universality and extensiveness, each sample code in the sample code set can be uniformly converted into the UTF-8 encoding format.

[0035] Then, the server can locate the code of each sample code in the sample code set to extract the vulnerability code, the repair code and the context code respectively. For each sample code in the sample code set, the server can locate the interval of the vulnerability code in the sample code according to the code line number of the vulnerability code, thereby extracting the vulnerability code located in the interval. Similarly, the server can locate the interval of the repair code in the sample code according to the code line number of the repair code, thereby extracting the repair code located in the interval.

[0036] As for the context code, the server can determine the upper context segment and the lower context segment according to the line number of the vulnerability code or the line number of the repair code. Specifically, the following formula can be used to determine the line number of the upper context segment and the line number of the lower context segment respectively: The code line number of the context above is: [0,z1=min(x1,y1)] The code line number of the following context is: [z2=max(x2,y2),-1], where -1 indicates the last line.

[0037] After obtaining the above line number, the server can extract the upper context segment PC=S[0,z1] and the lower context segment SC=S[z2,-1]. PC stands for the upper context segment, and SC stands for the lower context segment.

[0038] The server can extract vulnerability code, repair code and context code from the sample code in the above manner. After that, the server can store these code fragments in a preset database according to a certain storage structure, such as Figure 3 shown.

[0039] Figure 3 A schematic diagram of a storage structure of a code snippet provided in this specification.

[0040] For any code snippet, the server can calculate the hash value according to the code information contained in the code snippet to obtain the identifier of the code snippet. Figure 3 The storage structure shown also records the code identifier of the sample code to which the code snippet belongs. Figure 3 The code section in is used to record the code snippet itself. And, Figure 3 The code category to which the code snippet belongs is also recorded, where PC represents the upper context segment, SC represents the lower context segment, B represents the vulnerability code, and G represents the repair code.

[0041] S102: Identify and delete the annotation information in each sample code to obtain a desensitized code corresponding to each sample code.

[0042] Since the current code contains a large amount of annotation information, this annotation information will leak the semantics of the software code, causing the vulnerability detection model to no longer perform semantic analysis on the code itself, but only rely on the annotation information to identify vulnerabilities, resulting in distorted detection results.

[0043] To this end, in this specification, after obtaining the above sample code set, the server can identify the annotation information contained in each sample code contained in the sample code set to obtain the desensitized code corresponding to each sample code. The server can use a preset regular expression to identify and delete the annotation information in the sample code.

[0044] For example, for C / C++, comments are divided into single-line comments (starting with / / ) and multi-line comments (starting with / * and ending with * / ). The regular expression for deleting single-line comments in C / C++ is: re.sub(r' / / .*', '', code) In this regular expression, the start symbol for matching comments is / / ; and the symbol .* means matching all characters after / / in the line.

[0045] The C / C++ multi-line comment removal regular expression is: re.sub(r' / \*.*?\* / ', '', code, flags=re.DOTALL) In this regular expression, / \\*: This part matches the start tag of the comment / *; .*? matches any character (non-greedy mode) until * / is encountered; \* / matches the end of the comment * / ; flags=re.DOTALL allows . to match the newline character, thus supporting cross-line comments.

[0046] For Python, comments are divided into single-line comments (starting with #) and multi-line comments (starting and ending with single or double quotes, including multi-line strings).

[0047] Among them, the Python multi-line comment deletion regular expression is: re.compile( r""" (\"\"\"[^\"\"\"]*\"\"\" | # Matches a multi-line double-quoted string \'\'\'[^\'\'\']*\'\'\' | # Matches multi-line single-quoted strings \"[^\"]*\" | # Matches a single-line double-quoted string \'[^\']*\' ) # matches a single-quoted string """, re.VERBOSE ) Among them, the Python single-line comment deletion regular expression is: re.compile(r"#.*$", re.MULTILINE) In this regular expression, re.MULTILINE modifies the behavior of ^ and $ so that they match the beginning and end of each line, respectively, rather than the beginning and end of the entire string.

[0048] In order to further enhance the vulnerability detection and repair capabilities of the vulnerability detection model, in this specification, in addition to deleting the comments in the above sample code, the server can also modify some identifiers in the sample code based on the sample code with the comments deleted.

[0049] The so-called partial identifier here can be understood as information that does not change the original semantic characteristics of the code. In this specification, the so-called identifier can refer to things such as class names, variable names, and function names.

[0050] Therefore, after the server identifies and deletes the comment information in each sample code, it obtains the initial desensitized code corresponding to each sample code. Then, the server can identify the non-critical identifiers contained in the initial desensitized code corresponding to each sample code, and modify these non-critical identifiers to obtain the desensitized code corresponding to each sample code. Among them, the non-critical identifiers mentioned here include the class names of non-critical classes, the variable names of non-critical variables, and the function names of non-built-in functions. The modification of non-critical identifiers does not change the original semantic characteristics of the code.

[0051] Therefore, when the server extracts identifiers such as class names, variable names, function names, etc. in the above-mentioned initial desensitized code, it will detect whether these identifiers are key identifiers. The so-called key identifiers refer to those that will change the original semantic characteristics of the code after modification. Therefore, the modification of some identifiers in the initial desensitized code in this specification only plays a certain role in confusion, and is not intended to completely change the original semantic characteristics of the initial desensitized code.

[0052] Furthermore, for different programming languages, the server may use different regular expressions to modify non-critical identifiers in the above initial desensitized code.

[0053] For example, C / C++ is a statically typed language and a strongly typed language. The data type of a variable must be determined during compilation, and the rules are relatively clear. Class names, function names, and variable names can all be considered as a type of variable.

[0054] Therefore, the regular expressions used to match C / C++ class names, function names, and variable names are as follows: var_regex=re.compile(r"\b(int|double|float|char|bool|void|short|long|size_t|class)\s+([a-zA-Z_][a-zA-Z0-9_]*)\b").

[0055] The above expression matches common variable types (such as int, double, float, etc.) and captures the variable name.

[0056] Among them, \b: This is a word boundary marker. It does not match any actual characters, but indicates a position - the characters before and after this position are not part of the same word. Specifically, \b is often used to ensure that the match is a whole word rather than a part of a word. For example, when matching a variable name, using \b can avoid mistaking part of a variable name for part of another variable name.

[0057] (int|double|float|char|bool|void|short|long|size_t|class): This is a capture group defined, which contains multiple keywords separated by pipe characters |. Common data types in C / C++ and the keyword class are listed here. The vertical bar | means "or" in regular expressions, so this capture group will match any of the listed data types or class keywords.

[0058] \s+: This represents one or more whitespace characters. Whitespace characters include space, tab, etc. The plus sign + here means that the preceding element (in this case, whitespace characters) occurs at least once.

[0059] ([a-zA-Z_][a-zA-Z0-9_]*): This is another capturing group, used to match variable or class names. The first character must be a letter (uppercase or lowercase) or an underscore, which is represented by [a-zA-Z_]. The following part can contain zero or more letters, numbers, or underscores, which is represented by [a-zA-Z0-9_]*. The asterisk * means that the previous element (in this case, a set of letters, numbers, or underscores) can appear zero or more times.

[0060] For another example, Python is a dynamic language that determines data types at runtime. No type declaration is required before using a variable. For example, in Python, if variable a=1, then the type of a is integer. If a="abc", the type of a is string. Therefore, it is relatively responsible when implementing obfuscation.

[0061] To this end, the server can use a variety of regular expressions to match class, function, and variable names. In Python, when defining classes and functions, the iconic keywords class and def are prepended, so the C / C++ matching mode can be used; variable declarations are more obscure, and given the characteristics of the Python programming language, this manual treats the left side of the assignment statement as a variable.

[0062] Therefore, the regular expression used to match Python classes and functions is as follows: class_func_regex = re.compile(r"\b(class|def)\s+([a-zA-Z_][a-zA-Z0-9_]*)") Among them, \b: This is a word boundary marker, which ensures that the match is an independent word rather than a part of a word. For example, when matching class or def, using \b can avoid mistaking these keywords for parts of other words.

[0063] (class|def): This is a capture group, which contains two options separated by a vertical bar |, indicating an "or" relationship. Here it means matching one of the two keywords class or def. Class is used to define a class, while def is used to define a function.

[0064] \s+: This represents one or more whitespace characters. Whitespace characters include space, tab, etc. The plus sign + here means that the preceding element (in this case, whitespace characters) occurs at least once.

[0065] ([a-zA-Z_][a-zA-Z0-9_]*): This is another capture group, used to match class names or function names. The specific rules are as follows: [a-zA-Z_]: Indicates that the first character must be a letter (uppercase or lowercase) or an underscore.

[0066] [a-zA-Z0-9_]*: Indicates that the following characters can contain zero or more letters, numbers, or underscores. The asterisk * means that the previous element (in this case, a set of letters, numbers, or underscores) can appear zero or more times.

[0067] In particular, Python supports multiple value returns, so the regular expression needs to match all results on the left side of the assignment statement. The regular expression used to match multi-variable assignments is as follows: multi_assign_regex=re.compile( r"\b([a-zA-Z_][a-zA-Z0-9_]*)\s*,\s*([a-zA-Z_][a-zA-Z0-9_]*)(?:\s*,\s*([a-zA-Z_][a-zA-Z0-9_]*))*\s*=\s*([^=\n]+)" ) The explanation of each symbol in this regular expression refers to the symbol explanation of the regular expression above, which will not be described in detail here.

[0068] In addition, after obtaining the sample code, the server can extract the vulnerability code, repair code and context code contained in the sample code and execute the following operations: Figure 3 The storage structure shown is used for storage. Then, when the server performs the modification operation of the above non-critical identifier, it is necessary to first merge the vulnerability code, repair code and context code corresponding to the same code identifier before modifying the non-critical identifier. This is mainly to ensure that the name of each code snippet remains consistent during obfuscation.

[0069] On this basis, the server can store the code snippets contained in the desensitized code in a preset database, such as Figure 4 shown.

[0070] Figure 4 A schematic diagram of the storage structure used to store each code fragment in the desensitized code provided in this specification.

[0071] from Figure 4 As can be seen in the figure, the code snippet refers to the code snippet modified by the non-critical identifier (i.e., desensitized). The server can calculate the hash value of the actual code snippet modified by the non-critical identifier to obtain the code snippet identifier of the code snippet. The code identifier refers to the identification information of the sample code to which the code snippet belongs. The code category is used to indicate whether the code snippet is a vulnerability code, a repair code, a context upper segment, or a context lower segment. The code refers to the desensitized code snippet itself.

[0072] S103: For each desensitized code, adjust the desensitized code according to a preset adjustment strategy to obtain at least one enhanced code corresponding to the desensitized code.

[0073] In order to further enhance the vulnerability identification and repair capabilities of the vulnerability detection model, after obtaining the above-mentioned desensitized codes, the server can adjust the desensitized codes according to the preset adjustment strategy to obtain the enhanced codes.

[0074] Among them, this specification mainly introduces two adjustment strategies. One is to adjust the control flow in the code. The so-called control flow is used to represent the execution order and method of the code statements in the code. Then, adjusting the control flow means changing the execution order and method of the code statements in the above desensitized code. Since a desensitized code may involve multiple control flows, after the server adjusts the control flow in the desensitized code, at least one enhanced code can be obtained.

[0075] In this specification, the server can use a large language model to adjust the control flow in the above desensitized code. For example, the server can use the following prompt word template: <lang> \n{programming_language}\n< / lang> \n <code> \n{code_snippet}\n< / code> ,\n <code>The code in the tag is <lang>Written in the programming language, it is required to adjust the control flow in the code (if any) or replace it with an equivalent expression without changing the semantics and affecting the original logic execution, such as: changing the conditional branch if(x>0) to if(!(x<=0)), replacing a+b with b+a or replacing ++i in the loop with i+=1. Note that only the modified code is output.

[0076] {programming_language} and {code_snippet} are placeholders, which are filled in when generating the large language model prompt.

[0077] The server can use the above-mentioned prompt word template to input the desensitized code into the large language model. The large language model will perform intent recognition based on the content of the prompt word template, and then adjust the control flow of the obtained desensitized code to obtain at least one enhanced code and output it.

[0078] In addition, the server may also adjust the desensitized code by means of code expansion. The so-called code expansion is to add or subtract redundant codes (ie, increase redundant noise) in the desensitized code, thereby obtaining at least one enhanced code.

[0079] The server can also use a large language model to complete the code expansion process. For example, the server can use the following prompt word template: <lang> \n{programming_language}\n< / lang> \n <code> \n{code_snippet}\n< / code> ,\n <code>The code in the tag is <lang>It is required to write in the programming language without changing the semantics and affecting the original execution logic. <code>Add 10 to 20 lines of redundant code to the output. Note that only the modified code is output.

[0080] The prompt words in the prompt word template have been explained in the above content and will not be described in detail here.

[0081] In this specification, the server can use the method of adjusting the control flow alone to obtain at least one enhanced code, or can use the method of code expansion alone to obtain at least one enhanced code. Of course, both methods can be used at the same time, that is, the control flow of the desensitized code is adjusted, and the code expansion is also performed on the desensitized code. As for whether to adjust the control flow first or to expand the code first, it can be determined according to actual needs.

[0082] In addition, in practical applications, there may be other adjustment methods to obtain at least one enhanced code, which will not be described in detail here.

[0083] After the server obtains the enhanced code corresponding to each desensitized code in the above manner, the obtained enhanced code can be stored. Since the server can store each code fragment in a preset database in advance, the enhanced code can also be obtained for each code fragment, that is, for a code fragment, the server can use the above method to obtain at least one enhanced code corresponding to the code fragment.

[0084] After obtaining the enhanced codes of these code fragments, the server can store these enhanced codes according to a certain data storage structure, such as Figure 5 shown.

[0085] Figure 5 A schematic diagram of the storage structure of the enhanced code snippet obtained after adjustment provided in this specification.

[0086] exist Figure 5 In the code snippet identifier, the identifier of the code snippet corresponding to the enhanced code snippet can be obtained by the server by calculating the hash value. The code identifier refers to the identification information of the sample code to which the enhanced code snippet belongs. The code refers to the enhanced code snippet itself. The code category is used to indicate whether the enhanced code snippet specifically corresponds to the vulnerability code, repair code, or context code (that is, whether it is the enhanced code snippet obtained by adjusting the vulnerability code, repair code, or context code).

[0087] Figure 5 The adjustment strategy in is used to record the adjustment strategy adopted by the enhanced code snippet, where CFP means that only the adjustment strategy of adjusting the control flow is adopted to obtain the enhanced code snippet; Bloat means that only the adjustment strategy of code expansion is adopted to obtain the enhanced code snippet; All means that both the adjustment strategy of adjusting the control flow and the adjustment strategy of code expansion are adopted; None means that no adjustment strategy is adopted.

[0088] Therefore, in actual applications, the enhanced code snippets finally obtained by the server also contain some codes that have only been desensitized but not adjusted by any adjustment strategy.

[0089] S104: for each code identifier, combine the code fragments belonging to the code identifier to obtain multiple composite codes corresponding to the code identifier, and save them in a preset database.

[0090] After obtaining the above enhanced codes, the server may combine the code fragments in each enhanced code to obtain a plurality of composite codes.

[0091] During the combination process, the server is not limited to combining only the code fragments in the enhanced code, but can also combine the code fragments contained in the enhanced code and the original sample code (as mentioned above, the desensitized code finally obtained by the server also includes the code that has only been desensitized but not adjusted by the adjustment strategy, so the enhanced code here actually includes the code that has only been desensitized and the code that has been adjusted by the adjustment strategy) to obtain multiple composite codes.

[0092] In the specific combination process, the server can actually combine a code snippet in a code (including enhanced code and sample code) with one or more code snippets contained in other codes to obtain a composite code. In the entire combination process, it is necessary to ensure that the code snippets contained in the code corresponding to the same code identifier are combined. Moreover, it is necessary to ensure that the code categories of the code snippets contained in the obtained composite code are different. In other words, there will not be two or more vulnerability codes, repair codes or context codes in a composite code.

[0093] Furthermore, in the process of combining code fragments, the server may combine based on the vulnerability code as the core, and based on the repair code as the core. Combining based on the vulnerability code as the core means combining the adjusted vulnerability code with the context code (including adjusted and unadjusted context code), and combining the unadjusted vulnerability code with the context code (also including adjusted and unadjusted context code).

[0094] The following 16 situations can be obtained by combining the vulnerability code as the core: CFP(PC / SC)+None(B), CFP(PC / SC)+CFP(B), CFP(PC / SC)+Bloat(B), CFP(PC / SC)+All(B); Bloat(PC / SC)+None(B), Bloat(PC / SC)+CFP(B), Bloat(PC / SC)+Bloat(B), Bloat(PC / SC)+All(B); All(PC / SC)+None(B), All(PC / SC)+CFP(B), All(PC / SC)+Bloat(B), All(PC / SC)+All(B); None(PC / SC)+None(B), None(PC / SC)+CFP(B), None(PC / SC)+Bloat(B), None(PC / SC)+All(B).

[0095] Among them, CFP means that only the adjustment strategy of adjusting control flow is used to adjust the code snippet, Bloat means that only the adjustment strategy of code expansion is used to adjust the code snippet; All means that both the adjustment strategy of adjusting control flow and the adjustment strategy of code expansion are used; None means that no adjustment strategy is used.

[0096] PC represents the upper context, SC represents the lower context, and B represents the vulnerability code.

[0097] Taking the combination of CFP(PC / SC)+None(B) as an example, it means the composite code obtained by combining the context upper segment and context lower segment adjusted by the control flow adjustment strategy with the vulnerability code that has not been adjusted by any adjustment strategy. The remaining combination forms will not be described in detail here.

[0098] Similarly, the server can also get the following 16 situations by combining the repair code as the core: CFP(PC / SC)+None(G), CFP(PC / SC)+CFP(G), CFP(PC / SC)+Bloat(G), CFP(PC / SC)+All(G); Bloat(PC / SC)+None(G),Bloat(PC / SC)+CFP(G),Bloat(PC / SC)+Bloat(G), Bloat(PC / SC)+All(G); All(PC / SC)+None(G), All(PC / SC)+CFP(G), All(PC / SC)+Bloat(G), All(PC / SC)+All(G); None(PC / SC)+None(G),None(PC / SC)+CFP(G),None(PC / SC)+Bloat(G), None(PC / SC)+All(G).

[0099] Among them, PC, SC, CFP, Bloat, All, and None have been explained and will not be repeated here, and G means repair code.

[0100] Taking the combination of Bloat(PC / SC)+None(G) as an example, it represents a composite code obtained by combining the context upper segment and context lower segment with code expansion and the repair code without any adjustment.

[0101] It should be noted that in order to improve the vulnerability identification capability of the vulnerability detection model, the server can further count the line number intervals of the vulnerability code in the composite code, so that in the subsequent adjustment process of the vulnerability detection model, the vulnerability detection model can identify the line number of the vulnerability code contained in the input code and adjust the vulnerability detection model according to the actual line number of the vulnerability code.

[0102] To this end, in this specification, the server can count the vulnerability code line number interval, and specifically the following function can be used for counting: Vulnerable code line starts: len(PC) End line of the vulnerability code: len(PC) + len(Bad) The len() function is used to count the code line numbers.

[0103] In addition, the server can also mark the line numbers of the compound code. According to different programming languages, C / C++ adds comments at the end of each line of code: / / Line x; Python adds comments at the end of each line: # Line x.

[0104] Furthermore, the server may store the composite code obtained above according to a certain storage structure, such as Figure 6 shown.

[0105] Figure 6 A schematic diagram of a storage structure of a composite code provided in this specification.

[0106] exist Figure 6 In the composite code identifier, code identifier, code in Figure 6 The explanations in the Figure 6 From the above composite code, it can be seen that if the composite code is obtained by combining with the vulnerability code as the core, then the composite code contains the vulnerability code, and if the composite code is obtained by combining with the repair code as the core, then the composite code does not contain the repair code.

[0107] Therefore, the code type is used to indicate whether it is a composite code obtained by combining the vulnerability code as the core or a composite code obtained by combining the repair code as the core. When the code type is Bad, it indicates that the composite code contains the vulnerability code, and when the code type is Good, it indicates that there is no vulnerability code in the composite code.

[0108] Correspondingly, for Figure 6 The server can query the specific vulnerability type (such as the above-mentioned out-of-bounds write, out-of-bounds read, improper input validation, etc.) from the original sample code set according to the code identifier. If the composite code does not contain the vulnerability code, the server will not be able to query the corresponding vulnerability type from the original sample code set through the code identifier. At this time, the vulnerability type of the composite code is None.

[0109] Furthermore, for Figure 6 The same is true for the vulnerability line number in . If the composite code contains the vulnerability code, the line number where the vulnerability code is located is recorded here. If the composite code does not contain the vulnerability code, the line number where the vulnerability code is located is not recorded here.

[0110] Figure 6 The "list of corresponding identifiers of reference code snippets" in the above statement refers to the fact that the composite code contains enhanced code snippets. Figure 6 The storage structure shown can record the identifier corresponding to the enhanced code snippet. Since the composite code may contain multiple enhanced code snippets, Figure 6 The storage structure in records a list consisting of multiple identifiers.

[0111] Figure 6 The enhancement type in indicates the specific adjustment strategy used to obtain the enhanced code contained in the composite code. For example, for the combination of CFP(PC / SC)+None(B), the enhancement type used is: the context upper segment and context lower segment are adjusted using the adjustment strategy of adjusting the control flow, and the vulnerability code is not adjusted using any adjustment strategy.

[0112] It should be noted that the reason why Figure 6 The storage structure shown records various identifiers, mainly indicating the sources of the code snippets (including enhanced code snippets) contained in the composite code, so as to facilitate data management and subsequent data maintenance.

[0113] S105: constructing an evaluation set according to the composite code corresponding to each code identifier stored in the preset database, and adjusting a preset vulnerability detection model through the evaluation set, so as to perform code vulnerability detection through the adjusted vulnerability detection model.

[0114] The server can save the above composite codes in a preset database, and then build an evaluation set for adjusting the vulnerability detection model based on these composite codes.

[0115] Among them, the evaluation set constructed in this manual is mainly to improve the vulnerability identification and repair capabilities of the vulnerability detection model. Then, the construction of the evaluation set can be achieved from the following four angles, that is, by generating four different types of evaluation samples to construct the evaluation set. These four different types of evaluation samples can enhance the capabilities of the vulnerability detection model from different angles.

[0116] The first test sample: In this specification, the server can, for each code identifier, select a composite code from multiple composite codes corresponding to the code identifier as the question stem code, select at least part of the composite codes from other composite codes corresponding to the code identifier, and select at least one composite code from composite codes corresponding to other code identifiers.

[0117] Then, the server can construct a first evaluation sample based on the above-mentioned question stem code, at least part of the compound codes selected from other compound codes corresponding to the code identifier, at least one compound code selected from compound codes corresponding to other code identifiers, and a preset prompt word template, and then in the subsequent process, construct an evaluation set through the first evaluation sample.

[0118] From the above process of constructing the first evaluation sample, it can be seen that the server subsequently uses the first evaluation sample to adjust the vulnerability detection model, mainly to allow the vulnerability detection model to identify codes that are not semantically relevant to the question code. For the first evaluation sample, at least part of the compound codes selected from other compound codes corresponding to the code identifier belong to codes that are semantically relevant to the question code, and at least one compound code selected from compound codes corresponding to other code identifiers belongs to codes that are not semantically relevant to the question code.

[0119] In actual applications, the server can generate the first evaluation sample through the following prompt word template instance: <code>{Q}< / code> <options>[ <option> i. {S[i]}< / option> for i in range(N)]< / options> , you are a professional code analysis assistant who can understand and explain complex code logic. Now there is a long code <code>, please <options>Find the option with the least semantic relevance. Note that only the correct option that you think is suitable for the question should be output, and any content other than the option should not be output.

[0120] Among them, {Q}, {S[i]} are placeholders, [ <option> i. {S[i]}< / option> for i in range(N)] means generating N options.

[0121] The server can generate a large batch of first evaluation samples by repeatedly using the above examples.

[0122] The second test sample: For each code identifier, the server can select a composite code from multiple composite codes corresponding to the code identifier, and use the vulnerability code contained in the composite code as the question stem code. Then, based on the code identifier, the server can determine the repair code corresponding to the question stem code, and obtain at least part of the vulnerability code contained in other composite codes corresponding to the code identifier, as well as obtain code fragments contained in composite codes corresponding to other code identifiers.

[0123] Afterwards, the server can construct a second evaluation sample based on the question code, the repair code corresponding to the question code, at least part of the vulnerability code obtained from other composite codes corresponding to the code identifier, code snippets contained in the composite code corresponding to other code identifiers, and a preset prompt word template, so as to construct an evaluation set based on the second evaluation sample in a subsequent process.

[0124] From the above process of constructing the second evaluation sample, it can be seen that the server subsequently uses the second evaluation sample to adjust the vulnerability detection model, mainly to allow the vulnerability detection model to identify the repair code used to repair the question code from a large number of codes. For the second evaluation sample, the repair code corresponding to the question code is the correct option for repairing the question code, and the others are wrong options. In this way, after adjusting the vulnerability detection model using the second evaluation sample, the vulnerability detection model will have the ability to repair the code, and then can repair the identified vulnerability code in actual applications.

[0125] In actual applications, the server can generate the second evaluation sample through the following prompt word template instance: <code>{Q}< / code> <options>[ <option> i. {G[i]}< / option> for i in range(N)]< / options> , you are a professional code analysis assistant who can understand and explain complex code logic. Now there is a long piece of code with vulnerabilities <code>, please <options>Find the option to fix the vulnerability code in the question. Note that only the correct options that you think are suitable for the question are output, and any content other than the options is prohibited.

[0126] Each prompt word in the prompt word template has been explained in the previous example and will not be repeated here. The server can reuse the prompt word template to obtain a large number of second evaluation samples.

[0127] It should be noted that in order to further increase the perplexity of the vulnerability detection model, the adjustment strategy adopted for the context code in the correct option (i.e., the repair code corresponding to the question code) and the adjustment strategy adopted for the context code in the vulnerability code with the same code identifier as the question code can be required to be consistent with the adjustment strategy adopted in the question code. For example, assuming that the adjustment strategy adopted in the question code is CFP (PC / SC), then the composite code to which the repair code and the vulnerability code with the same code identifier as the question code belong also needs to adopt CFP (PC / SC).

[0128] In addition, in order to further increase the confusion, it is possible to require that the vulnerability code with the same code identifier as the question code adopts a different adjustment strategy from the question code. For example, assuming that the question code adopts the adjustment strategy of CFP (PC / SC), then for the composite code to which the vulnerability code with the same code identifier as the question code belongs, only CFP (PC / SC) + None (B), CFP (PC / SC) + Bloat (B) and CFP (PC / SC) + All (B) can be selected.

[0129] The third evaluation sample: For each code identifier, the server can select a vulnerability code contained in a composite code from multiple composite codes corresponding to the code identifier as the question code, read the code line number corresponding to the question code from the preset database, and randomly generate a number of code line numbers for the question code. Then, the server can generate a third evaluation sample based on the question code, the code line number corresponding to the question code read from the preset database, the number of code line numbers randomly generated for the question code, and the preset prompt word template, so as to construct an evaluation set based on the third evaluation sample in the subsequent process.

[0130] From the third test sample mentioned above, it can be seen that the server subsequently uses the third test sample to adjust the vulnerability detection model, mainly to enable the vulnerability detection model to accurately locate the location of the vulnerability code in the code, that is, to accurately determine the line number of the code line where the vulnerability code is located. For the third test sample, the line number corresponding to the question code read from the preset database is the correct option, and the randomly generated line number is the wrong option.

[0131] Among them, it is necessary to ensure that the code line number corresponding to the vulnerability code read from the preset database does not intersect with several randomly generated code line numbers in the line number range, and the code where the starting and ending line numbers are located is not empty.

[0132] In actual applications, the server can generate the third evaluation sample through the following prompt word template instance: <code>{Q}< / code> <options>[ <option> i. {L[i]}< / option> for i in range(N)]< / options> , you are a professional code analysis assistant who can understand and explain complex code logic. Now there is a long piece of code with vulnerabilities <code>, please <options>Find the option of the line number where the vulnerability is located. Note that only the correct options that you think meet the question are output, and any content other than the options is prohibited.

[0133] Each prompt word in the prompt word template has been explained in the previous example and will not be repeated here. The server can reuse the prompt word template to obtain a large number of third evaluation samples.

[0134] The fourth evaluation sample: For each code identifier, the server can select a vulnerability code contained in a composite code from multiple composite codes corresponding to the code identifier as the question code, read the vulnerability type information corresponding to the question code from the preset database, and randomly obtain vulnerability type information of several other vulnerability types. Then, the server can generate a fourth evaluation sample based on the question code, the vulnerability type information corresponding to the question code read from the preset database, the vulnerability type information of several other vulnerability types obtained randomly, and the preset prompt word template, so as to construct an evaluation set in the subsequent process based on the generated fourth evaluation sample.

[0135] From the fourth evaluation sample mentioned above, it can be seen that the server subsequently uses the fourth evaluation sample to adjust the vulnerability detection model, mainly to enable the vulnerability detection model to identify the vulnerability type corresponding to the question code. For the fourth evaluation sample, the vulnerability type information corresponding to the question code read from the preset database is the correct option, while the vulnerability type information of several other vulnerability types obtained randomly is a misplaced option.

[0136] In actual applications, the server may generate the fourth evaluation sample through the following prompt word template instance: <code>{Q}< / code> <options>[ <option> i. {T[i]}< / option> for i in range(N)]< / options> , you are a professional code analysis assistant who can understand and explain complex code logic. Now there is a long piece of code with vulnerabilities <code>, please <options>Select the option corresponding to the vulnerability type. Note that only the correct options that you think are suitable for the question will be output, and any content other than the options will not be output.

[0137] Each prompt word in the prompt word template has been explained in the previous example and will not be repeated here. The server can reuse the prompt word template to obtain a large number of fourth evaluation samples.

[0138] It can be seen from the four evaluation samples generated by the above server that the first evaluation sample is used to improve the vulnerability code identification capability of the vulnerability detection model, the second evaluation sample is used to improve the code repair capability of the vulnerability detection model, the third evaluation sample is used to improve the vulnerability code location capability of the vulnerability detection model, and the fourth evaluation sample is used to improve the vulnerability type identification capability of the vulnerability detection model.

[0139] Through the above-mentioned evaluation samples, an evaluation set can be constructed. Through the constructed evaluation set, the capability of the vulnerability detection model can be enhanced from different angles, thereby providing a basis for accurate vulnerability identification and repair of the vulnerability detection model in practical applications.

[0140] Further, in this specification, after constructing the above evaluation set, the server can store the evaluation samples in the evaluation set, which can be specifically done as follows: Figure 7 The storage structure shown is implemented.

[0141] Figure 7 A schematic diagram of the storage structure of an evaluation set provided in this specification.

[0142] exist Figure 7 In , the evaluation sample identifier can be obtained by calculating the hash value, and the evaluation sample type is used to mark whether the evaluation sample in an evaluation set belongs to the first evaluation sample, the second evaluation sample, the third evaluation sample or the fourth evaluation sample.

[0143] The code identifier is used to indicate the code identifier corresponding to the question code in the assessment sample, and the assessment sample text refers to the text obtained after filling in the text according to the preset prompt word template. Figure 7 The storage structure shown also records the correct answer options corresponding to the assessment samples.

[0144] It should be noted that, in order to facilitate retrieval and management, the database storing various data in this specification may be a relational database.

[0145] The server can adjust the vulnerability detection model through the evaluation set constructed in the above manner, wherein the adjustment adopted can be an existing conventional model adjustment method, and this specification does not limit the specific adjustment method. Once the adjustment is completed, the server can deploy the vulnerability detection model to perform the code vulnerability detection task through the vulnerability detection model.

[0146] It can be seen from the above method that after obtaining the sample code set, the annotation information in each sample code contained in the sample code set can be deleted. This ensures that the subsequently constructed evaluation set will not leak the software code semantics to the vulnerability detection model, thereby enabling the vulnerability detection model to output results that are consistent with its actual vulnerability identification capabilities.

[0147] Moreover, by adjusting and modifying the control flow and non-critical identifiers of the sample code, the final evaluation sample can provide a certain degree of recognition difficulty for the vulnerability detection model, so that after the model adjustment stage, the capability of the vulnerability detection model can be significantly enhanced.

[0148] Furthermore, by adjusting the desensitized code, multiple composite codes can be obtained. Through these composite codes, a variety of evaluation samples can be obtained, which can enhance the vulnerability detection model's ability to identify and locate code vulnerabilities, analyze vulnerability types, and repair code from multiple angles. In subsequent practical applications, the recognition accuracy of the vulnerability detection model and the code repair effect can be significantly improved.

[0149] The above is a code vulnerability detection method provided by one or more embodiments of the present application. Based on the same idea, the present application also provides a corresponding code vulnerability detection device, such as Figure 8 shown.

[0150] Figure 8 A schematic diagram of a code vulnerability detection device provided in an embodiment of the present application specifically includes: The acquisition module 801 is used to acquire a sample code set, wherein the sample code set includes various sample codes, and for each sample code, the sample code corresponds to a code identifier, and the sample code includes various code fragments, and the code category of each code fragment includes: vulnerability code, repair code corresponding to the vulnerability code, and context code of the vulnerability code; An identification module 802 is used to identify and delete the annotation information in each sample code to obtain a desensitized code corresponding to each sample code; The adjustment module 803 is used to adjust each desensitization code according to a preset adjustment strategy to obtain at least one enhanced code corresponding to the desensitization code, wherein the at least one enhanced code is the same as the code identifier of the desensitization code; A combining module 804 is used to combine the code fragments belonging to each code identifier to obtain multiple composite codes corresponding to the code identifier, and store them in a preset database. The code categories of the code fragments included in a composite code are different. The construction module 805 is used to construct an evaluation set according to the composite code corresponding to each code identifier stored in the preset database, and adjust the preset vulnerability detection model through the evaluation set to perform code vulnerability detection through the adjusted vulnerability detection model.

[0151] Optionally, the adjustment module 803 is specifically used to adjust the control flow of each desensitized code to obtain at least one enhanced code corresponding to the desensitized code, and the control flow is used to represent the execution order and method of code statements in the desensitized code.

[0152] Optionally, the adjustment module 803 is specifically configured to, for each desensitized code, add a redundant code to the desensitized code to obtain at least one enhanced code corresponding to the desensitized code.

[0153] Optionally, the construction module 805 is specifically used to, for each code identifier, select a compound code from multiple compound codes corresponding to the code identifier as the question stem code, select at least part of the compound codes from other compound codes corresponding to the code identifier, and select at least one compound code from the compound codes corresponding to other code identifiers; construct a first evaluation sample according to the question stem code, at least part of the compound codes selected from other compound codes corresponding to the code identifier, at least one compound code selected from the compound codes corresponding to other code identifiers, and a preset prompt word template, so as to construct an evaluation set according to the first evaluation sample, and the first evaluation sample is used to enable the vulnerability detection model to identify codes that are semantically irrelevant to the question stem code.

[0154] Optionally, the construction module 805 is specifically used to, for each code identifier, select a composite code from multiple composite codes corresponding to the code identifier, and use the vulnerability code contained in the composite code as the question stem code, determine the repair code corresponding to the question stem code according to the code identifier, obtain at least part of the vulnerability code contained in other composite codes corresponding to the code identifier, and obtain code fragments contained in the composite code corresponding to other code identifiers; construct a second evaluation sample according to the question stem code, the repair code corresponding to the question stem code, at least part of the vulnerability code obtained from other composite codes corresponding to the code identifier, obtain code fragments contained in the composite code corresponding to other code identifiers, and a preset prompt word template, so as to construct an evaluation set according to the second evaluation sample, and the second evaluation sample is used to enable the vulnerability detection model to select a repair code for repairing the question stem code.

[0155] Optionally, for each composite code, the preset database also stores the code line number corresponding to the vulnerability code contained in the composite code; The construction module 805 is specifically used to, for each code identifier, select a vulnerability code contained in a composite code from multiple composite codes corresponding to the code identifier as the question stem code, read the code line number corresponding to the question stem code from the preset database, and randomly generate a number of code line numbers for the question stem code; generate a third evaluation sample according to the question stem code, the code line number corresponding to the question stem code read from the preset database, a number of code line numbers randomly generated for the question stem code, and a preset prompt word template, so as to construct an evaluation set according to the third evaluation sample, and the third evaluation sample is used to enable the vulnerability detection model to determine the code line number of the question stem code.

[0156] Optionally, for each composite code, the preset database also stores vulnerability type information corresponding to the vulnerability code contained in the composite code; The construction module 805 is specifically used to, for each code identifier, select a vulnerability code contained in a composite code from multiple composite codes corresponding to the code identifier as a question stem code, read the vulnerability type information corresponding to the question stem code from the preset database, and randomly obtain vulnerability type information of several other vulnerability types; generate a fourth evaluation sample according to the question stem code, the vulnerability type information corresponding to the question stem code read from the preset database, the vulnerability type information of several other vulnerability types obtained randomly, and a preset prompt word template, so as to construct an evaluation set according to the fourth evaluation sample, and the fourth evaluation sample is used to enable the vulnerability detection model to identify the vulnerability type corresponding to the question stem code.

[0157] The present application also provides a computer-readable storage medium, which stores a computer program, which can be used to execute the above Figure 1 A code vulnerability detection method is provided.

[0158] The present application also provides Fig. 9 The schematic structure diagram of the electronic device shown in FIG. Fig. 9 As shown, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory, and may also include other hardware required for the business. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to achieve the above Figure 1 The code vulnerability detection method described.

[0159] Of course, in addition to software implementation, this application does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is to say, the executor of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.

[0160] In the 1990s, it was very clear whether the improvement of a technology was a hardware improvement (for example, improvements to the circuit structure of diodes, transistors, switches, etc.) or a software improvement (improvement of the method flow). However, with the development of technology, many improvements in the method flow today can be regarded as direct improvements in the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that the improvement of a method flow cannot be implemented with a hardware entity module. For example, a programmable logic device (PLD) (such as a field programmable gate array (FPGA)) is such an integrated circuit whose logical function is determined by the user's programming of the device. Designers can "integrate" a digital system on a PLD by programming themselves, without having to ask chip manufacturers to design and produce dedicated integrated circuit chips. Moreover, nowadays, instead of manually making integrated circuit chips, this kind of programming is mostly implemented by "logic compiler" software, which is similar to the software compiler used when developing and writing programs, and the original code before compilation must also be written in a specific programming language, which is called hardware description language (HDL). There is not only one HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also know that it is only necessary to program the method flow slightly in the above-mentioned hardware description languages ​​and program it into the integrated circuit, and then it is easy to obtain the hardware circuit that implements the logic method flow.

[0161] The controller may be implemented in any suitable manner, for example, the controller may take the form of a microprocessor or processor and a computer-readable medium storing a computer-readable program code (e.g., software or firmware) executable by the (micro)processor, a logic gate, a switch, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller, examples of which include but are not limited to the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320, and the memory controller may also be implemented as part of the control logic of the memory. It is also known to those skilled in the art that, in addition to implementing the controller in a purely computer-readable program code manner, the controller may be implemented in the form of a logic gate, a switch, an application-specific integrated circuit, a programmable logic controller, and an embedded microcontroller by logically programming the method steps. Therefore, such a controller may be considered as a hardware component, and the devices for implementing various functions included therein may also be considered as structures within the hardware component. Or even, the devices for implementing various functions may be considered as both software modules for implementing the method and structures within the hardware component.

[0162] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0163] For the convenience of description, the above device is described by dividing it into various units according to its functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0164] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0165] This specification is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0166] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0167] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0168] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0169] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0170] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0171] It should also be noted that the term "includes", "comprising" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of further restrictions, an element defined by the sentence "comprising a ..." does not exclude the existence of other identical elements in the process, method, commodity or device including the element.

[0172] It should be understood by those skilled in the art that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0173] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0174] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0175] The above description is only an embodiment of the present specification and is not intended to limit the present specification. For those skilled in the art, the present specification may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present specification shall be included in the scope of the claims of the present specification.< / options> < / code> < / options> < / code> < / options> < / code> < / options> < / code> < / code> < / lang> < / code> < / lang> < / code>

Claims

1. A code vulnerability detection method, characterized in that: include: Acquire a sample code set, wherein the sample code set includes various sample codes, and for each sample code, the sample code corresponds to a code identifier, and the sample code includes various code fragments, and the code category of each code fragment includes: vulnerability code, repair code corresponding to the vulnerability code, and context code of the vulnerability code; Identify and delete the annotation information in each sample code to obtain a desensitized code corresponding to each sample code; For each desensitization code, the desensitization code is adjusted according to a preset adjustment strategy to obtain at least one enhanced code corresponding to the desensitization code, wherein the at least one enhanced code is the same as the code identifier of the desensitization code; For each code identifier, the code fragments belonging to the code identifier are combined to obtain multiple composite codes corresponding to the code identifier, and the composite codes are stored in a preset database, and the code categories of the code fragments contained in a composite code are all different; According to the composite code corresponding to each code identifier stored in the preset database, an evaluation set is constructed, and the preset vulnerability detection model is adjusted through the evaluation set to perform code vulnerability detection through the adjusted vulnerability detection model.

2. The method according to claim 1, characterized in that For each desensitized code, the desensitized code is adjusted according to a preset adjustment strategy to obtain at least one enhanced code corresponding to the desensitized code, specifically including: For each desensitized code, the control flow of the desensitized code is adjusted to obtain at least one enhanced code corresponding to the desensitized code, wherein the control flow is used to represent the execution order and method of the code statements in the desensitized code.

3. The method according to claim 1, characterized in that For each desensitized code, the desensitized code is adjusted according to a preset adjustment strategy to obtain at least one enhanced code corresponding to the desensitized code, specifically including: For each desensitized code, a redundant code is added to the desensitized code to obtain at least one enhanced code corresponding to the desensitized code.

4. The method according to claim 1, characterized in that According to the composite code corresponding to each code identifier stored in the preset database, an evaluation set is constructed, specifically including: For each code identifier, select one compound code from multiple compound codes corresponding to the code identifier as the question stem code, select at least part of the compound codes from other compound codes corresponding to the code identifier, and select at least one compound code from compound codes corresponding to other code identifiers; A first evaluation sample is constructed based on the question stem code, at least part of the compound code selected from other compound codes corresponding to the code identifier, at least one compound code selected from compound codes corresponding to other code identifiers, and a preset prompt word template, and an evaluation set is constructed based on the first evaluation sample, wherein the first evaluation sample is used to enable the vulnerability detection model to identify codes that are semantically irrelevant to the question stem code.

5. The method according to claim 1, characterized in that According to the composite code corresponding to each code identifier stored in the preset database, an evaluation set is constructed, specifically including: For each code identifier, select a composite code from multiple composite codes corresponding to the code identifier, and use the vulnerability code contained in the composite code as the question stem code, determine the repair code corresponding to the question stem code according to the code identifier, obtain at least part of the vulnerability code contained in other composite codes corresponding to the code identifier, and obtain code snippets contained in composite codes corresponding to other code identifiers; A second evaluation sample is constructed based on the question stem code, the repair code corresponding to the question stem code, at least part of the vulnerability code obtained from other composite codes corresponding to the code identifier, code snippets contained in the composite code corresponding to other code identifiers, and a preset prompt word template, and an evaluation set is constructed based on the second evaluation sample, and the second evaluation sample is used to enable the vulnerability detection model to select the repair code for repairing the question stem code.

6. The method according to claim 1, characterized in that For each composite code, the preset database also stores the code line number corresponding to the vulnerability code contained in the composite code; According to the composite code corresponding to each code identifier stored in the preset database, an evaluation set is constructed, specifically including: For each code identifier, a vulnerability code contained in a composite code is selected from multiple composite codes corresponding to the code identifier as the question stem code, the code line number corresponding to the question stem code is read from the preset database, and a number of code line numbers for the question stem code are randomly generated; A third evaluation sample is generated based on the question stem code, the code line number corresponding to the question stem code read from the preset database, a number of code line numbers randomly generated for the question stem code, and a preset prompt word template, so as to construct an evaluation set based on the third evaluation sample, and the third evaluation sample is used to enable the vulnerability detection model to determine the code line number of the question stem code.

7. The method according to claim 1, characterized in that For each composite code, the preset database also stores vulnerability type information corresponding to the vulnerability code contained in the composite code; According to the composite code corresponding to each code identifier stored in the preset database, an evaluation set is constructed, specifically including: For each code identifier, a vulnerability code contained in a composite code is selected from multiple composite codes corresponding to the code identifier as a question stem code, vulnerability type information corresponding to the question stem code is read from the preset database, and vulnerability type information of several other vulnerability types is randomly obtained; A fourth evaluation sample is generated based on the question code, the vulnerability type information corresponding to the question code read from the preset database, the vulnerability type information of several other vulnerability types obtained randomly, and a preset prompt word template, so as to construct an evaluation set based on the fourth evaluation sample, and the fourth evaluation sample is used to enable the vulnerability detection model to identify the vulnerability type corresponding to the question code.

8. A code vulnerability detection device, characterized in that: include: An acquisition module is used to acquire a sample code set, wherein the sample code set includes various sample codes, and for each sample code, the sample code corresponds to a code identifier, and the sample code includes various codes, and the code category of each code fragment includes: a vulnerability code, a repair code corresponding to the vulnerability code, and a context code of the vulnerability code; An identification module, used to identify and delete the annotation information in each sample code to obtain a desensitized code corresponding to each sample code; An adjustment module, configured to adjust each desensitization code according to a preset adjustment strategy to obtain at least one enhanced code corresponding to the desensitization code, wherein the at least one enhanced code is the same as the code identifier of the desensitization code; A combination module is used to combine the code fragments belonging to each code identifier to obtain multiple composite codes corresponding to the code identifier, and store them in a preset database. The code categories of the code fragments contained in a composite code are different. The construction module is used to construct an evaluation set according to the composite code corresponding to each code identifier stored in the preset database, and adjust the preset vulnerability detection model through the evaluation set to perform code vulnerability detection through the adjusted vulnerability detection model.

9. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method described in any one of claims 1 to 7 is implemented.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Source code detection method and device, computer equipment and storage medium

    CN116257856A

  • Code vulnerability detection large model construction method and device and electronic equipment

    CN118171291A

  • Aligning Natural Language to Linking Code Snippets to Perform a Complicated Task

    US20170060540A1

  • Code vulnerability remediation

    US20210124830A1