A Machine Learning-Based Code Vulnerability Detection Method and System
Through a machine learning-based method, combining statement source code and context source code, a code vulnerability detection model is built, which solves the problem of inaccurate detection results in the existing technology that ignores the code context, and achieves higher vulnerability detection accuracy.
Patent Information
- Application Number
- CN202510405325.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-04-02
AI Technical Summary
Existing source code vulnerability detection technologies generally ignore the impact of the code context content on the existence of the vulnerability, resulting in limited accuracy of the detection results.
Using a machine learning-based method, by extracting statement source code and context source code from the source code, combining data acquisition algorithms and machine learning algorithms, a code vulnerability detection model is built and a vulnerability classification label for code statements is predicted.
By taking into account the impact of the code context content on the existence of vulnerabilities, the accuracy of code vulnerability detection results is significantly improved, and the missed rate and resource consumption are reduced.
Smart Images

Figure CN119902962B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data processing, and particularly relates to a code vulnerability detection method and system based on machine learning. Background Art
[0002] Software, as an indispensable part of the information society, plays an increasingly important role. It is not only closely related to individuals' daily lives but also to the development of society. However, software is a double-edged sword. While providing convenient services for individuals and society, the potential vulnerabilities in it may also cause great losses to individuals and society. Vulnerabilities in software are often inevitable. On the one hand, it is difficult to ensure that there are no problems during the design, development, and deployment of software. On the other hand, due to commercial benefits, the software development cycle cannot be too long, further increasing the risk of software having vulnerabilities. In order to reduce the vulnerabilities in software and improve software quality, software vulnerability detection technology has emerged. Vulnerability detection technology is to check the source code of software or the execution process of software, and judge whether the software has vulnerabilities according to experience, known vulnerability patterns, and the execution results of software.
[0003] Source code vulnerability detection technology is an important means to improve software quality and security, reduce the number of vulnerabilities, and lower the cost of vulnerability repair. It can be divided into the following two methods according to whether software needs to be executed: (1) Static analysis, that is, analyzing the source code without running the software to detect code defects. However, this depends on experts to formulate rules, and there are problems such as strong subjectivity and imperfect rules; (2) Dynamic test analysis, that is, executing the target program and monitoring its running status to discover vulnerabilities. However, it has problems such as high false negative rate, large resource consumption, and difficulty in discovering complex vulnerabilities due to low analysis efficiency and low code coverage.
[0004] Currently, although there are existing source code vulnerability detection technologies for static analysis such as "CN116541843A - A Code Vulnerability Detection Method Based on Machine Learning", "CN111259394A - A Fine-Grained Source Code Vulnerability Detection Method Based on Graph Neural Network", and "CN117272309A - A Source Code Vulnerability Detection Method and Device Based on Graph Neural Network and Explainable Model", these technologies generally ignore the impact of code context on the existence of vulnerabilities, resulting in the need to improve the accuracy of code vulnerability detection results. Summary of the Invention
[0005] The purpose of the present invention is to provide a code vulnerability detection method and system based on machine learning to solve the problem that the accuracy improvement of code vulnerability detection results is limited due to the general neglect of the impact of code context on the existence of vulnerabilities in existing source code vulnerability detection technologies.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] In a first aspect, a machine learning-based code vulnerability detection method is provided, including:
[0008] Obtain the source code and extract the statement source code of at least one code statement from the source code;
[0009] For each code statement in the at least one code statement, also extract the corresponding context source code from the source code;
[0010] For each code statement, use a data collection algorithm to collect the historical test sample data from all historical test sample data of the same type of code statements that match the corresponding statement, and add the collection result to the corresponding sample dataset, where the historical test sample data includes the code features of the same type of code statements used as model input items and the historical test post-code vulnerability classification label values of the same type of code statements used as model output items, and the code features refer to the numerical features obtained by converting the statement source code and context source code of the corresponding statement;
[0011] For each code statement, apply the corresponding sample dataset to calibrate and verify the modeling of an artificial intelligence model based on a machine learning algorithm to obtain a corresponding code vulnerability detection model, where the code vulnerability detection model is used to output the predicted value of the test post-code vulnerability classification label of the corresponding statement after inputting the statement source code and context source code of the corresponding statement;
[0012] For each code statement, convert the corresponding statement source code and context source code into the corresponding code features, and then input the code features into the corresponding code vulnerability detection model to output the predicted value of the test post-code vulnerability classification label of the corresponding statement;
[0013] Summarize the predicted values of the test post-code vulnerability classification labels indicating the presence of code vulnerabilities in the corresponding code statements and the code statements corresponding to the predicted values of the test post-code vulnerability classification labels to obtain the vulnerability detection result of the source code and give an alarm reminder.
[0014] Based on the above invention content, a new solution for predicting the classification labels of post - test code vulnerabilities based on statement source code and context source code is provided. That is, first, extract the statement source code and context source code of each code statement from the source code. Then, for each statement, use a data collection algorithm to collect sample data from all historical test sample data corresponding to it, and add the collection result to the corresponding sample data set. Next, for each statement, apply the corresponding data set to calibrate and verify the modeling of an artificial intelligence model based on a machine learning algorithm to obtain the corresponding code vulnerability detection model, and apply the model to predict the corresponding classification label value of the post - test code vulnerability based on the corresponding statement source code and context source code. Finally, summarize the vulnerability detection results and give an alarm reminder. In this way, by taking into account the influence of code context content on the existence of vulnerabilities during the code vulnerability detection process, the accuracy of the code vulnerability detection results can be effectively improved, which is convenient for practical application and promotion.
[0015] In a possible design, the at least one code statement includes a first code statement for outputting text, a second code statement for obtaining user input, a third code statement for executing different code blocks according to different conditions, a fourth code statement for looping through elements in an iterable object, a fifth code statement for continuously looping and executing a code block when a condition is true, a sixth code statement for importing available program modules to use functions / classes in the available program modules, and / or a seventh code statement for temporarily not performing any operation.
[0016] In a possible design, for each of the code statements, using a data collection algorithm to collect the historical test sample data from all historical test sample data of the same - type code statements that match the corresponding statement, and adding the collection result to the corresponding sample data set includes:
[0017] For a certain code statement among the at least one code statement, randomly collect several of the historical test sample data from all historical test sample data of the same - type code statements that match the corresponding statement, and add the several historical test sample data obtained by random collection to the corresponding sample data set which is initially an empty set. Among them, the historical test sample data includes the code features of the same - type code statement used as model input items and the historical classification label values of the post - test code vulnerabilities of the same - type code statement used as model output items. The code features refer to the numerical features converted based on the statement source code and context source code of the corresponding statement.
[0018] Process and analyze the sample data set corresponding to the certain code statement to determine the sample data gap direction corresponding to the certain code statement.
[0019] Configure a data collection algorithm according to the direction of the sample data gap, and use the data collection algorithm to collect the historical test sample data from all historical test sample data of the same type of code statements that match the certain code statement;
[0020] Determine whether the newly collected historical test sample data matches the direction of the sample data gap. If so, add the newly collected historical test sample data to the sample data set corresponding to the certain code statement; otherwise, discard adding the newly collected historical test sample data to the sample data set corresponding to the certain code statement.
[0021] In a possible design, configuring a data collection algorithm according to the direction of the sample data gap, and using the data collection algorithm to collect the historical test sample data from all historical test sample data of the same type of code statements that match the certain code statement includes the following steps S331 to S339:
[0022] S331. According to the direction of the sample data gap, initialize the initial position corresponding to the initial feature value of the code feature at a position around the data gap, and then execute step S332, where the initial position and the position around the data gap are respectively located in the multi-dimensional feature space constructed based on the code feature, and the position around the data gap is determined according to the direction of the sample data gap;
[0023] S332. According to the initial feature value of the code feature, use the direct binary search algorithm to search from all historical test sample data of the same type of code statements that match the certain code statement for the first historical test sample data whose model input item matches the initial feature value, and then execute step S333;
[0024] S333. Initialize the feature sequence number variable to 1, and then execute step S334;
[0025] S334. Based on the current feature value of the code feature, change the feature value of the th feature in the code feature to obtain the new feature value of the code feature, and then execute step S335;
[0026] S335. According to the new feature value of the code feature, use the direct binary search algorithm to search from all historical test sample data of the same type of code statements that match the certain code statement for the second historical test sample data whose model input item matches the new feature value, and then execute step S336;
[0027] S336. Compare the historical post - test code vulnerability classification label values that are in the second historical test sample data and the first historical test sample data and are used as model output items. If the historical post - test code vulnerability classification label value in the second historical test sample data is different from the historical post - test code vulnerability classification label value in the first historical test sample data, then execute step S337; otherwise, execute step S338;
[0028] S337. Update the first historical test sample data to the second historical test sample data, and if the feature sequence number variable is less than the total number of dimensions of the code features, then increment the feature sequence number variable by 1, and then return to execute step S334; otherwise, execute step S339;
[0029] S338. Restore the feature value of the th feature to the value before the change, and if the feature sequence number variable is less than the total number of dimensions of the code features, then increment the feature sequence number variable by 1, and then return to execute step S334; otherwise, execute step S339;
[0030] S339. Take the first historical test sample data as the acquisition result and end this acquisition.
[0031] In a possible design, process and analyze the sample data set corresponding to the certain code statement to determine the sample data gap direction corresponding to the certain code statement, including:
[0032] According to the domain of definition of the post - test code vulnerability classification label values of the same - type code statements that match the certain code statement, perform statistical analysis on the sample data set corresponding to the certain code statement to determine the distribution interval of the post - test code vulnerability classification label values in the sample data set within the domain of definition;
[0033] Exclude the distribution interval from the domain of definition to obtain the distribution missing interval in the domain of definition;
[0034] Determine the sample data gap direction corresponding to the certain code statement according to the distribution missing interval.
[0035] In a possible design, for each code statement, apply the corresponding sample data set to calibrate and verify the modeling of an artificial intelligence model based on a machine learning algorithm to obtain the corresponding code vulnerability detection model, including:
[0036] Obtain multiple artificial intelligence models based on different machine learning algorithms;
[0037] Apply the sample data set corresponding to the statement to a pair of statements and models among the at least one code statement and the multiple artificial intelligence models, optimize the hyperparameters of the corresponding model based on an optimization algorithm, and obtain the corresponding model hyperparameters for minimizing the objective function wherein the objective function has the following calculation formula:
[0038]
[0039] In the formula, represents the output error situation index value of the artificial intelligence model, represents the calculation required duration index value of the artificial intelligence model;
[0040] Select an artificial intelligence model with the minimum objective function from the multiple artificial intelligence models;
[0041] Import the optimization search result of the hyperparameters of the artificial intelligence model and the model parameters obtained during the optimization process and corresponding to the optimization search result into the artificial intelligence model to obtain a code vulnerability detection model corresponding to the code statement in the pair of statements and models, where the code vulnerability detection model is used to output the predicted value of the post-test code vulnerability classification label of the corresponding statement after inputting the statement source code and the context source code of the corresponding statement.
[0042] In a possible design, for a pair of statements and models among the at least one code statement and the multiple artificial intelligence models, apply the sample data set corresponding to the statement, optimize the hyperparameters of the corresponding model based on an optimization algorithm, and obtain the corresponding model hyperparameters for minimizing the objective function The optimization search result includes the following steps S421 to S429:
[0043] S421. For a pair of statements and models among the at least one code statement and the multiple artificial intelligence models, initialize the optimization algorithm parameters corresponding to and including the maximum number of iterations and randomly generate an initial search value array for the set of parameters to be optimized, and then execute step S422, where the set of parameters to be optimized includes all the hyperparameters of the artificial intelligence model in the pair of statements and models, and the initial search value array of the set of parameters to be optimized is expressed as follows:
[0044]
[0045] In the formula, Represents less than or equal to a positive integer represents the total number of parameters in the set of parameters to be optimized represents the th parameter in the set of parameters to be optimized and the corresponding initial search value represents the th parameter in the set of parameters to be optimized and the upper limit of the parameter search space corresponding thereto represents the th parameter in the set of parameters to be optimized and the lower limit of the parameter search space corresponding thereto represents a fractional random generation function
[0046] S422. Import the initial search value array of the set of parameters to be optimized as model hyperparameters into the artificial intelligence model in the pair of statements and the model, obtain the first artificial intelligence model, and then apply the sample data set corresponding to the code statement in the pair of statements and the model to perform model training and testing on the first artificial intelligence model, obtain the first output error situation index value and the first calculation required duration index value, and finally import the first output error situation index value and the first calculation required duration index value into the objective function and use the output result as the fitness corresponding to the initial search value array, and then execute step S423, where the objective function has the following calculation formula
[0047]
[0048] In the formula represents the output error situation index value of the artificial intelligence model represents the calculation required duration index value of the artificial intelligence model
[0049] S423. Use the initial search value array of the set of parameters to be optimized as the current optimal search value array, and initialize the current iteration count , and also initialize the forbidden change countdown value of each parameter in the set of parameters to be optimized to zero, and then execute step S424
[0050] S424. Based on the current optimal search value array, perform independent increment processing and decrement processing on each parameter in the set of parameters to be optimized and with a current forbidden change countdown value of zero, obtain new search value arrays of the set of parameters to be optimized, and then execute step S425, where represents the total number of parameters in the set of parameters to be optimized and with a current forbidden change countdown value of zero, and there is ;
[0051] S425. For each of the arrays in the new search value arrays, import the corresponding array as a model hyperparameter into the artificial intelligence model in the certain pair of statements and model, obtain the second artificial intelligence model of the corresponding array, then apply the sample data set corresponding to the code statement in the certain pair of statements and model to train and test the second artificial intelligence model, obtain the second output error situation index value and the second calculation required duration index value of the corresponding array, and finally import the second output error situation index value and the second calculation required duration index value into the objective function
[0052] and use the output result as the fitness of the corresponding array, and then execute step S426;
[0053] S426. For each of the arrays, subtract the fitness of the corresponding array from the fitness corresponding to the current optimal search value array to obtain the fitness difference value of the corresponding array, and when the fitness difference value is greater than zero and is the largest fitness difference value in this time, update the forbidden change countdown value of the corresponding unique variable parameter to a positive integer value positively correlated with the fitness difference value, and then execute step S427, where the unique variable parameter refers to the parameter in the set of parameters to be optimized and used to obtain the corresponding array through increase processing or decrease processing;
[0054] S427. Determine whether there is any parameter in the set of parameters to be optimized with the current forbidden change countdown value being zero. If not, decrement the current forbidden change countdown value of each parameter in the set of parameters to be optimized by 1, and then return to execute step S427. Otherwise, execute step S428;
[0054] S427. Determine whether there is any parameter in the set of parameters to be optimized with the current forbidden change countdown value being zero. If not, decrement the current forbidden change countdown value of each parameter in the set of parameters to be optimized by 1, and then return to execute step S427. Otherwise, execute step S428; new search value arrays and corresponding to the minimum fitness, and then execute step S429. Otherwise, directly execute step S429;
[0055] S429. Increment the current iteration count by 1, and determine whether the current iteration count reaches the maximum iteration count . If so, use the current optimal search value array as the optimal search result corresponding to the certain pair of statements and model and used to minimize the objective function. Otherwise, return to execute step S424.
[0056] In a possible design, the multiple artificial intelligence models include artificial intelligence models based on graph neural networks, support vector machines, K-nearest neighbor methods, stochastic gradient descent methods, multivariable linear regression, multi-layer perceptrons, decision trees, backpropagation neural networks, and / or radial basis function networks.
[0057] In a second aspect, a machine learning-based code vulnerability detection device is provided, including a statement code extraction unit, a context code extraction unit, a sample data collection unit, a vulnerability detection modeling unit, a detection model application unit, and a detection result summary unit;
[0058] The statement code extraction unit is configured to obtain source code and extract the statement source code of at least one code statement from the source code;
[0059] The context code extraction unit is communicatively connected to the statement code extraction unit and is configured to, for each code statement in the at least one code statement, further extract the corresponding context source code from the source code;
[0060] The sample data collection unit is communicatively connected to the statement code extraction unit and is configured to, for each code statement, collect the historical test sample data from all historical test sample data of the same type of code statements matching the corresponding statement by using a data collection algorithm, and add the collection result to the corresponding sample data set, where the historical test sample data includes the code features of the same type of code statements used as model input items and the historical test post-code vulnerability classification label values of the same type of code statements used as model output items, and the code features refer to the numerical features obtained by converting the statement source code and context source code based on the corresponding statement;
[0061] The vulnerability detection modeling unit is communicatively connected to the sample data collection unit and is configured to, for each code statement, apply the corresponding sample data set to calibrate and verify the modeling of an artificial intelligence model based on a machine learning algorithm to obtain a corresponding code vulnerability detection model, where the code vulnerability detection model is configured to output the predicted value of the test post-code vulnerability classification label of the corresponding statement after inputting the statement source code and context source code of the corresponding statement;
[0062] The detection model application unit is respectively communicatively connected to the statement code extraction unit, the context code extraction unit, and the vulnerability detection modeling unit, and is configured to, for each code statement, convert the corresponding statement source code and context source code into the corresponding code features, and then input the code features into the corresponding code vulnerability detection model to output the predicted value of the test post-code vulnerability classification label of the corresponding statement;
[0063] The detection result summarization unit is communicatively connected to the detection model application unit, and is configured to summarize the predicted values of the post-test code vulnerability classification labels indicating the presence of code vulnerabilities in the corresponding code statements and the code statements corresponding to the predicted values of the post-test code vulnerability classification labels, so as to obtain the vulnerability detection result of the source code and give an alarm reminder.
[0064] In a third aspect, the present invention provides a computer system, including a memory, a processor, and a transceiver that are communicatively connected in sequence. Among them, the memory is used to store computer programs, the transceiver is used to send and receive messages, and the processor is used to read the computer programs and execute the machine learning-based code vulnerability detection method as described in the first aspect or any possible design in the first aspect.
[0065] In a fourth aspect, the present invention provides a computer-readable storage medium, on which instructions are stored. When the instructions are run on a computer, the machine learning-based code vulnerability detection method as described in the first aspect or any possible design in the first aspect is executed.
[0066] In a fifth aspect, the present invention provides a computer program product, including a computer program or instructions. When the computer program or the instructions are executed by a computer, the machine learning-based code vulnerability detection method as described in the first aspect or any possible design in the first aspect is implemented.
[0067] The beneficial effects of the above solution:
[0068] (1) The present invention creatively provides a new solution for predicting the post-test code vulnerability classification labels based on the statement source code and the context source code. That is, first extract the statement source code and the context source code of each code statement from the source code, then for each statement, use the data acquisition algorithm to collect sample data from all the corresponding historical test sample data, and add the collection result to the corresponding sample data set. Then, for each statement, use the corresponding data set to calibrate and verify the artificial intelligence model based on the machine learning algorithm to obtain the corresponding code vulnerability detection model, and apply the model to predict the corresponding post-test code vulnerability classification label prediction value based on the corresponding statement source code and context source code. Finally, summarize the vulnerability detection results and give an alarm reminder. In this way, by taking into account the influence of the code context content on the existence of vulnerabilities in the code vulnerability detection process, the accuracy of the code vulnerability detection results can be effectively improved;
[0069] (2) Optimize the model hyperparameters through the aforementioned new heuristic algorithm steps. During the optimization process, different numbers of temporary lockings of the search results in each best neighborhood search direction can be performed based on different fitness difference values, which is conducive to quickly searching for the optimal model hyperparameters. By using a parameter adjustment method that first increases / decreases the parameter to be adjusted quantitatively and then increases / decreases it randomly, it is also possible to avoid falling into local optimal solutions, further facilitating the rapid and accurate acquisition of the optimal search results and facilitating practical application and promotion. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0071] Figure 1 It is a flowchart of a code vulnerability detection method based on machine learning provided by an embodiment of the present application.
[0072] Figure 2 It is a structural diagram of a code vulnerability detection system based on machine learning provided by an embodiment of the present application.
[0073] Figure 3 It is a structural diagram of a computer system provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0074] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the present invention in combination with the drawings and the description of the embodiments or the prior art. Obviously, the following description of the drawing structure is only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other embodiments can be obtained based on these embodiments. It should be noted here that the description of these embodiment modes is used to help understand the present invention, but does not constitute a limitation to the present invention.
[0075] It should be understood that although terms such as first and second may be used herein to describe various objects, these objects should not be limited by these terms. These terms are only used to distinguish one object from another. For example, the first object can be called the second object, and similarly, the second object can be called the first object, without departing from the scope of the exemplary embodiments of the present invention.
[0076] It should be understood that for the term "and / or" that may appear in this text, it is merely a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, B exists alone, or both A and B exist simultaneously. Another example, A, B, and / or C can represent any one of A, B, and C or any combination of them. For the term " / and" that may appear in this text, it is a description of another association object relationship, indicating that two relationships can exist. For example, A / and B can represent two situations: A exists alone or both A and B exist simultaneously. In addition, for the character " / " that may appear in this text, it generally indicates that the associated objects before and after are in an "or" relationship.
[0077] Embodiment
[0078] As Figure 1 shown, the code vulnerability detection method provided in the first aspect of this embodiment and based on machine learning can be, but is not limited to, executed by a computer device with certain computing resources, such as a cloud server, a personal computer (Personal Computer, PC, referring to a multi-purpose computer suitable for personal use in terms of size, price, and performance; desktop computers, laptops, small laptops, tablet computers, and ultrabooks all belong to personal computers), a smart phone, a personal digital assistant (Personal Digital Assistant, PDA), or a wearable device and other electronic devices. As Figure 1 shown, the code vulnerability detection method can be, but is not limited to, including the following steps S1 to S6.
[0079] S1. Obtain the source code and extract the statement source code of at least one code statement from the source code.
[0080] In the step S1, the source code is the object for code vulnerability detection, which generally but not limitedly consists of multiple code statements, etc. Therefore, the statement source code of the at least one code statement can be conventionally extracted from the source code. The code statement is the smallest unit for code vulnerability detection in this embodiment, which can consist of one piece of code or multiple pieces of code; specifically, the at least one code statement includes but not limited to a first code statement for outputting text (such as a print statement), a second code statement for obtaining user input (such as an input statement), a third code statement for executing different code blocks according to different conditions (such as an if statement), a fourth code statement for looping through elements in an iterable object (such as a for statement), a fifth code statement for continuously looping through and executing a code block when the condition is true (such as a while statement), a sixth code statement for importing available program modules to use functions / classes in the available program modules (such as an import statement, and the aforementioned available program modules are exemplified by Python modules) and / or a seventh code statement for temporarily not performing any operation (such as a pass statement), etc. In addition, in order to avoid a large amount of noise information in the feature data passed to the subsequent model due to recognition errors, which may interfere with the final judgment result, after obtaining the source code, the identifiers in the source code can be identified with reference to the code standardization technical means in "CN116541843A - A Machine Learning-Based Code Vulnerability Detection Method", and the identifiers are standardized to obtain a standardized source code, and finally the statement source code of the at least one code statement is extracted from the standardized source code.
[0081] S2. For each code statement in the at least one code statement, the corresponding context source code is also extracted from the source code.
[0082] In the step S2, the context source code refers to the source code before and / or after the corresponding code statement, and these source codes can be limited by a preset number of lines (such as extracting the first 10 source codes and / or the last 10 source codes), or determined according to the front-back logical relationship with the corresponding code statement (such as extracting all the previous source codes with the same identifier as the corresponding code statement and / or all the subsequent source codes with the same identifier as the corresponding code statement), and they can also be conventionally extracted.
[0083] S3. For each of the code statements, use a data collection algorithm to collect the historical test sample data from all the historical test sample data of the same type of code statements that match the corresponding statement, and add the collection result to the corresponding sample dataset. Among them, the historical test sample data includes, but is not limited to, the code features of the same type of code statements used as model input items and the historical test post-code vulnerability classification label values of the same type of code statements used as model output items. The code feature refers to a numerical feature obtained by converting the statement source code and the context source code of the corresponding statement.
[0084] In step S3 above, all the historical test sample data of the same type of code statements that match the corresponding statement are exemplified as follows: For a certain code statement, if the corresponding statement is a for statement, the historical test sample data must be the sample data obtained from the historical test of the for statement, rather than the sample data obtained from the historical test of an if statement or other statements. Generally speaking, the data collection algorithm can adopt the direct binary search method (DBSMethod, also known as the binary search algorithm, which is a method for efficiently finding a specific element in an ordered array; its core idea is to locate the position of the target element by continuously narrowing the search range). However, it is also considered that the data distribution for the training of a single model neural network at the current stage must meet certain distribution conditions, that is, the distribution for each result in all the sample data must be relatively uniform. Otherwise, the trained model cannot achieve the basic prediction function. Therefore, preferably, for each of the code statements, use a data collection algorithm to collect the historical test sample data from all the historical test sample data of the same type of code statements that match the corresponding statement, and add the collection result to the corresponding sample dataset, including but not limited to the following steps S31 to S34.
[0085] S31. For a certain code statement among the at least one code statement, randomly collect several of the historical test sample data from all the historical test sample data of the same type of code statements that match the corresponding statement, and add the several historical test sample data obtained by random collection to the corresponding sample dataset that is initially an empty set. Among them, the historical test sample data includes, but is not limited to, the code features of the same type of code statements used as model input items and the historical test post-code vulnerability classification label values of the same type of code statements used as model output items, etc. The code feature refers to a numerical feature obtained by converting the statement source code and the context source code of the corresponding statement.
[0086] In the step S31, the code features can specifically be obtained by converting the statement source code and the context source code of the corresponding statement by using the text vector conversion algorithm in "CN116541843A - A Machine Learning - Based Code Vulnerability Detection Method", or can be obtained by processing the statement source code and the context source code of the corresponding statement by using the code graph feature extraction means in "CN111259394A - A Fine - Grained Source Code Vulnerability Detection Method Based on Graph Neural Network", and can also be obtained by processing the statement source code and the context source code of the corresponding statement by using other existing feature extraction means. In addition, the historical post - test code vulnerability classification label value can specifically be marked by performing historical dynamic test analysis on the statement source code and the context source code of the corresponding statement. And the specific code vulnerability classification label value can include, but is not limited to, a first label value (e.g., represented by "1") for indicating that there is a code vulnerability in the corresponding code statement and a second label value (e.g., represented by "0") for indicating that there is no code vulnerability in the corresponding code statement, which will not be elaborated here.
[0087] S32. Process and analyze the sample data set corresponding to the certain code statement, and determine the sample data gap direction corresponding to the certain code statement.
[0088] In the step S32, specifically, processing and analyzing the sample data set corresponding to the certain code statement to determine the sample data gap direction corresponding to the certain code statement includes, but is not limited to, the following steps S321 - S323.
[0089] S321. According to the domain of the post - test code vulnerability classification label values of the same - type code statements matching the certain code statement, perform statistical analysis on the sample data set corresponding to the certain code statement, and determine the distribution interval of the post - test code vulnerability classification label values in the sample data set within the domain.
[0090] In the step S321, for example, assume that there are the following 5 post - test code vulnerability classification label values for the same - type code statements matching the certain code statement: 4, 3, 2, 1, and 0 (i.e., this zero value is used to indicate that there is no code vulnerability in the corresponding code statement, and other non - zero values are used to indicate different code vulnerability situations in the corresponding code statement). Then the domain can be, for example, {4, 3, 2, 1, 0}. Then, through conventional statistical analysis means, it can be determined, for example, that the distribution interval of the post - test code vulnerability classification label values in the domain is {4, 3} and {0}.
[0091] S322. Eliminate the distribution interval in the domain to obtain the distribution missing interval in the domain.
[0092] In step S322, based on the specific example in step S321 above, the missing distribution interval can be obtained as {1, 0}.
[0093] S323. Determine the sample data gap direction corresponding to a certain code statement according to the missing distribution interval.
[0094] S33. Configure a data acquisition algorithm according to the sample data gap direction, and use the data acquisition algorithm to collect the historical test sample data from all historical test sample data of the same type of code statements that match a certain code statement.
[0095] In step S33, specifically, configuring a data acquisition algorithm according to the sample data gap direction, and using the data acquisition algorithm to collect the historical test sample data from all historical test sample data of the same type of code statements that match a certain code statement includes, but is not limited to, the following steps S331 to S339.
[0096] S331. According to the sample data gap direction, initialize the initial position corresponding to the initial feature value of the code feature at a position around the data gap, and then execute step S332, where the initial position and the position around the data gap are respectively located in the multi-dimensional feature space constructed based on the code feature, and the position around the data gap is determined according to the sample data gap direction.
[0097] In step S331, assume that the multi-dimensional feature space is a three-dimensional space, and the sample data gap direction is the direction pointing to a certain coordinate position in this three-dimensional space. Then, the code feature value closest to this certain coordinate position in the sample data set corresponding to a certain code statement can be used as the position around the data gap and the initial position.
[0098] S332. According to the initial feature value of the code feature, use the direct binary search algorithm to search for the first historical test sample data whose model input item matches this initial feature value from all historical test sample data of the same type of code statements that match a certain code statement, and then execute step S333.
[0099] S333. Initialize the feature sequence number variable to 1, and then execute step S334.
[0100] S334. Based on the current feature value of the code feature, change the feature value of the th feature in the code feature to obtain the new feature value of the code feature, and then execute step S335.
[0101] S335. According to the new eigenvalue of the code feature, use the direct binary search algorithm to search for the second historical test sample data whose model input item matches this new eigenvalue from all historical test sample data of the same type of code statements that match the certain code statement, and then execute step S336.
[0102] S336. Compare the historical test result code vulnerability classification label values used as the model output items in the second historical test sample data and the first historical test sample data. If the historical test result code vulnerability classification label value in the second historical test sample data is different from that in the first historical test sample data, execute step S337; otherwise, execute step S338.
[0103] S337. Update the first historical test sample data to the second historical test sample data, and if the feature sequence number variable is less than the total number of dimensions of the code feature, increment the feature sequence number variable by 1, and then return to execute step S334; otherwise, execute step S339.
[0104] S338. Restore the eigenvalue of the th feature to the value before the change, and if the feature sequence number variable is less than the total number of dimensions of the code feature, increment the feature sequence number variable by 1, and then return to execute step S334; otherwise, execute step S339.
[0105] S339. Take the first historical test sample data as the acquisition result and end this acquisition.
[0106] S34. Determine whether the newly acquired historical test sample data matches the sample data gap direction. If so, add the newly acquired historical test sample data to the sample data set corresponding to the certain code statement; otherwise, do not add the newly acquired historical test sample data to the sample data set corresponding to the certain code statement.
[0107] In the step S34, determining whether the newly collected historical test sample data matches the sample data gap direction may include, but is not limited to: according to the domain of definition of the post-test code vulnerability classification label value of the same type of code statement that matches a certain code statement, performing statistical analysis on the sample data set corresponding to the certain code statement and the newly collected historical test sample data again, determining the new distribution interval of the post-test code vulnerability classification label value in the domain of definition, and then if the new distribution interval expands (which means the missing distribution interval is shrinking at this time), it can be determined that the newly collected historical test sample data matches the sample data gap direction, otherwise it is determined that the newly collected historical test sample data does not match the sample data gap direction.
[0108] S4. For each of the code statements, applying the corresponding sample data set to calibrate and verify the modeling of an artificial intelligence model based on a machine learning algorithm to obtain a corresponding code vulnerability detection model, where the code vulnerability detection model is used to output the predicted value of the post-test code vulnerability classification label of the corresponding statement after inputting the statement source code and the context source code of the corresponding statement.
[0109] In the step S4, the machine learning algorithm is a core artificial intelligence algorithm that specifically studies how a computer simulates or implements human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve its own performance, and is the fundamental way to make a computer intelligent; specifically, the machine learning algorithm preferably but is not limited to using machine learning algorithms based on graph neural networks, support vector machines, K-nearest neighbor methods, stochastic gradient descent methods, multivariable linear regression, multi-layer perceptrons, decision trees, backpropagation neural networks, or radial basis function networks, etc., in order to quickly and accurately find the rules in the data. Thus, based on a certain amount of the sample data set, a verified code vulnerability detection model can be trained through a conventional calibration and verification modeling method (the specific process includes the calibration process and the verification process of the model, that is, first comparing the model simulation results with the measured data, and then adjusting the model parameters according to the comparison results to make the simulation results coincide with the actual situation). Considering that different machine learning algorithms have different prediction accuracies and different calculation times required for prediction, therefore, in order to train corresponding and optimal code vulnerability detection models that can balance detection accuracy and detection speed for different code statements based on different machine learning algorithms, preferably, for each of the code statements, applying the corresponding sample data set to calibrate and verify the modeling of an artificial intelligence model based on a machine learning algorithm to obtain a corresponding code vulnerability detection model, including but not limited to the following steps S41 to S44.
[0110] S41. Obtain multiple artificial intelligence models based on different machine learning algorithms.
[0111] In the step S41, specifically, the multiple artificial intelligence models include but are not limited to artificial intelligence models based on graph neural networks, support vector machines, K-nearest neighbor methods, stochastic gradient descent methods, multivariable linear regression, multi-layer perceptrons, decision trees, backpropagation neural networks, and / or radial basis function networks, etc.; they can be constructed conventionally based on the prior art.
[0112] S42. For a pair of statements and models among the at least one code statement and the multiple artificial intelligence models, apply the sample data set corresponding to the statement, and optimize the hyperparameters of the corresponding model based on an optimization algorithm to obtain the corresponding model hyperparameters and use them to minimize the objective function to obtain the optimal search result for minimizing, where the objective function has the following calculation formula:
[0113]
[0114] In the formula, represents the output error situation index value of the artificial intelligence model, represents the calculation required duration index value of the artificial intelligence model.
[0115] In the step S42, the model hyperparameters refer to the parameters preset in the learning model, and these parameters cannot be directly learned through the data in the standard model training process, but need to be manually set to optimize the performance of the model. For example, there are the number of hidden layer nodes, learning rate, and batch size, etc. The output error situation index value is used to reflect the detection accuracy of the model. The smaller its value, the higher the detection accuracy of the model. It specifically includes but is not limited to the average deviation value, variance, and / or average error rate of the overall sample, etc., and can be routinely statistically obtained during the model training and testing process. The calculation required duration index value is used to reflect the detection speed of the model. The shorter its value, the faster the detection speed of the model, and can be routinely timed during the model testing process. The objective function is used as a comprehensive index that takes into account both detection accuracy and detection speed. The smaller its result, the better the training model can comprehensively achieve the optimal in the two target dimensions of detection accuracy and detection speed, so as to select the best code vulnerability detection model that takes into account both detection accuracy and detection speed subsequently. In addition, specifically, the optimization algorithm can be but is not limited to using particle swarm optimization algorithm, Newton optimization algorithm, genetic optimization algorithm, grey wolf optimization algorithm, whale optimization algorithm, or tuna swarm optimization algorithm, etc.
[0116] S43. Select the one with the smallest objective function from the multiple artificial intelligence models An artificial intelligence model.
[0117] In the step S43, specifically, according to the fitness corresponding to the current optimal search value array of each artificial intelligence model, if it is found that the fitness corresponding to the current optimal search value array of a certain model is the minimum value, then the model is taken as the one with the minimum objective function. An artificial intelligence model.
[0118] S44. Import the optimization search results of the hyperparameters of the artificial intelligence model and the model parameters obtained during the optimization process and corresponding to the optimization search results into the artificial intelligence model to obtain a code vulnerability detection model corresponding to the code statement in the pair of statements and the model, wherein the code vulnerability detection model is used to output the post-test code vulnerability classification label prediction value of the corresponding statement after inputting the statement source code and context source code of the corresponding statement.
[0119] In step S44, the optimization search result of the hyperparameters of the artificial intelligence model is the current optimal search value array of the artificial intelligence model. Also, considering that some hyperparameters need to be integers (such as the number of hidden layer nodes), before the optimization search result of the hyperparameters of the artificial intelligence model is imported into the model as the model hyperparameters, it is necessary to round the corresponding parameter values, such as rounding the hidden layer node values.
[0120] S5. For each code statement, the corresponding statement source code and context source code are converted into the corresponding code features, and then the code features are input into the corresponding code vulnerability detection model, and the corresponding post-test code vulnerability classification label prediction value is output.
[0121] S6. Summarize the post-test code vulnerability classification label prediction value used to indicate that the corresponding code statement has a code vulnerability and the code statement corresponding to the post-test code vulnerability classification label prediction value, obtain the vulnerability detection result of the source code and give an alarm reminder.
[0122] In step S6, for example, if the at least one code statement includes statement A, statement B, statement C, and statement D, and the predicted value of the post-test code vulnerability classification label for statement B is 1, and the predicted value of the post-test code vulnerability classification label for statement C is 2, then the value 1 can be bound to statement B and the value 2 can be bound to statement C, and used as the vulnerability detection result of the source code, and output to the human-computer interaction interface for alarm reminder, so that software developers can repair the vulnerability in time. In addition, in order to achieve the purpose of automatically detecting and patching vulnerabilities, the vulnerability detection result of the source code can also be uploaded to the DeepSeek platform to automatically repair the code vulnerability, and the code vulnerability detection model is reapplied to check the statement source code and context source code of the code statement after vulnerability repair until the output result indicates that there is no predicted value of the post-test code vulnerability classification label indicating the existence of code vulnerability in the corresponding code statement.
[0123] Based on the code vulnerability detection method described in the foregoing steps S1 to S6, a new solution for predicting the post-test code vulnerability classification label based on the statement source code and context source code is provided, that is, first extract the statement source code and context source code of each code statement from the source code, then for each statement, use the data acquisition algorithm to collect sample data from all corresponding historical test sample data, and add the collection result to the corresponding sample data set, and then for each statement, use the corresponding data set to calibrate and verify the artificial intelligence model based on the machine learning algorithm to obtain the corresponding code vulnerability detection model, and apply the model to predict the corresponding post-test code vulnerability classification label prediction value based on the corresponding statement source code and context source code, and finally summarize the vulnerability detection results and give an alarm reminder. In this way, by taking into account the influence of the code context content on the existence of vulnerabilities during the code vulnerability detection process, the accuracy of the code vulnerability detection result can be effectively improved, which is convenient for practical application and promotion.
[0124] Based on the technical solution of the foregoing first aspect, this embodiment also provides a possible design one for optimizing the model hyperparameters based on a new heuristic algorithm, that is, for a pair of statements and models among the at least one code statement and the multiple artificial intelligence models, use the sample data set of the corresponding statement to optimize the hyperparameters of the corresponding model based on the optimization algorithm to obtain the corresponding model hyperparameters and the optimal search result for minimizing the objective function including but not limited to the following steps S421 to S429.
[0125] S421. For a pair of statements and models among the at least one code statement and the multiple artificial intelligence models, initialize the corresponding one that includes the maximum number of iterations The optimized algorithm parameters, and randomly generate an initial search value array for the set of parameters to be optimized, and then execute step S422, where the set of parameters to be optimized includes all hyperparameters of the artificial intelligence model in the pair of statements and the model, and the initial search value array of the set of parameters to be optimized is expressed as follows:
[0126]
[0127] In the formula, represents a positive integer less than or equal to ; represents the total number of parameters in the set of parameters to be optimized, represents the initial search value corresponding to the th parameter in the set of parameters to be optimized, represents the upper limit of the parameter search space corresponding to the th parameter, represents the lower limit of the parameter search space corresponding to the th parameter, represents a pure decimal random generation function.
[0128] S422. Import the initial search value array of the set of parameters to be optimized as model hyperparameters into the artificial intelligence model in the pair of statements and the model to obtain a first artificial intelligence model, and then apply the sample data set corresponding to the code statement in the pair of statements and the model to train and test the first artificial intelligence model to obtain a first output error situation index value and a first calculation required duration index value, and finally import the first output error situation index value and the first calculation required duration index value into the objective function and use the output result as the fitness corresponding to the initial search value array, and then execute step S423, where the objective function has the following calculation formula:
[0129]
[0130] In the formula, represents the output error situation index value of the artificial intelligence model, represents the calculation required duration index value of the artificial intelligence model.
[0131] In step S422, considering that some hyperparameters need to be integers (such as the number of hidden layer nodes), it is necessary to round the corresponding parameter values before importing the initial search value array as model hyperparameters into the model. For example, round the value of the number of hidden layer nodes. The specific process of the foregoing application for calibrating and validating the first artificial intelligence model using the sample data set corresponding to the pair of statements and the code statements in the model is a conventional method, which will not be elaborated here. In addition, when obtaining the first output error situation index value and the first calculation required duration index value, it is also necessary to record the model parameters corresponding to the initial search value array and obtained through model training for use in subsequent step S44.
[0132] S423. Use the initial search value array of the set of parameters to be optimized as the current optimal search value array, and initialize the current iteration count , and also initialize the forbidden change countdown value of each parameter in the set of parameters to be optimized to zero, and then execute step S424.
[0133] S424. Based on the current optimal search value array, perform independent increase and decrease operations on each parameter in the set of parameters to be optimized and with the current forbidden change countdown value of zero, to obtain new search value arrays of the set of parameters to be optimized, and then execute step S425, where represents the total number of parameters in the set of parameters to be optimized and with the current forbidden change countdown value of zero, and there is .
[0134] In step S424, the increase operation or the decrease operation needs to be carried out within the corresponding parameter search space, and can be a quantitative step-by-step increase / decrease, or an indefinite random increase / decrease. It can also first perform a quantitative step-by-step increase / decrease on each parameter in the set of parameters to be optimized and with the current forbidden change countdown value of zero, and then if it is found that the update of the current optimal search value array is not completed (that is, step S429 is directly executed in subsequent step S428), then perform an indefinite random increase / decrease on each parameter in the set of parameters to be optimized and with the current forbidden change countdown value of zero, so as to avoid falling into a local optimal solution. For example, if there are the following five parameters in the set of parameters to be optimized: parameter A, parameter B, parameter C, parameter D, parameter E, and parameter F, where the current forbidden change countdown values of parameter A, parameter B, parameter E, and parameter F are zero respectively (that is If the value is 4), then based on the current optimal search value array, independent increase and decrease processes can be performed on parameter A respectively to obtain two different new search value arrays of the set of parameters to be optimized (one obtained based on the increase process and the other based on the decrease process); based on the current optimal search value array, independent increase and decrease processes can be performed on parameter B respectively to obtain two different new search value arrays of the set of parameters to be optimized; based on the current optimal search value array, independent increase and decrease processes can be performed on parameter E respectively to obtain two different new search value arrays of the set of parameters to be optimized; based on the current optimal search value array, independent increase and decrease processes can be performed on parameter F respectively to obtain two different new search value arrays of the set of parameters to be optimized; thus, 2×4 = 8 new search value arrays can be obtained, and each array represents a neighborhood search direction in the search space.
[0135] S425. For each of the new search value arrays, import the corresponding array as the model hyperparameters into the artificial intelligence model in the certain pair of statements and model, obtain the second artificial intelligence model corresponding to the array, then apply the sample data set corresponding to the code statements in the certain pair of statements and model to train and test the second artificial intelligence model, obtain the second output error situation index value and the second calculation required duration index value corresponding to the array, and finally import the second output error situation index value and the second calculation required duration index value into the objective function and use the output result as the fitness corresponding to the array, and then execute step S426.
[0136] In step S425, the specific technical details can be obtained by conventional derivation in the aforementioned step S422 and will not be elaborated here. In addition, when obtaining the second output error situation index value and the second calculation required duration index value, it is also necessary to record the corresponding data and the model parameters obtained through model training for use in the subsequent step S44.
[0137] S426. For each of the arrays, subtract the fitness corresponding to the array from the fitness corresponding to the current optimal search value array to obtain the fitness difference value corresponding to the array. When the fitness difference value is greater than zero and is the largest fitness difference value in this time, update the forbidden change countdown value of the corresponding unique change parameter to a positive integer value positively correlated with the fitness difference value, and then execute step S427, where the unique change parameter refers to the parameter in the set of parameters to be optimized that is used to obtain the corresponding array through the increase or decrease process.
[0138] In the step S426, the fitness difference value being greater than zero and being the maximum fitness difference value in this time means that the best search value array has been obtained in the neighborhood search direction corresponding to the corresponding parameter in this time. Therefore, it is necessary to update the forbidden change countdown value of this corresponding parameter to a positive integer that is positively correlated with this fitness difference value to temporarily lock the search result in this neighborhood search direction. In addition, continuing with the example in the above step S424, among the eight new search value arrays corresponding to parameter A, parameter B, parameter E, and parameter F, if the fitness difference value of a certain new search value array corresponding to parameter F is greater than zero and is the maximum fitness difference value in this time (i.e., the largest among the eight fitness difference values in this time), then update the forbidden change countdown value of parameter F from zero to a positive integer value that is positively correlated with this fitness difference value, while the forbidden change countdown values of parameter A, parameter B, and parameter E remain zero.
[0139] S427. Determine whether there is any parameter in the set of parameters to be optimized whose current forbidden change countdown value is zero. If not, decrement the current forbidden change countdown value of each parameter in the set of parameters to be optimized by 1, and then return to execute step S427. Otherwise, execute step S428.
[0140] In the step S427, the current forbidden change countdown values of all parameters in the set of parameters to be optimized are not zero, which means that it is impossible to return to execute step S424 subsequently. Therefore, they need to be decremented together until the current forbidden change countdown value of at least one parameter is zero.
[0141] S428. Determine whether there is any fitness difference value greater than zero among the fitness difference values of the respective arrays. If so, update the current optimal search value array to the array corresponding to the minimum fitness among the new search value arrays, and then execute step S429. Otherwise, directly execute step S429.
[0142] S429. Increment the current iteration count by 1, and determine whether the current iteration count has reached the maximum iteration count . If so, use the current optimal search value array as the model hyperparameters corresponding to the pair of statements and the model and as the optimal search result for minimizing the objective function . Otherwise, return to execute step S424.
[0143] Based on the foregoing, it is possible to design a first method. By using the steps of the new heuristic algorithm described above to optimize the model hyperparameters, during the optimization process, different iterative lockings of the search results in each best neighborhood search direction can be performed based on different fitness difference values, which is conducive to quickly searching for the optimal model hyperparameters. By using a parameter adjustment method that first increases or decreases the parameter to be adjusted quantitatively and then increases or decreases it randomly, it is also possible to avoid falling into a local optimal solution, which is further conducive to quickly and accurately obtaining the optimal search result.
[0144] As Figure 2 shown, in the second aspect of this embodiment, a virtual system for implementing the code vulnerability detection method described in the first aspect or the first possible design is provided, including a statement code extraction unit, a context code extraction unit, a sample data collection unit, a vulnerability detection modeling unit, a detection model application unit, and a detection result summary unit.
[0145] The statement code extraction unit is configured to obtain the source code and extract the statement source code of at least one code statement from the source code.
[0146] The context code extraction unit is communicatively connected to the statement code extraction unit and is configured to extract the corresponding context source code from the source code for each code statement in the at least one code statement.
[0147] The sample data collection unit is communicatively connected to the statement code extraction unit and is configured to, for each code statement, collect the historical test sample data from all historical test sample data of the same type of code statements that match the corresponding statement by using a data collection algorithm, and add the collection result to the corresponding sample data set. The historical test sample data includes the code features of the same type of code statements used as model input items and the historical test post-code vulnerability classification label values of the same type of code statements used as model output items. The code features refer to the numerical features obtained by converting the statement source code and context source code of the corresponding statement.
[0148] The vulnerability detection modeling unit is communicatively connected to the sample data collection unit and is configured to, for each code statement, apply the corresponding sample data set to calibrate and verify the modeling of an artificial intelligence model based on a machine learning algorithm to obtain a corresponding code vulnerability detection model. The code vulnerability detection model is configured to output the predicted value of the post-test code vulnerability classification label of the corresponding statement after inputting the statement source code and context source code of the corresponding statement.
[0149] The detection model application unit is communicatively connected to the statement code extraction unit, the context code extraction unit, and the vulnerability detection modeling unit respectively. For each code statement, it is used to convert the corresponding statement source code and context source code into the corresponding code features, and then input the code features into the corresponding code vulnerability detection model to output the predicted value of the post-test code vulnerability classification label corresponding thereto.
[0150] The detection result summarization unit is communicatively connected to the detection model application unit, and is used to summarize the predicted value of the post-test code vulnerability classification label indicating that the corresponding code statement has a code vulnerability and the code statement corresponding to the predicted value of the post-test code vulnerability classification label, so as to obtain the vulnerability detection result of the source code and give an alarm reminder.
[0151] The working process, working details and technical effects of the foregoing device provided in the second aspect of this embodiment can be referred to the machine learning-based code vulnerability detection method described in the first aspect or possibly designed, and will not be elaborated herein.
[0152] As Figure 3 shown, the third aspect of this embodiment provides a computer system for executing the machine learning-based code vulnerability detection method described in the first aspect or possibly designed, including a memory, a processor, and a transceiver that are communicatively connected in sequence. Among them, the memory is used to store computer programs, the transceiver is used to send and receive messages, and the processor is used to read the computer programs and execute the machine learning-based code vulnerability detection method described in the first aspect or possibly designed. Specifically, for example, the memory may include, but is not limited to, random access memory (RAM), read-only memory (ROM), flash memory, first input first output (FIFO), and / or first input last output (FILO), etc.; the processor may include, but is not limited to, a microprocessor of the STM32F105 series. In addition, the computer system may further include, but is not limited to, a power supply module, a display screen, and other necessary components.
[0153] The working process, working details and technical effects of the foregoing computer system provided in the third aspect of this embodiment can be referred to the machine learning-based code vulnerability detection method described in the first aspect or possibly designed, and will not be elaborated herein.
[0154] In the fourth aspect of this embodiment, a computer-readable storage medium storing instructions for the machine learning-based code vulnerability detection method described in the first aspect or any possible design is provided. That is, instructions are stored on the computer-readable storage medium, and when the instructions are run on a computer, the machine learning-based code vulnerability detection method described in the first aspect or any possible design is executed. Among them, the computer-readable storage medium refers to a carrier for storing data, and may include, but is not limited to, computer-readable storage media such as floppy disks, optical discs, hard disks, flash memories, USB flash drives, and / or Memory Sticks. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices.
[0155] For the working process, working details, and technical effects of the aforementioned computer-readable storage medium provided in the fourth aspect of this embodiment, reference may be made to the machine learning-based code vulnerability detection method described in the first aspect or any possible design, which will not be elaborated herein.
[0156] In the fifth aspect of this embodiment, a computer program product is provided, including a computer program or instructions, and when the computer program or the instructions are executed by a computer, the machine learning-based code vulnerability detection method described in the first aspect or any possible design is implemented. Among them, the computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices.
[0157] Finally, it should be noted that the above are only preferred embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A code vulnerability detection method based on machine learning, characterized in that: include: Obtaining source code, and extracting statement source code of at least one code statement from the source code; For each code statement in the at least one code statement, extracting a corresponding context source code from the source code; For each code statement, a data collection algorithm is used to collect the historical test sample data from all historical test sample data of the same type of code statements matching the corresponding statement, and the collection result is added to the corresponding sample data set, wherein the historical test sample data contains code features of the same type of code statements used as model input items and historical post-test code vulnerability classification label values of the same type of code statements used as model output items, and the code features refer to numerical features obtained by converting the statement source code and context source code of the corresponding statement; For each of the code statements, the corresponding sample data set is applied to calibrate and verify the artificial intelligence model based on the machine learning algorithm to obtain the corresponding code vulnerability detection model, which specifically includes: obtaining multiple artificial intelligence models based on different machine learning algorithms; for a pair of statements and models in the at least one code statement and the multiple artificial intelligence models, the sample data set of the corresponding statement is applied to optimize the hyperparameters of the corresponding model based on the optimization algorithm to obtain the corresponding model hyperparameters and use them to make the objective function Minimize the optimization search results, where the objective function The calculation formula is as follows: In the formula, Indicates the output error indicator value of the artificial intelligence model, Indicates the time required for calculation of the artificial intelligence model; Selects the one with the smallest objective function from the multiple artificial intelligence models an artificial intelligence model; importing the optimization search results of the hyperparameters of the artificial intelligence model and the model parameters obtained in the optimization process and corresponding to the optimization search results into the artificial intelligence model, and obtaining a code vulnerability detection model corresponding to the code statement in the pair of statements and the model, wherein the code vulnerability detection model is used to output the predicted value of the post-test code vulnerability classification label of the corresponding statement after inputting the statement source code and the context source code of the corresponding statement; For each of the code statements, the corresponding statement source code and context source code are converted into the corresponding code features, and then the code features are input into the corresponding code vulnerability detection model, and the corresponding post-test code vulnerability classification label prediction value is output; The post-test code vulnerability classification label prediction value used to indicate the existence of code vulnerabilities in corresponding code statements and the code statements corresponding to the post-test code vulnerability classification label prediction value are summarized to obtain the vulnerability detection result of the source code and give an alarm reminder.
2. The code vulnerability detection method according to claim 1, characterized in that: The at least one code statement includes a first code statement for outputting text, a second code statement for obtaining user input, a third code statement for executing different code blocks according to different conditions, a fourth code statement for looping through elements in an iterable object, a fifth code statement for continuously looping through code blocks when a condition is true, a sixth code statement for importing an available program module so as to use a function / class in the available program module, and / or a seventh code statement for temporarily not performing any operation.
3. The code vulnerability detection method according to claim 1, characterized in that: For each code statement, a data collection algorithm is used to collect the historical test sample data from all historical test sample data of the same type of code statements that match the corresponding statement, and the collection results are added to the corresponding sample data set, including: For a certain code statement in the at least one code statement, a number of the historical test sample data are randomly collected from all the historical test sample data of the same type of code statements matching the corresponding statement, and the number of the historical test sample data randomly collected are added to a corresponding sample data set which is initially an empty set, wherein the historical test sample data contains code features of the same type of code statements used as model input items and historical post-test code vulnerability classification label values of the same type of code statements used as model output items, and the code features refer to numerical features obtained by converting the statement source code and context source code of the corresponding statement; Processing and analyzing the sample data set corresponding to the certain code statement to determine a gap direction of the sample data corresponding to the certain code statement; configuring a data collection algorithm according to the sample data gap direction, and using the data collection algorithm to collect the historical test sample data from all historical test sample data of the same type of code statements that match the certain code statement; Determine whether the newly collected historical test sample data matches the sample data gap direction. If so, add the newly collected historical test sample data to the sample data set corresponding to the certain code statement; otherwise, abandon adding the newly collected historical test sample data to the sample data set corresponding to the certain code statement.
4. The code vulnerability detection method according to claim 3, characterized in that: Configuring a data collection algorithm according to the sample data gap direction, and using the data collection algorithm to collect the historical test sample data from all historical test sample data of the same type of code statements matching the certain code statement, including the following steps S331 to S339: S331. According to the direction of the gap in the sample data, the initial position corresponding to the initial feature value of the code feature is initialized at the position around the data gap, and then step S332 is executed, wherein the initial position and the position around the data gap are respectively located in the multidimensional feature space constructed based on the code feature, and the position around the data gap is determined according to the direction of the gap in the sample data; S332. According to the initial feature value of the code feature, a direct binary search algorithm is used to search for the first historical test sample data whose model input item matches the initial feature value from all historical test sample data of the same type of code statements that match the certain code statement, and then execute step S333; S333. Set the feature number variable Initialize to 1, and then execute step S334; S334. Based on the current feature value of the code feature, change the first The feature value of the feature is obtained, and a new feature value of the code feature is obtained, and then step S335 is executed; S335. According to the new feature value of the code feature, the direct binary search algorithm is used to search for the second historical test sample data whose model input item matches the new feature value from all historical test sample data of the same type of code statements that match the certain code statement, and then execute step S336; S336. Compare the historical post-test code vulnerability classification label values in the second historical test sample data and the first historical test sample data and used as model output items. If the historical post-test code vulnerability classification label value in the second historical test sample data is different from the historical post-test code vulnerability classification label value in the first historical test sample data, execute step S337; otherwise, execute step S338; S337. Update the first historical test sample data to the second historical test sample data, and if the feature number variable is less than the total number of dimensions of the code feature, then the feature sequence variable Increment by 1, then return to step S334, otherwise, go to step S339; S338. The feature value of the feature is restored to the value before the change, and if the feature sequence variable is less than the total number of dimensions of the code feature, then the feature sequence variable Increment by 1, then return to step S334, otherwise, go to step S339; S339. Take the first historical test sample data as the collection result and end this collection.
5. The code vulnerability detection method according to claim 3, characterized in that: Processing and analyzing the sample data set corresponding to the certain code statement to determine the direction of the sample data gap corresponding to the certain code statement includes: According to the definition domain of the post-test code vulnerability classification label values of the same type of code statements matching the certain code statement, statistical analysis is performed on the sample data set corresponding to the certain code statement to determine the distribution interval of the post-test code vulnerability classification label values in the sample data set in the definition domain; Eliminating the distribution interval in the definition domain to obtain a distribution missing interval in the definition domain; According to the distribution missing interval, a sample data gap direction corresponding to the certain code statement is determined.
6. The code vulnerability detection method according to claim 1, characterized in that: For a pair of statements and models in the at least one code statement and the multiple artificial intelligence models, the sample data set of the corresponding statement is applied, and the hyperparameters of the corresponding model are optimized based on the optimization algorithm to obtain the corresponding model hyperparameters and use them to make the objective function The minimized optimized search results include the following steps S421 to S429: S421. For a pair of statements and models in the at least one code statement and the plurality of artificial intelligence models, initialize the corresponding and maximum number of iterations The optimization algorithm parameters are randomly generated, and an initial search value array of the parameter set to be optimized is then executed, wherein the parameter set to be optimized includes all hyperparameters of the artificial intelligence model in the pair of statements and models, and the initial search value array of the parameter set to be optimized is It is expressed as follows: In the formula, Indicates less than or equal to A positive integer, represents the total number of parameters in the parameter set to be optimized, Indicates the first The initial search value corresponding to the parameter, Indicates that The upper limit of the parameter search space corresponding to the parameters, Indicates that The lower limit of the parameter search space corresponding to the parameters, Represents a purely decimal random generator function; S422. Import the initial search value array of the parameter set to be optimized as a model hyperparameter into the artificial intelligence model in the pair of statements and the model to obtain a first artificial intelligence model, and then apply the sample data set corresponding to the code statement in the pair of statements and the model to perform model training and testing on the first artificial intelligence model to obtain a first output error situation index value and a first calculation time required index value, and finally import the first output error situation index value and the first calculation time required index value into the objective function , and the output result is used as the fitness corresponding to the initial search value array, and then step S423 is executed, wherein the objective function The calculation formula is as follows: In the formula, Indicates the output error indicator value of the artificial intelligence model, Indicates the time required for the calculation of the artificial intelligence model; S423. Use the initial search value array of the parameter set to be optimized as the current optimal search value array, and initialize the current number of iterations , and also initialize the prohibited change countdown value of each parameter in the parameter set to be optimized to zero, and then execute step S424; S424. Based on the current optimal search value array, each parameter in the parameter set to be optimized and whose current forbidden change countdown value is zero is independently increased and decreased to obtain the parameter set to be optimized. A new search value array is generated, and then step S425 is executed, wherein: represents the total number of parameters in the set of parameters to be optimized whose current forbidden change countdown value is zero, and ; S425. For the For each array in the new search value array, the corresponding array is imported as a model hyperparameter into the artificial intelligence model in the pair of statements and the model to obtain a second artificial intelligence model of the corresponding array, and then the sample data set corresponding to the code statement in the pair of statements and the model is applied to perform model training and testing on the second artificial intelligence model to obtain a second output error situation index value and a second calculation time required index value of the corresponding array, and finally the second output error situation index value and the second calculation time required index value are imported into the objective function , and use the output result as the fitness of the corresponding array, and then execute step S426; S426. For each array, the fitness corresponding to the current optimal search value array is subtracted from the fitness of the corresponding array to obtain the fitness difference value of the corresponding array, and when the fitness difference value is greater than zero and is the maximum fitness difference value this time, the forbidden change countdown value of the corresponding unique variable parameter is updated to a positive integer value positively correlated with the fitness difference value, and then step S427 is executed, wherein the unique variable parameter refers to the parameter in the parameter set to be optimized and is used to obtain the parameter of the corresponding array by increasing or decreasing the processing; S427. Determine whether there is any parameter in the set of parameters to be optimized whose current prohibited change countdown value is zero. If not, decrement the current prohibited change countdown value of each parameter in the set of parameters to be optimized by 1, and then return to step S427, otherwise, execute step S428; S428. Determine whether there is any fitness difference value greater than zero in the fitness difference values of each array. If so, update the current optimal search value array to the value in the array. The array in the new search value array and corresponding to the minimum fitness, then execute step S429, otherwise directly execute step S429; S429. Make the current number of iterations Add 1 to the current number of iterations. Whether the maximum number of iterations has been reached If so, the current optimal search value array is used as the model hyperparameter corresponding to the certain pair of statements and the model and used to make the objective function Minimize the optimized search results, otherwise return to execute step S424.
7. The code vulnerability detection method according to claim 1, characterized in that: The multiple artificial intelligence models include artificial intelligence models based on graph neural networks, support vector machines, K nearest neighbor method, stochastic gradient descent method, multivariate linear regression, multilayer perceptron, decision tree, back propagation neural network and / or radial basis function network.
8. A code vulnerability detection system based on machine learning, characterized in that: It includes a statement code extraction unit, a context code extraction unit, a sample data collection unit, a vulnerability detection modeling unit, a detection model application unit and a detection result summary unit; The statement code extraction unit is used to obtain source code and extract statement source code of at least one code statement from the source code; The context code extraction unit is communicatively connected to the statement code extraction unit, and is used to extract corresponding context source code from the source code for each code statement in the at least one code statement; The sample data collection unit is communicatively connected to the statement code extraction unit, and is used for collecting the historical test sample data from all historical test sample data of the same type of code statements matching the corresponding statement using a data collection algorithm for each code statement, and adding the collection result to the corresponding sample data set, wherein the historical test sample data contains code features of the same type of code statements used as model input items and historical post-test code vulnerability classification label values of the same type of code statements used as model output items, and the code features refer to numerical features obtained by converting the statement source code and context source code of the corresponding statement; The vulnerability detection modeling unit is communicatively connected to the sample data collection unit, and is used to apply the corresponding sample data set to each code statement, perform calibration verification modeling on the artificial intelligence model based on the machine learning algorithm, and obtain the corresponding code vulnerability detection model, specifically including: obtaining multiple artificial intelligence models based on different machine learning algorithms; for a certain pair of statements and models in the at least one code statement and the multiple artificial intelligence models, applying the sample data set of the corresponding statement, optimizing the hyperparameters of the corresponding model based on the optimization algorithm, obtaining the corresponding model hyperparameters and using them to make the objective function Minimize the optimization search results, where the objective function The calculation formula is as follows: In the formula, Indicates the output error indicator value of the artificial intelligence model, Indicates the time required for calculation of the artificial intelligence model; Selects the one with the smallest objective function from the multiple artificial intelligence models an artificial intelligence model; importing the optimization search results of the hyperparameters of the artificial intelligence model and the model parameters obtained in the optimization process and corresponding to the optimization search results into the artificial intelligence model, and obtaining a code vulnerability detection model corresponding to the code statement in the pair of statements and the model, wherein the code vulnerability detection model is used to output the predicted value of the post-test code vulnerability classification label of the corresponding statement after inputting the statement source code and the context source code of the corresponding statement; The detection model application unit is respectively connected to the statement code extraction unit, the context code extraction unit and the vulnerability detection modeling unit, and is used to convert the corresponding statement source code and the context source code into the corresponding code features for each code statement, and then input the code features into the corresponding code vulnerability detection model, and output the corresponding post-test code vulnerability classification label prediction value; The detection result summary unit is communicatively connected to the detection model application unit, and is used to summarize the post-test code vulnerability classification label prediction value indicating that a corresponding code statement has a code vulnerability and the code statement corresponding to the post-test code vulnerability classification label prediction value, obtain the vulnerability detection result of the source code and give an alarm reminder.
9. A computer system, characterized in that: It includes a memory, a processor and a transceiver which are communicatively connected in sequence, wherein the memory is used to store a computer program, the transceiver is used to send and receive messages, and the processor is used to read the computer program and execute the code vulnerability detection method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Fine-grained source code vulnerability detection method based on graph neural network
CN111259394A
Code vulnerability detection method based on machine learning
CN116541843A
Source code vulnerability detection method and device based on graph neural network and interpretable model
CN117272309A
Web vulnerability scanning method, system and device and medium
CN119203163A