A Directed Vulnerability Mining Method and System Based on Serial Ensemble Learning
Through the targeted vulnerability mining method based on serial integrated learning, the problems of low vulnerability mining efficiency and high false alarm rate in the existing technology are solved, and accurate vulnerability classification and fine-grained analysis are achieved in massive data scenarios.
Patent Information
- Application Number
- CN202211251160.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-13
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-10-13
AI Technical Summary
Existing vulnerability mining technologies are limited by the data set size and efficiency, making it difficult to batch process and explore potential vulnerabilities in the source code, resulting in high labor costs.
The directional vulnerability mining method based on serial ensemble learning is adopted, and the code training set is labeled, data preprocessed, sensitive function positioning and program control flow diagram analysis is carried out to form a uniform training set module, and a weak classifier of CART decision tree is trained to form the final strong classifier to achieve vulnerability mining.
It reduces the false alarm rate and missed alarm rate of traditional vulnerability mining, improves the accuracy and fine-grainedness of vulnerability classification, and can effectively judge vulnerability classification in massive data scenarios.
Smart Images

Figure CN115510455B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of computer vulnerability mining technology and network security technology, and mainly relates to a directional vulnerability mining method and system based on serial ensemble learning. Background Art
[0002] Under the information age, system vulnerabilities are often sensitive issues of the system. During the system coding process, if some serious vulnerabilities occur, they will be exploited by attackers, resulting in unpredictable consequences. Nowadays, domestic vulnerability mining often uses offline tools. Security personnel manually upload and import source code, and then use the tools to obtain vulnerability mining reports. Obviously, this method is limited by the size of the dataset and efficiency issues. How to use artificial intelligence methods to batch process and mine potential vulnerabilities in source code, thereby greatly reducing labor costs, is an important issue that needs to be solved in the current network security field.
[0003] Currently, there are related static mining technologies at home and abroad. For example, the solution proposed by Liu (Liu, B., et al. ADiff: Cross-version binary code similarity detection with DNN. in ASE 2018 - Proceedings of the 33rd ACM / IEEE International Conference on Automated Software Engineering. 2018. Montpellier, France. ACM.) uses DNN to predict C / C++ vulnerability source code. However, the vulnerability sample distribution of terminal devices is uneven, and there are a series of massive data vulnerability phenomena. Therefore, applying the serial ensemble learning method in the field of vulnerability mining is an innovation and improvement.
[0004] Serial ensemble learning is a method that uses a number of weak classifiers with strong correlation relationships to continuously train and finally obtain a strong classification function. Due to its serial ensemble learning mode, it can greatly reduce the residuals and thus fit by minimizing the squared loss function. Strictly speaking, the serial ensemble learning method is the result of a strong classifier formed by continuously training multiple weak classifiers. Therefore, the strong classifier obtained by training through the method of minimizing residuals must also produce better results than ordinary classifiers. Before performing serial ensemble learning, it is often necessary to process the data. After cleaning operations such as removing annotations and collecting stop words on the used data, it is marked using the method of tagging, and then sensitive word positioning is performed based on the processed data, and the CFG flow of this slice is obtained based on the sensitive words. The CFG flow is a representation method of a program, and the CFG graph includes information on control dependence and data dependence. Suppose there are two instructions, namely instruction A and instruction B. The so-called control dependence means that instruction B will only be executed after instruction A is executed, that is, A -> B; while data dependence refers to multiple statements that modify the same resource. The used word segmentation and vectorization algorithms use the Word2vec algorithm, which represents the content in the form of word vectors. For example, the sensitive statement memcpy is represented as [0.843, -0.125, 0.734, -0.3450, 0.6540], and it can also be used to calculate the similarity with other statements. Through training, a word can be mapped to a short vector of a fixed length, thus converting it into a language that machine learning can understand. Summary of the Invention
[0005] The present invention precisely aims at the deficiencies existing in the existing vulnerability mining technology, and provides a targeted vulnerability mining method and system based on serial ensemble learning. After tagging the code training set, a tagged training set is formed, and the vulnerable code training set among them is extracted for data preprocessing. Sensitive function positioning is performed on the preprocessed vulnerable code to obtain statements containing sensitive functions; the program control flow graph CFG is used to obtain the program slice related to this statement, and based on the number of codes in the vulnerability training set, a code training set without sensitive statements is used to mix with it to form a uniform training set module; the training set samples with initial weights are sent into the weak classifier of the CART decision tree for training, and the weight coefficient is adjusted and re-learned by calculating whether the classification error rate and the number of iterations meet the requirements, and the final strong classifier is formed in the way of weighted integration to realize the classification of test samples and complete the vulnerability mining. This method takes into account the context dependence relationship of the code and reduces the false positive rate and false negative rate of traditional vulnerability mining.
[0006] To achieve the above object, the technical solution adopted by the present invention is: a targeted vulnerability mining method based on serial ensemble learning, including the following steps:
[0007] S1, Determination of the source code training set: Obtain a batch of CWE vulnerability source codes from Git and the CWE official website, and organize them into a CWE vulnerability source code training set according to the numbers;
[0008] S2, Data preprocessing: Perform data preprocessing on the CWE vulnerability source code training set obtained in step S1 to obtain a vulnerability code training set from which sensitive words can be extracted;
[0009] S3, Sensitive word location: Locate sensitive words in the preprocessed vulnerability code training set. The sensitive word location targets the top 20 common vulnerability numbers on the CWE official website, and use joern to obtain all sensitive words in the vulnerability code training set and their line numbers;
[0010] S4, Obtaining vulnerability code slices: Obtain the control flow graph of the entire function code for the vulnerability code obtained in step S3, and extract the control flow information and data flow information related to sensitive words based on the control flow graph, and combine them to obtain a code slice based on the CFG control flow graph;
[0011] S5, Obtaining the dataset to be processed: According to the code slices of the CFG control flow graph obtained in step S4, mix the non-vulnerable code training set with a number of vulnerability code training sets similar to the code slices formed by the CFG control flow graph to form a dataset to be processed by the serial integrated learning module;
[0012] S6, Processing by the serial integrated learning module: Send the mixed slices obtained in step S5 into the serial integrated learning module for training. After training, obtain a weak classifier and calculate the classification error rate. Adjust the weights of the training set according to whether the number of inspection iterations and the error rate meet the requirements, and then perform weighted summation. Form a strong classifier through serial learning. Extract the test source code data through sensitive word location and the control flow program slices of the program statements where the sensitive words are located. After word segmentation and vectorization operations, send them into the strong classifier to implement the vulnerability classification and discrimination of the test source code, and complete the mining of vulnerabilities.
[0013] As an improvement of the present invention, the data preprocessing in step S2 at least includes removing code comments, replacing user-defined functions with the common function main(), and detecting unclosed symbols; in the data preprocessing step, use checkmax to judge the proportion of the code training set containing vulnerabilities, divide the dataset into a training set containing vulnerabilities and a training set without vulnerabilities, and mark the training set containing vulnerabilities as positive samples and the training set without sensitive functions as negative samples.
[0014] As another improvement of the present invention, in step S5, the mixing ratio range of the non-vulnerable code training set and the code slices formed by the CFG control flow graph is 1:1 - 1:1.2, and the optimal is 1:1.
[0015] As a further improvement of the present invention, step S6 further includes:
[0016] S61: Assign weights to the data set. The initial sample weights are all equal, and the weight of each sample is 1 / m, that is, M 1i = 1 / m, where i = 1…m;
[0017] S62: Send the entire data set into the serial ensemble learning module for training to obtain a CART decision tree weak classifier;
[0018] S63: Calculate the classification error rate and weight coefficient of the weak classifier, and at the same time set the maximum number of iterations and the minimum classification error rate. The classification error rate where, e n is the classification error rate, G n (x i ) is the classification result, y i is the label value, and M represents the sample weight coefficient; the weight coefficient A takes values in [0,1] as the classification error rate;
[0019] S64: If the classification error rate of the classifier is greater than the minimum classification error rate or the number of model iterations has not reached the maximum iteration threshold, then modify the weight coefficient of the sample and send it into the serial ensemble learning module for training again, and calculate the classification error calculation rate of the classifier after updating the sample weights until the classification error rate of the classifier is less than the minimum classification error rate or the number of model iterations reaches the maximum iteration threshold, and the process ends;
[0020] S65: Combine the weight coefficients of each weak classifier and use the weighted summation method to obtain the final strong classifier. The summation formula is where G n (x i ) outputs values of {1, -1}.
[0021] To achieve the above object, the technical solution adopted by the present invention is also: A directional vulnerability mining system based on serial ensemble learning, including a computer program, and when the computer program is executed by a processor, it implements the steps of the method described in claim 1.
[0022] Compared with the prior art, the present invention first preprocesses the publicly available dataset, then performs CFG flow slicing on the processed vulnerability source code, and then assigns weights to the training set and continuously adjusts the weights based on the training results, thereby realizing a serial ensemble learning method with strong classification function. The present invention improves the overfitting problem caused by the huge amount of data in the prior art during vulnerability mining, and can thus accurately and effectively perform vulnerability classification and judgment in the application scenario of massive data, improving the accuracy and fine-grainedness of traditional vulnerability mining; secondly, the code slicing based on the CFG flow takes into account the context dependence of the code, improving the fine-grainedness of slicing. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 It is a flowchart of the steps of the directed vulnerability mining method based on serial ensemble learning of the present invention;
[0024] Figure 2 It is the CFG flow graph obtained in Embodiment 2 of the present invention; DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] The present invention will be further clarified below in conjunction with the drawings and specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and not to limit the scope of the present invention.
[0026] Embodiment 1
[0027] A directed vulnerability mining method based on serial ensemble learning, as Figure 1 shown, includes the following steps:
[0028] S1, determination of the source code training set: Obtain a batch of CWE vulnerability source codes from Git and the CWE official website, and organize them into a CWE vulnerability source code training set according to the numbers;
[0029] S2, data preprocessing: Perform data preprocessing on the CWE vulnerability source code training set obtained in step S1, remove code comments, and replace user-defined functions with the common function main(), to obtain a vulnerability code training set from which sensitive words can be extracted;
[0030] S3, sensitive word location: Use joern to obtain the CFG flow information of each sample in the preprocessed training set, and based on the line where the sensitive word is located, obtain other control flow and data flow information related to the sensitive word along the CFG flow, and combine the statement where the sensitive word is located with other program statements to obtain a vulnerability code slice of the CFG flow; The sensitive word location is for the top 20 common vulnerability numbers on the CWE official website, such as CWE119, etc.;
[0031] S4, Vulnerability code slice acquisition: Obtain the control flow graph of the entire function code for the vulnerability code obtained in step S3, and extract the control flow information and data flow information related to sensitive words based on the control flow graph. After combination, obtain the code slice based on the CFG control flow graph;
[0032] S5, Dataset to be processed acquisition: According to the code slice of the CFG control flow graph obtained in step S4, mix the vulnerability-free code training set with a number of training sets similar to the vulnerability code training set and the code slice formed by the CFG control flow graph to form a dataset to be processed by the serial ensemble learning module; The mixing ratio ranges from 1:1 to 1:1.2, aiming to balance the dataset to improve the training effect of subsequent serial ensemble learning;
[0033] S6, Serial ensemble learning module processing: Send the mixed slice obtained in step S5 into the serial ensemble learning module for training. After training, obtain a weak classifier and calculate the classification error rate. Adjust the weight of the training set or form the final strong classifier according to whether the number of inspection iterations meets the requirements to complete the vulnerability mining. The specific steps are as follows:
[0034] S61: First, enter the training set weight allocation module, that is, assign weights to each sample, and the initial sample weights are equal;
[0035] S62: Send the entire training set into the learner for learning to obtain a weak classifier. The discrimination ability of the weak classifier is limited, so a strong classifier needs to be obtained in the subsequent steps;
[0036] S63: Calculate the classification error rate of the weak classifier, and at the same time calculate the weight coefficient of the weak classifier to prepare for the final strong classifier. At the same time, set the maximum number of iterations and the minimum classification error rate;
[0037] S64: If the classification error rate of the classifier is greater than the minimum classification error rate or the number of model iterations has not reached the maximum iteration threshold, modify the weight coefficient of the sample to highlight the weight of the misclassified sample, and send it into the serial ensemble learning module for training again. Similar to the attention mechanism, the classifier pays more attention to the samples with large sample weights. Calculate the classification error rate of the classifier after updating the sample weights until the classification error rate of the classifier is less than the minimum classification error rate or the number of model iterations reaches the maximum iteration threshold, and the process ends;
[0038] S65: Combine the weight coefficients of each weak classifier and use the weighted summation method to obtain the final strong classifier.
[0039] Example 2:
[0040] This embodiment uses a sample from the SARD dataset in the United States. Taking the ConfigExecProg function in X976 as an example, the program slice of ConfigExecProg is as follows:
[0041]
[0042]
[0043] The processing steps are as follows:
[0044] 1. After the above processing steps, the sensitive word execProg is obtained using joern. Here, no return value is detected for execProg, and it is not correctly released within the declared scope, which may lead to a memory leak. Therefore, the line where execProg is located is the sensitive statement line and is marked as 4.
[0045] 2. Use the joern tool to obtain the CFG flow of this program fragment. As shown in the following figure, it can be seen that the forward control flow related to exeProg is only ConfigExecProg and int i = 0. Therefore, the CFG flow vulnerability slice is int i = 0; execProg = strdup(string);. Therefore, it is extracted and stored in json format. ConfigExecProg{FunctionId: "X976", Childnum: "4", Dangerfunction: "execProg", IsCfgnode: "yes", Dangersentence: "execProg = strdup(string)"} Figure 2 shown. Through the following figure, it can be seen that the forward control flow related to exeProg is only ConfigExecProg and int i = 0. Therefore, the CFG flow vulnerability slice is int i = 0; execProg = strdup(string);. Therefore, it is extracted and stored in json format. ConfigExecProg{FunctionId: "X976", Childnum: "4", Dangerfunction: "execProg", IsCfgnode: "yes", Dangersentence: "execProg = strdup(string)"}
[0046] 3. Then mix the normal non-vulnerable program with it in a ratio of 1:1 to form a training set that can be processed by the serial integration module. Use the English word segmentation in the jieba library for word segmentation and the Word2vec algorithm for vectorization processing to prepare for the subsequent serial integration learning processing module.
[0047] 4. Here, the serial integration learning module takes the AdaBoost algorithm as an example. The training set in the form after the above processing is: D = {(x1, y1), (x2, y2)...}, where x1 = [0.88, 0.765...] such vectorized values, with a total of m samples. Specifically as follows:
[0048] 4.1 First, it is necessary to initialize the weights of the training set. The weight of each sample is 1 / m. Since it is the first round of iteration, the entire training set M 1i = D, where i = 1,..., mi, that is, M 1i = 1 / m.
[0049] 4.2 Assume the number of classifiers is N, that is, it needs to go through N iterations to terminate, or the classification error rate of the classifier is lower than the threshold after N times. At this time, n = 1, …, N, then the weight of the training set is expressed as M ni , send the training set to the classifier for training to obtain the weak classifier G n (x i ), and the output value is {1, -1}.
[0050] 4.3 Calculate the error rate of the CART decision tree weak classifier where I(G n (x) ≠ y i ) takes a value of 0 or 1. If the classification is correct, it is 0; otherwise, it is 1. Calculate the weight coefficient of the weak classifier where A takes values in [0, 1], then it means that the smaller e n , the larger a n . Conversely, the smaller, the larger the weight of the better classifier, which is convenient for the subsequent weighted summation.
[0051] 4.4 If the minimum acceptable error rate is e min , then when e n > e min , modify the weight of the training set Here, i = 1, …, m, n = 1, …, N - 1, G n (x i ) is the classification result predicted by the classifier, taking values {-1, 1}, y i is the true classification result of the training set, taking values {-1, 1}, a n is the weight coefficient of the weak classifier, Z n is the normalization factor, so that Therefore Send it to the weak classifier again for discrimination until the termination condition is met.
[0052] 4.5 Finally, all weak classifiers are weighted to obtain for the classification and discrimination of the test set.
[0053] For the above multiple processes, multiple intermediate weak classifier models are formed. Finally, the strong classifier formed by the weak classifiers is tested. The evaluation indicators of the classifier mainly adopt three indicators: false positive rate FPR, false negative rate FNR, and accuracy A. Among them, FPR is calculated as FNR is calculated as Let FP be the number of program slices that are not vulnerable but are detected as vulnerable, TP be the number of vulnerable program slices that are detected as vulnerable, FN be the number of vulnerable program slices that are detected as having no vulnerabilities, and TN be the number of program slices that are not vulnerable and are detected as not vulnerable. The experimental results are shown in the following table. The classification accuracy of this method on SARD reaches 85.372%.
[0054] Dataset FPR (%) FNR (%) A(%) SARD 15.32 17.34% 85.372 Vuldeepecker 18.91 22.31% 83.390 SySeVR 18.13 21.27% 84.670
[0055] In summary, the present invention can accurately and effectively perform vulnerability classification and judgment in the application scenario of massive data, improving the accuracy and fine-grainedness of traditional vulnerability mining; secondly, the code slicing based on CFG flow considers the context dependence relationship of the code, improving the fine-grainedness of slicing.
[0056] It should be noted that the above content only illustrates the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. For those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements all fall within the protection scope of the claims of the present invention.
Claims
1. A method for targeted vulnerability mining based on serial ensemble learning, characterized in that, it includes the following steps: S1, determination of source code training set: Obtain a batch of CWE vulnerability source codes and organize them into a CWE vulnerability source code training set according to the numbers; S2, data preprocessing: Perform data preprocessing on the CWE vulnerability source code training set obtained in step S1 to obtain a vulnerability code training set from which sensitive words can be extracted; S3, sensitive word location: Locate sensitive words in the preprocessed vulnerability code training set to obtain all sensitive words in the vulnerability code training set and their line numbers; S4, obtaining vulnerability code slices: Obtain the control flow graph of the entire function code of the vulnerability code obtained in step S3, and extract the control flow information and data flow information related to the sensitive words based on the control flow graph, and combine them to obtain code slices based on the CFG control flow graph; S5, obtaining the dataset to be processed: According to the code slices of the CFG control flow graph obtained in step S4, mix the non-vulnerable code training set with a number similar to the number of the vulnerability code training set with the code slices formed by the CFG control flow graph to form a dataset to be processed by the serial ensemble learning module; S6, processing by the serial ensemble learning module: Send the mixed slices obtained in step S5 into the serial ensemble learning module for training. After training, obtain weak classifiers and calculate the classification error rate. Adjust the weight of the training set according to whether the number of inspection iterations and the error rate meet the requirements, and then perform weighted summation. Form a strong classifier through serial learning. Extract the test source code data through sensitive word location and the control flow program slices of the program statements where the sensitive words are located. After word segmentation and vectorization operations, send them into the strong classifier to realize the vulnerability classification and discrimination of the test source code, and complete the vulnerability mining; This step specifically includes: S61: Assign weights to the data set. The initial sample weights are all equal, and the weight of each sample is 1 / m, that is, M 1i= 1 / m, where i = 1...m; S62: Send the entire dataset into the serial ensemble learning module for training to obtain CART decision tree weak classifiers; S63: Calculate the classification error rate and weight coefficient of the weak classifier, and simultaneously set the maximum number of iterations and the minimum classification error rate. The classification error rate where G n (x i ) is the classification result, y i is the label value, and M is the sample weight coefficient; the weight coefficient A of the weak classifier takes values in [0, 1]; S64: If the classification error rate of the classifier is greater than the minimum classification error rate or the number of model iterations has not reached the maximum iteration threshold, then modify the weight coefficient of the sample, send it into the serial ensemble learning module for training again, and calculate the classification error calculation rate of the classifier after updating the sample weight until the classification error rate of the classifier is less than the minimum classification error rate or the number of model iterations reaches the maximum iteration threshold, and the process ends; S65: Combine the weight coefficients of each weak classifier and use weighted summation to obtain the final strong classifier. The summation formula is where G n (x i ) has an output value of {1, -1}.
2. The method for targeted vulnerability mining based on serial ensemble learning as described in claim 1, characterized in that: The data preprocessing in step S2 at least includes the removal of code comments, replacing user-defined functions with the common function main(), and detection of unclosed symbols; In the data preprocessing step, use checkmax to judge the proportion of the code training set containing vulnerabilities, divide the dataset into a training set containing vulnerabilities and a training set without vulnerabilities, mark the training set containing vulnerabilities as positive samples, and mark the training set without sensitive functions as negative samples.
3. The method for targeted vulnerability mining based on serial ensemble learning as described in claim 1, characterized in that: In the step S5, the mixing ratio range of the vulnerability-free code training set and the code slice formed by the CFG control flow graph is 1:1 - 1:1.
2.
4. The directed vulnerability mining method based on serial ensemble learning according to claim 3, characterized in that: in the step S5, the mixing ratio of the vulnerability-free code training set and the code slice formed by the CFG control flow graph is 1:
1.
5. A directed vulnerability mining system based on serial ensemble learning, including a computer program, characterized in that when the computer program is executed by a processor, the steps of the method according to claim 1 are implemented.
Citation Information
Patent Citations
Abnormal intrusion detection ensemble learning method and apparatus based on Wiener process
CN103716204A
Network intrusion detection method based on active learning and migration learning
CN109462610A