Software vulnerability parallel mining method for optimizing machine learning

Through the optimized machine learning method, KNN deep clustering algorithm and Euro-style distance calculation are used to establish a multi-class vulnerability information resource library, solving the problem of automatic software vulnerability mining, and achieving high-precision and low false alarm rate vulnerability mining effect.

CN120046154AActive Publication Date: 2025-05-27CHINESE PEOPLES LIBERATION ARMY UNIT 61660
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202411968709.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-27
Estimated Expiration
2044-12-30

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively and automatically mine software vulnerabilities, resulting in the threat of software security.

Method used

Using an optimized machine learning method, a multi-class vulnerability information resource library is established through KNN deep clustering algorithm and Euclidean distance calculation to realize parallel operation and mining of software vulnerabilities.

Benefits of technology

The feature classification accuracy of software vulnerabilities is improved, the F1 value is maintained (above 0.88), and the false alarm rate of mining is reduced, realizing the parallel mining task of multiple software vulnerabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046154A_ABST
    Figure CN120046154A_ABST
Patent Text Reader

Abstract

The invention relates to a software vulnerability parallel mining method for optimizing machine learning, and belongs to the field of vulnerability mining. According to the method, vulnerability code functions, parameters, target variables and the like in software vulnerability historical data are mapped into a symbol name domain, normalization processing is carried out on the vulnerability code functions, the parameters, the target variables and the like, a multi-class vulnerability information resource library composed of a base class, a subclass and a code class is established, and vulnerability code naming is unified. A KNN algorithm in machine learning is used for deeply clustering and mining software features, and the Euclidean distance between any two is introduced and calculated. And defining the key node to which the feature with the maximum similarity belongs as the software vulnerability. According to the method, the feature classification precision is high, the F1 value of software vulnerability mining can be kept above 0.88, and the mining false alarm rate is low; through operation of a parallelization tool, parallel mining tasks of various software vulnerabilities can be realized in a specific scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of vulnerability mining, and particularly relates to a method for parallel mining of software vulnerabilities that optimizes machine learning. Background Art

[0002] With the popularization and development of computers, software support is becoming increasingly indispensable. The security and usability of software are extremely important for systems and users. When there are vulnerabilities in software, illegal persons can attack authorized access modules through operations, resulting in the leakage of internal confidential information and even system paralysis. In order to improve the stability of software, software developers have studied and obtained a series of strict security design guidelines and operation processes. However, due to the complexity of the software itself, in a multi-source application environment, vulnerabilities can always be mined from perspectives or dimensions that developers never considered, seriously threatening the security of network systems.

[0003] Therefore, the present invention introduces the KNN (K-Nearest Neighbor) deep clustering algorithm in machine learning algorithms to learn software programming rules from massive data, enabling the vulnerability mining method to perform self-decision-making, classification, and global prediction using existing rules. Map the vulnerability code functions, parameters, target variables, etc. in the historical data of software vulnerabilities to the symbolic name domain, and perform normalization processing on them. Establish a multi-class vulnerability information resource library composed of base classes, subclasses, and code classes, unify the vulnerability coding and naming, and after constructing the vulnerability information resource library, introduce the Euclidean distance to complete the parallel running and mining work of multi-class vulnerabilities. Summary of the Invention

[0004] (1) Technical Problems to be Solved

[0005] The technical problem to be solved by the present invention is how to provide a method for parallel mining of software vulnerabilities that optimizes machine learning to solve the problem of automatic mining of software vulnerabilities.

[0006] (2) Technical Solutions

[0007] In order to solve the above technical problems, the present invention proposes a method for parallel mining of software vulnerabilities that optimizes machine learning, and the method includes the following steps:

[0008] S1. Normalization Processing of Software Data

[0009] Perform feature mapping and normalization preprocessing on the code functions, parameters, and target variable information in the software package, quantify the occurrence frequencies of key nodes in different types of vulnerabilities in the software package, and obtain corresponding feature values;

[0010] S2. Algorithm Design: Taking the eigenvalue as the input, introducing machine learning to complete the classification of software data features; using the KNN (K-Nearest Neighbor) deep clustering algorithm in machine learning to deeply cluster and mine software features, and then using the Euclidean distance to calculate the similarity between the clustered features and the corpus labels of multiple types of vulnerability information resources, and defining the key node to which the feature with the maximum similarity belongs as the software vulnerability;

[0011] S3. Evaluating the detection results through performance evaluation metrics to intelligently complete the detection of software vulnerabilities.

[0012] (III) Beneficial Effects

[0013] The present invention proposes a method for parallel mining of software vulnerabilities by optimizing machine learning, mapping vulnerability code functions, parameters, target variables, etc. in software vulnerability historical data to the symbolic name domain, performing normalization processing on them, establishing a multi-type vulnerability information resource library composed of base classes, subclasses, and code classes, and uniformly naming vulnerability encodings. Using the KNN algorithm in machine learning to deeply cluster and mine software features, and introducing the calculation of the Euclidean distance between any two of them. Defining the key node to which the feature with the maximum similarity belongs as the software vulnerability. It is proved by simulation that the proposed method has a high feature classification accuracy, and the F1 value of software vulnerability mining can be maintained above 0.88, with a low false alarm rate of mining. Therefore, through the operation of a parallelization tool, multiple software vulnerability parallel mining tasks can be realized in a specific scenario. Description of the Drawings

[0014] Figure 1 is the flow chart of the present invention;

[0015] Figure 2 is a schematic diagram of a processed data set sample (input training sample of the algorithm). Detailed Embodiments

[0016] To make the objectives, contents, and advantages of the present invention clearer, the following further describes in detail the specific embodiments of the present invention with reference to the drawings and embodiments.

[0017] 1. Technical Preparation

[0018] Database: Selecting 4 groups of NASA software vulnerability data sets of the open-source database Tera-PROMISE as experimental samples.

[0019] Language and Platform: python, Sklearn machine learning library.

[0020] 2. Implementation Process

[0021] S1. Normalization processing of software data.

[0022] Extracting code functions, parameters, and target variables from the software package,

[0023] Based on the software vulnerability levels and types, it can be known that the definitions of multiple types of software vulnerabilities are abstract values and cannot be used as quantitative data to be input into machine learning algorithms for parallel mining. Therefore, an existing dataset of multiple types of vulnerabilities is collected and mapped to the feature vector space in sequence to clarify the attribute feature values corresponding to different types of vulnerabilities. Considering the diversity of the thoughts and behaviors of different developers and coders, information such as code functions, parameters, and target variables in software packages is subjected to feature mapping and normalization preprocessing, and the occurrence frequencies of key nodes in different types of vulnerabilities in the software packages are quantified to obtain corresponding feature values, with a unified format for facilitating subsequent calculations. In a certain embodiment, the software package uses 4 groups of NASA software vulnerability datasets of the open-source database Tera-PROMISE.

[0024] S11. First, filter out non-ASCII bytes in the software package that have no connection with software vulnerabilities, map information such as code functions, parameters, and target variables to symbolic names one by one, perform unified coding naming, and after normalization processing, project them into the feature vector space to obtain corresponding feature values;

[0025] S12. Quantify the occurrence frequencies of key nodes in different types of vulnerabilities through natural text language to obtain corresponding feature values;

[0026] S13. Use the feature values of information such as code functions, parameters, and target variables and the feature values of key nodes together as the feature values of software vulnerability data;

[0027] Taking the code function as an example, the specific steps of S11 are as follows:

[0028] Suppose the code function sequence of any software obtained by scanning the system is X = {x 1 , x 2 , …, x n}, and the set representation of the key nodes of any source code file is A, A = {α 1 , α 2 , …, α m}, where n and m both represent non-zero integers; the mapping process is expressed as:

[0029] ψ(x) = λ(x, α) · v α,A

[0030]

[0031] In the formula, v α,A represents the term frequency parameter of the inverse document, which can eliminate a part of the similarity of software vulnerability sequences and reduce the influence of key nodes of non-occurring vulnerabilities on vulnerability feature extraction. The calculation formula is:

[0032]

[0033] Among them, df α represents the number of times a file in which the key node α appears, and tf α,A represents the ratio between the number of times the key node α appears in the set A and the total number N of key nodes. α

[0034] S2. Algorithm design: Using the eigenvalue of software vulnerability data as the input, machine learning is introduced to complete the classification of software data features. Software data covers all the information required for parallel vulnerability mining. However, software programming is complex and the information dimension is high. To improve the performance of the mining algorithm, the KNN (K-Nearest Neighbor) deep clustering algorithm in machine learning is used. The KNN algorithm deeply clusters to mine software features, and then the Euclidean distance is used to calculate the similarity between the features after clustering and the corpus labels of multiple types of vulnerability information resources. The key node to which the feature with the maximum similarity belongs is defined as the software vulnerability.

[0035] The deep clustering process mainly includes three steps: software feature specification, feature dimensionality reduction, and feature clustering. At the same time of feature extraction, interference noise and duplicate and missing data also need to be removed.

[0036] S21. Normalize the eigenvalue;

[0037] If the number of software samples to be mined is M, and the eigenvalue of the β-th feature in each sample is expressed as Q=(q β1 ,q β2 ,…,q βM ), the normalization process is expressed as:

[0038]

[0039] represents the average value of the β-th feature, and σ represents the standard deviation of the β-th feature. Since the processed features conform to the normal distribution, so σ = 1.

[0040] S22. Use deep learning to reduce the dimensionality of the normalized high-dimensional eigenvalue to obtain low-dimensional features, improving the performance of the mining algorithm. This step includes: input layer, output layer, and hidden layer. The dimensionality reduction process is

[0041] ζ = f θ (q) = ω(Wq + b)

[0042] p = g θ' (ζ) = ω'(W'ζ + b')

[0043] q represents the input feature, p represents the feature after dimensionality reduction, f θ ​(q) represents the non - linear function for the input to the hidden layer, which is used to process the uncorrelated features of the input. ζ represents the intermediate hidden variable for the input to the hidden layer, g θ' (ζ) represents the non - linear function from the hidden layer to the output layer, which is used to process the uncorrelated features of the hidden layer input. ω, ω' describe the activation functions of different layers, W, W' describe the weights of different layers, and b, b' describe the offset vectors of different layers.

[0044] S23. The present invention is a parallel mining algorithm based on multiple types of vulnerabilities, which deeply clusters the low - dimensional features as the features of the "feature source". j M is any set of feature types in the sample to be mined, o α represents the clustering center obtained by the KNN algorithm. The deep clustering can be described as:

[0045]

[0046] sim(p, o α ) represents the cosine distance between p and o α , and η represents the clustering process.

[0047] The classification mining algorithm can be used to classify and identify the nearest neighbor feature sets of different "feature sources".

[0048] S24. The present invention uses the Euclidean distance to calculate the similarity between the feature x of the "feature source" and the corpus label y in the multi - type vulnerability information resource library, and defines the key node to which the feature with the maximum similarity belongs as the software vulnerability. Among them, the multi - type vulnerability information resource library is a resource library composed of base classes, sub - classes, and code classes. The calculation formula is:

[0049]

[0050] Among them, n is the corresponding dimension.

[0051] S3. The detection results are evaluated through performance evaluation indicators to intelligently complete the detection of software vulnerabilities.

[0052] Evaluating the mining results can be described as a binary classification problem, that is, no error in mining is a positive example, and having an error is a negative example. The specific calculation formulas for each index are:

[0053] 1) Precision describes the ratio of the number of correctly mined software vulnerabilities to the total number of mined ones. The calculation formula is

[0054]

[0055] 2) Recall rate describes the ratio of the number of correctly mined software vulnerabilities to the number of real software vulnerabilities. The calculation formula is

[0056]

[0057] 3) The F1 value describes the harmonic mean between recall and precision. Since there is a non-linear relationship between the two, there may be a situation where the recall is high but the precision is low. The larger the F1 value, the smaller the inconsistency between the two values. The calculation process is as follows

[0058]

[0059] In the formula, the true positive example TP represents the data that is actually a vulnerability and has been discovered. The true negative example TN represents the data that is actually a vulnerability but has not been discovered. The false positive example FP represents the data that is not actually a vulnerability but has been discovered. The false negative example FN represents the data that is actually a vulnerability but has not been discovered either.

[0060] In the results, the F1 value showed a certain oscillation as the number of experiments increased, and was generally stable between [0.86, 0.89], indicating that the recall rate of this technical method is high and the mining is accurate, and the difference between the two values is small and evenly harmonized.

[0061] Beneficial effects:

[0062] To enhance the security of software, the present invention proposes a method for parallel mining of software vulnerabilities by optimizing machine learning. Map the vulnerability code functions, parameters, target variables, etc. in the historical data of software vulnerabilities to the symbolic name domain, and perform normalization processing on them. Establish a multi-class vulnerability information resource library composed of base classes, sub-classes, and code classes, and unify the vulnerability coding and naming. Use the KNN algorithm in machine learning to deeply cluster and mine software features, and introduce the calculation of the Euclidean distance between any two of them. Define the key node to which the feature with the largest similarity belongs as the software vulnerability. Through simulation, it is proved that the method of the present invention has a high feature classification accuracy, and the F1 value of software vulnerability mining can be maintained above 0.88, and the mining false alarm rate is low. Therefore, through the operation of the parallelization tool, it is possible to realize the parallel mining tasks of multiple software vulnerabilities in a specific scenario.

[0063] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and deformations can be made, and these improvements and deformations should also be regarded as the protection scope of the present invention.

Claims

1. A software vulnerability parallel mining method for optimizing machine learning, characterized in that: The method comprises the following steps: S1. Software data normalization processing The code functions, parameters, and target variable information in the software package are feature mapped and normalized, and the occurrence frequency of key nodes in different types of vulnerabilities in the software package is quantified to obtain the corresponding feature values; S2. Algorithm design: Take the feature value as input, introduce machine learning, and complete the classification of software data features; use the KNN (K-NearestNeighbor) deep clustering algorithm in machine learning to deep cluster and mine software features, and then use the Euclidean distance to calculate the similarity between the clustered features and the labels of the corpus of multiple vulnerability information resource libraries, and define the key node to which the feature with the maximum similarity belongs as a software vulnerability; S3. Evaluate the detection results through performance evaluation indicators and complete the detection of software vulnerabilities intelligently.

2. The software vulnerability parallel mining method for optimizing machine learning according to claim 1, characterized in that: The software package mentioned is the NASA software vulnerability dataset.

3. The software vulnerability parallel mining method for optimizing machine learning according to claim 1, characterized in that: The S1 specifically includes: S11, first filter the non-ASCII bytes in the software package that are not related to the software vulnerability, map the code function, parameter, and target variable information to the symbol name one by one, unify the encoding and naming, and after normalization, project them into the feature vector space to obtain the corresponding feature value; S12. By quantifying the frequency of occurrence of key nodes in different types of vulnerabilities through natural text language, the corresponding feature values ​​can be obtained; S13. The characteristic values ​​of information such as code functions, parameters, and target variables and the characteristic values ​​of key nodes are taken together as characteristic values ​​of software vulnerability data.

4. The software vulnerability parallel mining method for optimizing machine learning according to claim 3, characterized in that: For the code function, the S11 includes: Assume that the code function sequence of any software obtained by scanning the system is X=x1,x2,…,x n} , the set of key nodes corresponding to any source code file is represented as A, A = {α1, α2, …, α m }, n and m are both integers not equal to 0; the mapping process is expressed as: ψ(x)=λ(x,α)·v α,A In the formula, v α,A The term frequency parameter representing the inverse document can eliminate the similarity of some software vulnerability sequences and reduce the impact of non-vulnerability key nodes on vulnerability feature extraction; the calculation formula is: Among them, df α It is expressed as the number of files where the key node α appears, tf α,A Represented as the number of occurrences of key node α in set A df α The ratio between the number of critical nodes and the total number N.

5. The software vulnerability parallel mining method for optimizing machine learning according to any one of claims 1 to 4, characterized in that: The S2 specifically includes: S21, normalizing the eigenvalues; S22. Use deep learning to reduce the normalized high-dimensional feature values ​​to obtain low-dimensional features; S23, deep clustering of low-dimensional features as "feature source" features; S24. Use Euclidean distance to calculate the similarity between the features of the "feature source" and the corpus labels in the multi-class vulnerability information resource library, and define the key node to which the feature with the maximum similarity belongs as a software vulnerability, where the multi-class vulnerability information resource library is a resource library composed of base classes, subclasses, and code classes.

6. The software vulnerability parallel mining method for optimizing machine learning according to claim 5, characterized in that: The S21 includes: If the number of software samples to be mined is M, the characteristic value of the βth feature in each sample is expressed as Q = (q β1 ,q β2 ,…,q βM ), the normalization process is expressed as: represents the mean value of the β-th feature, and σ represents the standard deviation of the β-th feature.

7. The method for parallel software vulnerability mining for optimizing machine learning according to claim 6, characterized in that: The S22 includes: an input layer, an output layer and a hidden layer, and the dimension reduction process is: ζ=f θ (q)=ω(Wq+b) p=g θ' (ζ)=ω'(W'ζ+b') Among them, q represents the input feature, p represents the feature after dimensionality reduction, and f θ (q) represents the nonlinear function input to the hidden layer, which is used to process the irrelevant features of the input, ζ represents the intermediate latent variable input to the hidden layer, and g θ' (ζ) represents the nonlinear function from the hidden layer to the output layer, which is used to process the uncorrelated features of the hidden layer input; ω, ω' describe the activation functions of different layers, W, W' describe the weights of different layers, and b, b' describe the offset vectors of different layers.

8. The method for parallel software vulnerability mining for optimizing machine learning according to claim 7, characterized in that: The S23 includes: based on a parallel mining algorithm for multiple types of vulnerabilities, deep clustering is performed using low-dimensional features as features of "feature sources"; M is any set of feature types in the sample to be mined, o α Represents the cluster center obtained by the KNN algorithm. The deep clustering is described as: sim(p,o α ) indicates p, o α The cosine distance between them, η represents the clustering process, and the classification mining algorithm can classify and identify the nearest neighbor feature sets of different "feature sources".

9. The software vulnerability parallel mining method for optimizing machine learning according to claim 8, characterized in that: The S24 includes: using the Euclidean distance to calculate the similarity between the feature x of the "feature source" and the corpus label y in the multi-class vulnerability information resource library, and the calculation formula is: Among them, n is the corresponding dimension.

10. The software vulnerability parallel mining method for optimizing machine learning according to claim 1, characterized in that: S3 includes: the evaluation mining results can be described as a two-classification problem, that is, mining error-free positive examples and error-free negative examples. The specific calculation formula of each indicator is: Precision describes the ratio of the number of correctly discovered software vulnerabilities to the total number of discovered vulnerabilities. The calculation formula is: The recall rate describes the ratio of the number of correctly discovered software vulnerabilities to the number of real software vulnerabilities. The calculation formula is: The F1 value describes the harmonic mean between recall and precision. The larger the F1 value, the smaller the inconsistency between the two values. The calculation process is: In the formula, true positive TP indicates data whose actual result is a vulnerability and is mined, true negative TN indicates data whose actual result is a vulnerability and is not mined, false positive FP indicates data whose actual result is not a vulnerability but is mined, and false negative FN indicates data whose actual result is a vulnerability and is not mined.

Citation Information

Patent Citations

  • Parallel vulnerability mining method based on open source library and text mining

    CN104166680A

  • Programming mode and mode matching based bug clustering method

    CN105045715A

  • Cross-project vulnerability detection model based on domain self-adaption

    CN115168865A

  • Network protocol vulnerability mining method based on mixed variation strategy

    CN115238822A

  • Code quality detection method and device, equipment and storage medium

    CN117608630A