Vulnerability severity level discrimination method and system based on clustering analysis

Through the cluster analysis method, the network vulnerability data is preprocessed and clustered, and the attack vector features are extracted and the vulnerability severity level is predicted using heat maps, which solves the problem of insufficient vulnerability evaluation speed and accuracy in the existing technology, and achieves more efficient and accurate vulnerability severity level analysis and judgment.

CN120050116APending Publication Date: 2025-05-27CHANGSHA UNIVERSITY OF SCIENCE AND TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510511554.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Existing vulnerability assessment methods are difficult to provide accurate and real-time threat severity annotations when facing a large number of vulnerabilities, resulting in lagging evaluation and response speeds, especially in the face of new unknown threats.

Method used

The vulnerability severity level discrimination method based on cluster analysis is adopted to accurately determine the vulnerability severity level by obtaining network vulnerability data, data preprocessing, clustering analysis, attack vector feature extraction and heat map weight prediction.

Benefits of technology

It improves the accuracy and efficiency of the analysis of the severity level of vulnerability, reduces subjective deviations in manual judgment, can automatically update vulnerability information in real time, adapt to the dynamically changing network security environment, and improves the overall network security protection level.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120050116A_ABST
    Figure CN120050116A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of network security vulnerability analysis, in particular to a vulnerability severity level judgment method and system based on clustering analysis. The method comprises the following steps: acquiring network vulnerability data; performing data preprocessing on the acquired network vulnerability data; performing feature extraction on the preprocessed data by using a convolutional neural network model; performing clustering analysis on the extracted features to obtain a vulnerability marking result; and performing scoring based on the obtained vulnerability marking result to obtain a scoring result. According to the method, the BERT benchmark attention is optimized by using the adaptive feature attention, and the vulnerability data is used for retraining supervised learning, so that the effect of accurately predicting the attack vector is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of network security vulnerability analysis, and in particular, to a method and system for determining the severity level of vulnerabilities based on cluster analysis. Background Art

[0002] With the rapid development of information technology, network security threats are becoming increasingly severe, and vulnerability management has become the core link to ensure the security of information systems. Due to the increasing complexity of software security threats, any software or device is vulnerable to targeted attacks by cybercriminals using vulnerabilities. Traditional vulnerability assessment methods have gradually revealed their limitations. Especially when a large number of vulnerabilities need to be marked for threat severity, existing systems are difficult to provide good accuracy and timeliness. Most current defense systems rely on manual updates and inefficient rule matching methods, which makes the assessment and response speed of vulnerabilities far lag behind that of attackers, and also results in particularly weak accuracy and robustness in the face of new and unknown threats.

[0003] Judging the severity level of vulnerabilities is a key technology to avoid losses caused by vulnerabilities in a timely manner. It is difficult to accurately determine the dynamic analysis and multi-dimensional correction of vulnerability risk factors through traditional static assessment or discussions among professionals. By using the global perception ability of real-time threat intelligence and the precise quantification advantage of environment adaptation, it makes up for the deficiencies of manual judgment in terms of timeliness, objectivity, and scenario adaptability. In a vulnerability analysis system with multi-source data collaboration, there are multiple functional nodes such as a vulnerability knowledge base, a dynamic scoring engine, an environment adaptation module, and an intelligent decision-making unit, which have heterogeneous data processing and model calculation capabilities. Through the multi-module collaborative analysis mechanism, the system can integrate the basic attributes of vulnerabilities, threat situation evolution data, and environmental characteristic parameters to achieve deep collaboration of risk quantification and resource scheduling, effectively improving the accuracy of vulnerability repair priority decision-making and the automatic response ability of security protection.

[0004] However, with the exponential growth of the number of vulnerabilities, the traditional manual-based vulnerability management method can no longer meet the requirements of real-time and precision. Judging the severity level of the potential harm of vulnerabilities provides a reference basis for the vulnerability repair priority. However, the existing general scoring system still has significant limitations in practical applications. The traditional general vulnerability system scoring is usually generated based on static data at the initial stage of vulnerability disclosure, and it is difficult to dynamically reflect the evolution of vulnerability exploitation techniques, changes in attack frequencies, or risk attenuation after the release of repair patches, resulting in the scoring results lagging behind the actual threat situation. There is an urgent need to achieve more accurate judgment and analysis of vulnerability severity levels through technical means.

[0005] To address the above problems, it is urgent to propose a precise vulnerability severity level judgment and analysis system based on the Common Vulnerability Scoring System. At the same time, a fully automated vulnerability severity level judgment system is designed. By clustering analysis, the internal connections of vulnerabilities are mined, and then natural language tools are used to vectorize vulnerability data for model supervised learning, effectively using software vulnerability severity level judgment and analysis to improve the efficiency and accuracy of software threat location. Summary of the Invention

[0006] To solve the above-mentioned problems, the present invention provides a method and system for discriminating vulnerability severity levels based on clustering analysis. It can be applied to the inefficient problems such as the lag in the evaluation and response speed, misjudgment, task backlog, and manpower consumption caused by a large number of software security threats and program vulnerabilities.

[0007] In the first aspect, a method for discriminating vulnerability severity levels based on clustering analysis provided by the present invention adopts the following technical solutions: A method for discriminating vulnerability severity levels based on clustering analysis includes: Obtain network vulnerability data; Perform data preprocessing on the obtained network vulnerability data; Use a clustering algorithm to perform clustering analysis on the preprocessed data to obtain reduced-dimensional clustering data; Use a convolutional neural network model to extract attack vector features from the clustering data; By repeatedly extracting attack vector features, predict the vulnerability severity level score based on the heatmap weight.

[0008] Further, the obtaining of network data includes obtaining a network vulnerability data set, including a large-scale vulnerability information data set and a small-scale vulnerability information data set; each vulnerability includes 8 attack vector strings and vulnerability information.

[0009] Further, the data preprocessing of the obtained network data includes data cleaning, natural language processing, and node vectorization. Among them, natural language processing includes using stemming and part-of-speech reduction. After natural language processing, data node vectorization is performed, and a BERT bidirectional encoder is used to process the data to obtain the data vector after node vectorization.

[0010] Furthermore, the clustering analysis of the preprocessed data using a clustering algorithm to obtain clustered data after dimensionality reduction includes the use of a K-means clustering algorithm, which is a clustering algorithm based on sample set partitioning. K-means clustering divides the sample set after data preprocessing into K subsets to form K classes, and divides the samples into K classes, miniaturizing the distance from the center of each sample to the class to which it belongs, so that each sample belongs to only one class, and determines the most appropriate number of fern clusters and number of dimensionality reduction by respectively calculating the error sum of squares rule, silhouette coefficient and Karlinski-Harabas index.

[0011] Furthermore, the attack vector feature extraction of the clustered data using a convolutional neural network model includes constructing a convolutional neural network based on a natural language processing pre-training model BERT and a text convolutional neural network, wherein the input attack vector data is replaced with a special tag using a masked language mechanism, the cross entropy loss between the actual tag and the predicted tag at the masked position is calculated, and the relationship between all tags in the sequence is calculated using a multi-head self-attention mechanism to assign a weight to each tag, wherein the cross entropy loss between the actual tag and the predicted tag at the masked position is: Marked with X .

[0012] Furthermore, the method of using a convolutional neural network model to extract attack vector features from the clustered data also includes performing node vectorization based on a transformer mechanism to obtain a vector representation of the attack vector, and then performing a convolution operation on the sequence using multiple adaptive convolution kernels to obtain a feature graph sequence. The distribution weights calculated using a multi-head attention mechanism are used as bias items for the convolution, and then sent to a pooling layer for a pooling operation to obtain high-dimensional vulnerability information features. Finally, after performing a nonlinear transformation on the features using an adaptive activation function, the copper leakage attack vector is multi-classified to obtain a unique hot probability.

[0013] Furthermore, the attack vector feature extraction is repeatedly performed to predict the vulnerability severity level score based on the heat map weight, including using three convolution kernels of different sizes to extract language model features of different lengths and performing a maximum pooling operation, wherein the category of each attack vector is obtained by one-hot probability calculation, and then the vulnerability score is calculated using a calculation formula, and the vulnerability severity level is judged based on the vulnerability score.

[0014] Furthermore, the vulnerability score is calculated using a calculation formula, including calculating the sum of attack vector categories by defining standardized fixed values ​​of attack vector categories, and then obtaining a basic score of the vulnerability severity level, which is expressed as: Among them, AV, AC, S, C, I, and A are attack vector categories, and their normalized values are defined by the vulnerability scoring system, and are reflected in the dataset. , , is the sum of this attack vector category, is the sum of the attack vector categories.

[0015] Furthermore, in the method of repeatedly extracting attack vector features and predicting the vulnerability severity level score based on the heatmap weight, it also includes calculating the exploitability evaluation score and the environmental impact evaluation score of the vulnerability according to the predicted attack vector categories and the calculated basic vulnerability severity level score, which is expressed as: .

[0016] In a second aspect, a vulnerability severity level discrimination system based on clustering analysis includes: A data acquisition module configured to acquire network vulnerability data; A preprocessing module configured to perform data preprocessing on the acquired network vulnerability data; A clustering module configured to perform clustering analysis on the preprocessed data using a clustering algorithm to obtain reduced-dimensional clustering data; A feature extraction module configured to extract attack vector features from the clustering data using a convolutional neural network model; A prediction module configured to repeatedly extract attack vector features and predict the vulnerability severity level score based on the heatmap weight.

[0017] In summary, the present invention has the following beneficial technical effects: 1. The present invention formulates a specific data preprocessing method, a vulnerability clustering method, and an adaptive attention mechanism for judging the vulnerability severity level, and can achieve higher accuracy in judging the vulnerability level based on the common vulnerability scoring system with a smaller number of benchmark labels. In this task, it can achieve the performance of a pre-trained model with a maximum of 200,000 preset label amounts compared to the benchmark labels of the comparative invention.

[0018] 2. The present invention innovatively proposes a new vulnerability severity level judgment model, uses a large natural language processing model to realize the judgment of the vulnerability severity level based on the common vulnerability scoring standard, combines a convolutional neural network, optimizes the BERT baseline attention using self-adaptive feature attention, and uses vulnerability data for retraining supervised learning to achieve the effect of accurately predicting attack vectors.

[0019] 3. While improving the efficiency of vulnerability assessment, it significantly enhances the accuracy of assessment results and reduces the subjective bias of manual judgment. At the same time, this method can automatically update vulnerability information in real time, adapt to the dynamically changing network security environment, further improve the timeliness and effectiveness of vulnerability severity level judgment, help enterprises or organizations make more reasonable and scientific decisions in the process of handling vulnerabilities in corresponding professional fields, and thus improve the overall network security protection level. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 It is a scatter diagram of the application of the present invention using cluster analysis; Figure 2 It is a heat map of the application of the present invention using cluster analysis; Figure 3 It is a schematic diagram of the error comparison between the method of the present invention and the comparative method; Figure 4 It is a schematic diagram of the error distribution of the method of the present invention Figure 5 It is a schematic diagram of the error distribution of the comparative method; Figure 6 It is a schematic diagram of the error distribution of the prediction of high-risk vulnerabilities by the method of the present invention; Figure 7 It is a schematic diagram of the comparison of the year density fitting curves between the method of the present invention and the real data; Figure 8 It is a schematic diagram of the architecture of the method of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0021] The present invention will be further described in detail below with reference to the accompanying drawings.

[0022] Example 1 Referring to Figures 1 - 8 , a method for discriminating the severity level of vulnerabilities based on cluster analysis in this embodiment includes: Obtain network vulnerability data; Perform data preprocessing on the obtained network vulnerability data; Use a clustering algorithm to perform cluster analysis on the preprocessed data to obtain dimension-reduced clustering data; Use a convolutional neural network model to extract attack vector features from the clustering data; Based on the heat map weight, predict the vulnerability severity level score by repeatedly extracting attack vector features.

[0023] Specifically, it includes the following steps: For the convenience of understanding the essence of the present invention, the main technical parameters involved in the present invention are first defined and described, as shown in Table 1, Table 1 Symbol Definition and Description of Main Technical Parameters Symbol Definition Interpretation Instructions SSE Sum of Squared Errors of Vulnerability Data Clustering Analysis <![CDATA[SC i > Silhouette Coefficient of Vulnerability Data <![CDATA[σ i > Average Distance between Different Categories of Clustering Results <![CDATA[L MASK > Cross - Entropy Loss between the Actual Mark and the Predicted Mark at the Masking Position Attention(Q, K, V) Three Layers of the Attention Mechanism <![CDATA[Conv(H, W V )]]> Convolution Result of Vulnerability Features <![CDATA[Feature i > Vulnerability Feature Vector <![CDATA[W v V + b v > Linear Transformation of Vulnerability Features ReLU(Z) Non - linear Activation of Vulnerability Features <![CDATA[P o-h > One - Hot Probability of Vulnerability Attack Vector <![CDATA[Loss c-e > Model Training Loss Rate Next, the software vulnerability scenarios applicable to this embodiment will be described.

[0024] Step 1: Obtain data.

[0025] Among them, to conduct a judgment on the severity level of CVSS vulnerabilities, public vulnerability characteristic information is required. Generally, software developers can submit vulnerability-related characteristic information through the National Vulnerability Database of the United States. The submitted information will go through its official channels, and specialized experts in the corresponding fields will conduct research on the actual severity level of the vulnerabilities, summarize relevant vulnerability characteristics, and publish them. CVSS is the Common Vulnerability Scoring System, a vulnerability assessment standard widely used in the field of computer security. We selected two datasets for experiments. One is a large-scale vulnerability information dataset that pays more attention to domain characteristics, and the other is a small-scale vulnerability information dataset that pays more attention to vulnerability details. Both datasets are based on the Common Vulnerability Scoring System 3.x standard, and each vulnerability includes eight attack vector strings and vulnerability-related information.

[0026] The present invention collected past public vulnerability information data (from January 2018 to December 2024) from the National Vulnerability Database of the United States. After data processing and data cleaning, a total of 126,607 valid data were extracted, including characteristic information such as vulnerability ID, vulnerability professional field, vulnerability description, vulnerability submitter, Common Vulnerability Scoring System score, and vulnerability database. We divided them into a training set of approximately 120,000 and a test set of approximately 6,000 as a large-scale dataset. Another dataset comes from the 7th China Computer Federation Open Source Contest. The training set contains 5,624 entries, and the test set contains 710 entries as a small-scale dataset.

[0027] Step 2: Data preprocessing.

[0028] When conducting data cleaning, it is usually necessary to remove some unnecessary or irrelevant data columns to make the dataset more concise and meet the analysis requirements. When conducting the task of judging the severity level of vulnerabilities, the present invention discards useless data columns such as "CVE-ID", "Issue_Url_old", "Issue_Url", and "Repo_new" in the vulnerability information of the dataset. Although these characteristic information may be useful in some scenarios, in many cases, they do not directly contribute to vulnerability analysis or the study of vulnerability relevance. Therefore, these columns can be selected to be discarded, and vulnerability characteristic description information, the field where the vulnerability is located, attack vectors, Common Vulnerability Scoring System scores, etc. are retained. Then, directly remove the label mixing errors or chaos and irrelevant garbled characters.

[0029] After the vulnerability data is cleaned, the present invention uses stemming and part-of-speech restoration for natural language processing. Stem extraction is the process of restoring a word to its basic form or root, which is widely used in natural language preprocessing. In the CVSS vulnerability severity level assessment, we perform stemming on CVE-related feature information to help extract key words from the vulnerability description, thereby simplifying the language expression of vulnerability feature information and improving analysis efficiency. CVE is a public vulnerability and exposure, which is public network vulnerability data. Part-of-speech restoration is the process of converting a word into its standard word form, which takes into account the contribution of the grammatical role of the word. Unlike stemming, part-of-speech restoration not only removes affixes, but also adjusts its morphology according to the grammatical role of the word in the sentence, ensuring that the word becomes its most basic and standardized form, which is also widely used in natural language preprocessing. The present invention combines the best-performing text preprocessing method. In order to select the added words, this paper sorts them according to the frequency of their appearance in the description and selects the top n words. In order to avoid redundancy, this paper only considers words that appear only in the description, and does not consider words that appear in the default vocabulary. Considering the presence of software versions and code snippets in some data descriptions, this paper uses regular expressions to filter numbers and special characters. This approach reduces the "noise" added by vocabulary, because the filtered data is irrelevant to category classification and may eliminate the importance of relevant added words.

[0030] Node vectorization is performed on the data after natural language processing so that the model can correctly recognize natural language and help the model understand semantics, context, and the relationship between words. The present invention uses a bidirectional encoder of the pre-trained model BERT. The processed data information is directly input into the model to obtain the data vector after node vectorization.

[0031] Step 3: Perform cluster analysis to find out the internal connections of the vulnerabilities.

[0032] Among them, the K-means clustering algorithm in the present invention is a clustering algorithm based on sample set division. K-means clustering divides the sample set that has undergone data preprocessing into K subsets to form K classes, and divides the samples into K classes, minimizing the distance from the center of each sample to the class to which it belongs, so that each sample belongs to only one class. In subsequent papers, three algorithms were applied to determine whether K-means clustering is effective as an indicator. The present invention reduces the source data set to two dimensions and then performs a principal component analysis algorithm on it, and performs standardization processing on it to find out the internal connections between the vulnerability data. The clustering evaluation rules used are as follows: Step 3.1 Calculate the error sum of squares rule.

[0033] Error square sum rule: The smaller the error square sum, the better the clustering effect.

[0034] Among them, represents the observed value, represents the predicted value before summing the samples separately. This formula represents the gap between the predicted value and the observed value of the CVE information.

[0035] Step 3.2 Calculate the silhouette coefficient.

[0036] The silhouette coefficient can be used to evaluate the impact of different algorithms or different algorithm running methods on the clustering results based on the same original data, and combines cohesion and separation. The value of the silhouette coefficient is normalized to [-1, 1], and the closer the value is to 1, the better the cohesion and separation.

[0037] Among them, represents the cohesion of the sample point, represents the other sample points belonging to the same category as the sample and the distance represents the distance. The calculation method of is similar to but it needs to traverse other clusters to obtain multiple values and select the minimum value as the final result. It can be seen that

[0038] Step 3.3 Calculate the Calinski-Harabasz index.

[0039] The Calinski-Harabasz index is defined as the ratio of between-component scatter to within-component scatter. The larger this score, the better the clustering effect. The separation degree of the dataset is measured by calculating the sum of the squares of the distances between the center points of each class and the center point of the dataset, and it is obtained from the ratio of separation degree to compactness.

[0040] Among them, the formula is expressed as: Among them is the Calinski-Harabasz index. is the between-class scatter matrix, representing the degree of dispersion between different class centers. is the within-class scatter matrix, representing the degree of dispersion of samples within the same class. is the total number of vulnerability data. is the number of classes.

[0041] Through these three comparison methods, the present invention can determine a better combination of the dimensionality reduction number and the number of clusters. However, considering the actual situation and combining with the corresponding dataset, when the dimensionality reduction is too low, clustering will result in too few data in a certain cluster of the clustering. Moreover, the invention mainly uses clustering for data analysis to provide a thinking direction and solution idea for the subsequent establishment of a prediction model. Therefore, the present invention selects the case where PCA is reduced to two dimensions and the number of clustering clusters is 4. PCA is the principal component analysis, which is an unsupervised learning method aiming to reduce high-dimensional vulnerability information data to a low-dimensional space and capture the connections between vulnerability information data. Using clustering helps the present invention analyze the correlation between each part of the dataset and the CVSS scoring rules. Through these three comparison methods, the present invention can determine a better combination of the dimensionality reduction number and the number of clusters.

[0042] Among them, directly use the sklearn toolkit to perform PCA(2) dimensionality reduction. Considering that the number of categories of vulnerability-related text data is small, too high a dimensionality reduction number will cause the loss of the actual meaning of the data. The present invention has tested PCA(1 to 4) during the experiment and found that the effect of PCA(2) is the best and will not overly lose the meaning of the original vulnerability data.

[0043] The sum of squared errors measures the sum of the squares of the distances from each sample inside the cluster to the cluster center, reflecting the tightness of the cluster. As the number of clusters increases, the value of the sum of squared errors will decrease; the silhouette coefficient evaluates the distances between each sample and its own cluster and the nearest cluster, thereby measuring the tightness and separation degree of the clustering. The larger the value, the better the clustering effect, but too high a value of this numerical will cause too small a distance between clusters; the Calinski-Harabasz index evaluates the ratio of the between-class difference to the within-class difference in the clustering result. The larger the value of the Calinski-Harabasz index, the better the clustering effect. When calculating the three comparison numerical values, the present invention calculates the number of clusters from 2 to 10. The experiment shows that the sum of squared errors reaches the lowest point when the number of clusters is 3 or 4, the silhouette coefficient reaches the maximum when the number of clusters is 4, and the Calinski-Harabasz index reaches the maximum when the number of clusters is 4. In summary, the present invention can judge that for this dataset, the effect is better when the number of clusters is 4 during vulnerability clustering.

[0044] Step 4, construct a model.

[0045] Among them, the specific implementation process of the model of the present invention is described, and the complete implementation steps included are as follows: When considering the evaluation method of the vector string of the Common Vulnerability Scoring System, most methods construct a separate deep learning classifier for each vector string, and then follow the multi-task learning paradigm to generate a unified form of submission features for a specific metric classifier. At the same time, the present invention analyzes and considers the idea of dynamically adjusting the overall vector string. On this basis, the present invention proposes that the single vector string metric weights obtained through clustering analysis can establish an adaptive close connection between classifiers.

[0046] To propose a method for exploring the close connection method observed in the vector string of the Common Vulnerability Scoring System of the present invention, the present invention utilizes the natural language processing pre-trained model BERT combined with a text convolutional neural network to extract the features of different metrics in the vector string, and allows the model to perform retraining and learning iteratively. Therefore, the present invention is divided into three modules, namely the masked language module, the adaptive attention mechanism module, and the feature extraction and prediction re-input module.

[0047] First, it is the masked language module. The present invention randomly replaces some of the metric labels in each input attack vector with special tokens, and the model will predict the tokens marked as based on the context. The objective function of the masked language module L_MASK is the cross-entropy loss between the actual token and the predicted token at the masked position: Then, it is the adaptive attention of the model. For each layer, including the L layers in the encoder, the output representation H of the previous layer is passed through the corresponding projection matrices , , layer by layer, allowing the model to focus on different parts of the input sequence when processing a certain token, and then obtaining the linear projections as the sorted queue Q, the key value K, and the value V. The attention mechanism can generally be described as mapping a query and a set of key-value pairs to an output. The multi-head self-attention mechanism and the position feed-forward fully connected network, and each sub-layer adopts a residual connection and layer normalization, that is, by calculating the relationship between all tokens in the sequence to assign weights to each token.

[0048] Among them, the compatibility is measured by the dot product and then scaled by the dimension and further normalized using the softmax function.

[0049] The third part is the feature extraction and prediction re-input module. The present invention innovatively uses a convolutional neural network for feature extraction in the attack vector prediction. In the input layer, a text token containing a vulnerability description is input , first, vectorize the text according to the Transformer model structure to obtain the vector representation of each attack vector. Then, input it into the convolutional layer to filter each attack vector, and adopt the method of hidden feature extraction. Take the attack vector output by the present invention as the input of the convolutional text neural network, and perform convolutional operations on the text sequence through multiple adaptive convolutional kernels to extract the features of the attack vector, and a sequence of feature maps can be obtained , use the clustering attack vector connection weights obtained in Section 3 as the bias term of the convolution: Next, enter the pooling layer. Through max pooling or average pooling operations, extract the most representative relevant input vulnerability text local features from the convolutional layer, that is, high-dimensional vulnerability information features. Then, the training weights of the model and the convolutional layer weights of TextCNN will be updated together. The high-dimensional vulnerability information features obtained through convolution will be re-input into the present invention, so that the pre-trained model can adaptively adjust its parameters according to the judgment of the vulnerability severity level, etc., to improve the performance on the task, where is the pooling result of the i-th convolutional kernel, is the number of convolutional kernels, and thus the feature vector , contains: Then is the fully connected layer. Send the pooled attack vector features into the fully connected layer for classification operations, and perform non-linear transformation through the adaptive activation function. The connected feature vector , where is the weight matrix of the fully connected layer, is the bias term, which can change dynamically according to the adaptive attention, is the attack vector feature after non-linear transformation: Finally is the output layer. Use the softmax function to perform multi-classification on each attack vector of the common vulnerability scoring system to obtain the one-hot probability: Among them, the one-hot probability is mainly applied in computer science and machine learning and is obtained from one-hot encoding. One-hot encoding is a method of converting categorical data into a numerical form. Its core idea is to convert each category into a fixed-length vector. In this vector, a tensor with a value of 1 represents the category, and the rest are 0. Thus, the unique probability can be calculated: Among them, is the probability of the attack vector category and is the logical score of this attack vector category and is the total number of this attack vector category.

[0050] The present invention uses a common cross-entropy loss function to measure the difference between the prediction and the true label as the loss function. Let be the true data label, L be the number of attack vector labels in the Common Vulnerability Scoring System, be the probability vector of the attack vector in the Common Vulnerability Scoring System predicted by the model, and the loss is: The dynamic parameters existing in the model, including the feature map dimension, the model learning rate, the convolution kernel length of the present invention, and the bias term parameter of the fully connected layer in the present invention, are updated by using the backpropagation algorithm of the Adam optimizer.

[0051] Next, the beneficial effects of the method of the present invention are introduced, and the experimental comparison data of the method of the present invention and the comparative method are provided.

[0052] The experimental scenario of the present invention is set as follows: Public vulnerability information data (from January 2018 to December 2024) was collected from the network. After data processing and data cleaning, a total of 126,607 valid data were extracted, including feature information such as vulnerability numbers, vulnerability professional fields, vulnerability descriptions, vulnerability submitters, scores of each item based on the Common Vulnerability Scoring System, and vulnerability libraries. We divided them into a training set of approximately 120,000 and a test set of approximately 6,000 as a large-scale data set. Another data set comes from the 7th China Computer Federation Open Source Innovation Contest. The training set contains 5,624 data, and the test set contains 710 entries as a small-scale data set. The other experimental parameters related to the present invention are as follows: (1) The overall model training uses 3 epochs, the sample batch size is 32, and the learning rate is 2e-5; (2) The optimizer is Adam; (3) The convolutional neural network uses a hidden layer size of 768, 12 adaptive convolution kernels, 4 filters, and the sizes are [3, 4, 5]; (4) Since we finally need to obtain specific predicted label values for vulnerability severity level judgment, random sample batches are not used.

[0053] Step 5: Use the model to repeatedly extract features from the data.

[0054] Among them, the features of the attack vector are extracted by performing convolution operations on the text sequence with multiple adaptive convolution kernels.

[0055] Step 5.1 Text Convolution The convolution operation is performed on the vulnerability data after node vectorization. The present invention uses three convolution kernels of different sizes to extract language model features of different lengths. The role of the convolution kernel is to perform local perception on each piece of text, similar to extracting patterns in a local area.

[0056] Among them is the vulnerability text vector data matrix, is the convolution kernel, is the bias term.

[0057] Step 5.2 Max Pooling For the in Step 4.2, perform max pooling operation, take the maximum value in the local area of each feature map to reduce the dimension of the features, so as to extract the most important features for the model's reinforcement learning.

[0058] Among them is the vulnerability data feature map, are the values in the pooling window.

[0059] Step 6, Automated Vulnerability Severity Judgment System.

[0060] Combine the models and vulnerability features proposed in Steps 4 and 5 and extract them repeatedly, and transfer the weights shown in the heat map into the parameters and input them into the model. Each sub-model predicts the eight indicators of AV, AC, PR, UI, S, C, I, and A in the CVSS vector string respectively, and then combines the weights obtained from the heat map in this article to perform label prediction on the vulnerability descriptions of the CVSS score dataset to be predicted. Each indicator has different labels, and different labels have different effects on the CVSS score. Then, through standardized calculation, the final vulnerability severity level score predicted by the prediction result is obtained. Among them, the vector string is the attack vector, which is the standard for determining the vulnerability level in the Common Vulnerability Scoring System.

[0061] Among them, the category of each attack vector is obtained from the one-hot probability. From the categories of the eight attack vectors, the score calculation formula proposed by the Common Vulnerability Scoring System can be used for calculation: Among them, AV, AC, S, C, I, and A are the attack vector categories, and their standardized values are defined by the vulnerability scoring system and have been reflected in the dataset. 、 、 are the sum of the attack vector categories, Is the sum of the attack vector categories.

[0062] Calculated through it Is the basic score of the vulnerability scoring system. This score has been standardized in the interval [0, 10], and this score is the vulnerability severity level score of a specific vulnerability. According to the vulnerability scoring system, a basic score of [0, 4) is defined as a low-risk vulnerability, [4, 7) as a medium-risk vulnerability, and [7, 10] as a high-risk vulnerability, and the determination of the vulnerability severity level is thus obtained.

[0063] As a further implementation method, Based on the sum of various categories of attack vectors predicted in the previous steps and the basic score of the vulnerability severity level calculated, the exploitability evaluation score and the environmental impact evaluation score of the vulnerability can be calculated.

[0064] Where PR and UI are attack vector categories, Is the basic score of the vulnerability calculated in the previous step, Is the exploitability evaluation score of the vulnerability, Is the environmental impact evaluation score of the vulnerability.

[0065] From the relevant actual values, overall understand the life cycle assessment, environmental assessment, etc. of a certain vulnerability. Then, according to the basic score, exploitability evaluation score, environmental impact evaluation score of the vulnerability in the actual year of the dataset and the relevant values predicted and calculated by the present invention, draw a density fitting curve of the vulnerability severity level over the years. By this, not only can we obtain the different professional fields, threat levels, etc. involved by the vulnerability in different years, but also we can obtain that the vulnerability severity level predicted by the present invention is accurate enough.

[0066] To objectively evaluate the effectiveness of the proposed method, the present invention selects four language models for comparison.

[0067] Figure 1 Is the scatter diagram of the application of clustering analysis in the present invention. Figure 1 In (a) is the scatter diagram after principal component analysis dimensionality reduction and standardization, Figure 1(b) in the figure is a scatter plot after K-means clustering, in which purple scatter points represent the severity of vulnerability information data, black scatter points represent the basic score connection clustering between vulnerability information data, red represents the environmental score connection clustering between vulnerability information data, and light brown represents the impact score connection clustering between vulnerability information data. The method of the present invention uses K-means clustering analysis. After importing the data set, data cleaning is performed to discard useless data columns such as "CVE-ID", "Issue_Url_old", "Issue_Url", and "Repo_new", leaving "vector string", "BaseScore", "ExploitabilityScore", and "ImpactScore". Considering that there are fewer data columns that need to be analyzed for correlation, the method of the present invention performs PCA (2) dimensionality reduction on the data, and then performs standardization processing, and obtains a scatter plot of the clustered data. Then, the clustering evaluation index is analyzed by combining the error square sum accuracy, silhouette coefficient, and Kalinsky-Harabas index. Considering that the general vulnerability scoring system has four score types, the present invention finally decides to use clustering category 4 for clustering. The data trends obtained after clustering are relatively uniform because the provided serial number data is used for clustering and dimensionality reduction without random shuffling, but it can be seen that the clustering effect is good.

[0068] Figure 2 This is the heat map of cluster analysis applied in this invention. After K-means cluster analysis, this paper associates the use of clustered data for heat map visualization data analysis to observe the relationship between each indicator in the attack vector of the general vulnerability scoring system and multi-dimensional data with different scores. From the heat map, it can be found that the degree of influence of the attack vector on the vulnerability severity score, that is, the weight of its indicator on the final vulnerability severity score.

[0069] Figure 3 Schematic diagram of error comparison between the method of the present invention and the comparative method. Figure 3 (a) is a schematic diagram showing the error comparison between the method of the present invention and the comparative method for basic scores of vulnerability severity level identification. Figure 3 (b) is a schematic diagram showing the error comparison between the method of the present invention and the comparative method for the environmental score of vulnerability severity level discrimination. Figure 3Figure (c) is a schematic diagram of the error comparison of the influence scores of the method of the present invention and the comparative method for judging the severity level of vulnerabilities. The experimental results show that on the complete and limited training data sets, the method of the present invention is significantly better than BERT and some pre-trained large language models that also perform target task training. In the task of judging the severity level of vulnerabilities based on the Common Vulnerability Scoring System, especially in multi-class prediction, the model CICNet of the method of the present invention has a significant improvement in accuracy, F1 score, precision, and average accuracy compared with other benchmark pre-trained large models. The highest F1 score of the task reaches 96.6%, and the average improvement in error analysis is about 40%.

[0070] Figure 4 and Figure 5 is a schematic diagram of the error distribution of the method of the present invention and the comparative method. Figure 4 In (a), it is a schematic diagram of the error distribution of the basic score of the method of the present invention for judging the severity level of vulnerabilities, Figure 4 in (b), it is a schematic diagram of the error distribution of the environmental score of the method of the present invention for judging the severity level of vulnerabilities, Figure 4 in (c), it is a schematic diagram of the error distribution of the influence score of the method of the present invention for judging the severity level of vulnerabilities. Figure 5 In (a), it is a schematic diagram of the error distribution of the basic score of the comparative method for judging the severity level of vulnerabilities, Figure 5 in (b), it is a schematic diagram of the error distribution of the environmental score of the comparative method for judging the severity level of vulnerabilities, Figure 5 in (c), it is a schematic diagram of the error distribution of the influence score of the comparative method for judging the severity level of vulnerabilities. It clearly shows the comparison gap between the mean squared error graph of the scores of the Common Vulnerability Scoring System predicted by the method of the present invention and the mean squared error graph of the scores of the selected comparative method. It is not difficult to see that the error of the severity level score predicted by the method of the present invention is small and the error distribution is relatively smooth, indicating that the method of the present invention has high prediction accuracy and strong stability. However, there is still room for improvement in the influence score of the method of the present invention. Considering that its calculation formula mainly comes from the three indicators C, I, and A in the attack vector, it can be seen that the prediction accuracy of the method of the present invention for these three indicators can still be improved.

[0071] Figure 6Schematic diagram of the prediction error distribution of high - risk vulnerabilities in the method of the present invention. The error of the vulnerability level predicted by the method of the present invention is small, and the error distribution is relatively smooth, indicating that the method of the present invention has high prediction accuracy and strong stability. However, there is still room for improvement in the impact score of the method of the present invention. Considering that its calculation formula mainly comes from the three indicators C, I, and A in the attack vector string of the Common Vulnerability Scoring System, it can also be seen from Figure 2 that the prediction accuracy of the method of the present invention for these three indicators is poor. In the Common Vulnerability Scoring System, vulnerabilities with a basic score in the range of [7, 10] are judged as high - risk vulnerabilities. The squared error for judging the severity of high - risk vulnerabilities is very small, and the prediction effect is very good. Thus, the method of the present invention can achieve a very good judgment effect on high - risk vulnerabilities without manual intervention, greatly improving the speed and accuracy of security detection.

[0072] Figure 7 Schematic diagram for comparing the year - density fitting curves of the method of the present invention and real data. Figure 7 In (a) is the year - fitting curve diagram of the basic score for judging the severity level of vulnerabilities by the method of the present invention, Figure 7 In (b) is the year - fitting curve diagram of the basic score for judging the severity level of vulnerabilities of real data. By plotting the actual year - density fitting curve and the predicted - year Common Vulnerability Scoring System score - density fitting curve, the differences between the actual year - density fitting curve and the predicted - year Common Vulnerability Scoring System score - density fitting curve are compared, and the severity levels of vulnerabilities are marked. It can be seen from the density fitting curve diagram of experimental data that the predicted - year density fitting curve is basically consistent with the actual year - density fitting curve in trend, indicating the accuracy and stability of the model of the present invention in capturing data patterns. The undulations of the predicted curve are in good synchronization with the actual curve, and the error range is within an acceptable range, indicating that the model of the present invention can not only reasonably simulate the past data distribution but also effectively infer future trends. This result further verifies the prediction ability of the model and proves the effectiveness of the experiment.

[0073] Embodiment 2 This embodiment provides a system for judging the severity level of vulnerabilities based on clustering analysis, including: A data acquisition module, configured to: A computer - readable storage medium, in which multiple instructions are stored, and the instructions are adapted to be loaded and executed by a processor of a terminal device for the method of judging the severity level of vulnerabilities based on clustering analysis.

[0074] A terminal device, including a processor and a computer - readable storage medium, the processor is used to implement each instruction; the computer - readable storage medium is used to store multiple instructions, and the instructions are adapted to be loaded and executed by the processor for the method of judging the severity level of vulnerabilities based on clustering analysis.

[0075] The above are all preferred embodiments of the present invention, and the protection scope of the present invention is not limited thereby. Therefore, all equivalent changes made according to the structure, shape, and principle of the present invention shall be covered within the protection scope of the present invention.

Claims

1. A vulnerability severity level determination method based on cluster analysis, characterized in that: include: Obtain network vulnerability data; Perform data preprocessing on the acquired network vulnerability data; Use clustering algorithm to perform cluster analysis on the preprocessed data to obtain cluster data after dimension reduction; Use the convolutional neural network model to extract attack vector features from clustered data; By repeatedly extracting attack vector features, the vulnerability severity score is predicted based on the heat map weight.

2. According to the method for determining the severity level of a vulnerability based on cluster analysis according to claim 1, it is characterized in that: The obtaining of network vulnerability data includes obtaining a network vulnerability data set, including a large-scale vulnerability information data set and a small-scale vulnerability information data set; each vulnerability includes 8 attack vector character strings and vulnerability information.

3. According to the method for determining the severity level of a vulnerability based on cluster analysis according to claim 2, it is characterized in that: The obtained network data is preprocessed, including data cleaning, natural language processing and node vectorization, wherein the natural language processing includes stem extraction and part-of-speech restoration, node vectorization of the data is performed after natural language processing, and the BERT bidirectional encoder is used to process the data to obtain the data vector after node vectorization.

4. According to the method for determining the severity level of a vulnerability based on cluster analysis according to claim 3, it is characterized in that: The method uses a clustering algorithm to perform cluster analysis on the preprocessed data to obtain clustered data after dimensionality reduction, including using a K-means clustering algorithm, which is a clustering algorithm based on sample set division. K-means clustering divides the sample set after data preprocessing into K subsets to form K classes, and divides the samples into K classes, minimizing the distance from the center of each sample to the class to which it belongs, so that each sample belongs to only one class, and determines the most appropriate number of fern clusters and number of dimensionality reduction by respectively calculating the error sum of squares rule, silhouette coefficient and Karlinsky-Harabas index.

5. According to the method for determining the severity level of vulnerabilities based on cluster analysis according to claim 4, it is characterized in that: The method uses a convolutional neural network model to extract attack vector features from the clustered data, including constructing a convolutional neural network based on a natural language processing pre-training model BERT and a text convolutional neural network, wherein a masked language mechanism is used to replace the input attack vector data with a special tag, and the cross entropy loss between the actual tag and the predicted tag at the masked position is calculated, and a multi-head self-attention mechanism is used to calculate the relationship between all tags in the sequence to assign a weight to each tag, wherein the cross entropy loss between the actual tag and the predicted tag at the masked position is: Marked with X .

6. According to the method for determining the severity level of vulnerabilities based on cluster analysis according to claim 5, it is characterized in that: The method uses a convolutional neural network model to extract attack vector features from clustered data, and also includes performing node vectorization based on a transformer mechanism to obtain a vector representation of the attack vector, and then using multiple adaptive convolution kernels to perform a convolution operation on the sequence to obtain a feature graph sequence. After using the distribution weight calculated by the multi-head attention mechanism as a bias item of the convolution, it is sent to the pooling layer for pooling operation to obtain high-dimensional vulnerability information features. Finally, after using an adaptive activation function to perform nonlinear transformation on the features, the copper leakage attack vector is multi-classified to obtain a unique hot probability.

7. The vulnerability severity level determination method based on cluster analysis according to claim 6 is characterized in that: The method repeatedly extracts attack vector features and predicts vulnerability severity scores based on heat map weights, including using three convolution kernels of different sizes to extract language model features of different lengths and performing maximum pooling operations, wherein the category of each attack vector is obtained by one-hot probability calculation, and then the vulnerability score is calculated using a calculation formula, and the vulnerability severity level is judged based on the vulnerability score.

8. According to the method for determining vulnerability severity based on cluster analysis according to claim 7, it is characterized in that: The vulnerability score is calculated using the calculation formula, including calculating the sum of the attack vector categories by defining the standardized fixed values ​​of the attack vector categories, and then obtaining the basic score of the vulnerability severity level, which is expressed as: AV, AC, S, C, I, and A are attack vector categories, which are defined by the vulnerability scoring system and have standardized values, which are reflected in the dataset. , , is the total of the attack vector category, is the sum of attack vector categories.

9. The vulnerability severity level determination method based on cluster analysis according to claim 8 is characterized in that: The method of repeatedly extracting attack vector features and predicting vulnerability severity scores based on heat map weights also includes calculating vulnerability exploitability evaluation scores and environmental impact evaluation scores based on the predicted attack vector categories and the calculated vulnerability severity basic scores, which are expressed as: , Among them, PR and UI are attack vector categories. The basic vulnerability score calculated in the previous step. Score the vulnerability exploitability evaluation. Assign a score to the vulnerability's environmental impact assessment.

10. A vulnerability severity level determination system based on cluster analysis, characterized in that: include: The data acquisition module is configured to acquire network vulnerability data; The preprocessing module is configured to perform data preprocessing on the acquired network vulnerability data; The clustering module is configured to perform cluster analysis on the preprocessed data using a clustering algorithm to obtain clustered data after dimension reduction; The feature extraction module is configured to extract attack vector features from the clustering data using a convolutional neural network model; The prediction module is configured to predict the vulnerability severity score based on the heat map weight by repeatedly performing attack vector feature extraction.