A network security situation assessment method based on denoising autoencoder kernel density estimation

The unsupervised evaluation of network situation data through the noise reduction self-encoding kernel density estimation method solves the problem of low accuracy of traditional methods in large-scale network environments, and achieves a more efficient network security situation evaluation.

CN114547608BActive Publication Date: 2025-08-22DALIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210108654.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-28
Publication Date
2025-08-22
Estimated Expiration
2042-01-28

AI Technical Summary

Technical Problem

Traditional network security situation evaluation methods are not very accurate in large-scale network environments, especially due to the high-dimensionality and nonlinear characteristics of network situation data, it is difficult to improve the evaluation efficiency and accuracy of existing models.

Method used

Unsupervised network security situation evaluation method based on noise reduction and self-coding kernel density estimation is adopted. Dimension reduction and feature extraction are performed through noise reduction and self-coding network, and threat occurrence probability is estimated by combining kernel density estimation to achieve network security situation evaluation without data labels.

Benefits of technology

The accuracy and recall rate of network security situation evaluation have been improved, and the accuracy and recall rate of the model in real network environment have been increased by 3.51% and 5.99% respectively, significantly improving the evaluation effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114547608B_ABST
    Figure CN114547608B_ABST
Patent Text Reader

Abstract

The present invention discloses a network security situation assessment method based on denoising autoencoder kernel density estimation, which belongs to the field of computer network security and comprises the following steps: obtaining network flow data as situation data; preprocessing the situation data; dividing the preprocessed situation data into training set data and test set data according to a proportion; constructing a denoising autoencoder network and a kernel code estimation model based on the training set data; inputting the test set data into the denoising autoencoder network and the kernel code estimation model in sequence to obtain the threat occurrence probability of the network flow data; performing a security assessment on the network situation based on the threat occurrence probability of the network flow data to determine the level of the network security situation; utilizing the ability of the denoising autoencoder network to process redundant information and nonlinear feature learning to reduce the dimension of the network situation data and extract potential situation features, and combining the advantage of parameter-free estimation to propose kernel density estimation to perform density probability estimation on the potential features to obtain the threat occurrence probability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer network security, and in particular to a network security situation assessment method based on denoising autoencoder kernel density estimation. Background Art

[0002] The rapid development of cyberspace has provided convenience and benefits to humanity, but it has also posed challenges to network security. While network architectures often incorporate security measures such as firewalls and intrusion detection systems to detect and prevent attacks, these often generate numerous alerts and false positives. Network analysts struggle to effectively interpret these alerts, leading to the need for effective network security situation assessment methods to quantitatively analyze and evaluate the security posture of network systems. This approach, combined with a comprehensive understanding of the threat landscape, provides a comprehensive understanding of network security status, and aids decision-making for network managers.

[0003] Network situational awareness (CSA) is a key research technology in the information security field, transforming passive defense into active awareness. This concept is an extension of battlefield situational awareness. Based on the situational awareness model proposed by Base, it is essential for situation assessment research. Currently, many theories applied to network security situation assessment include fuzzy theory, evidence theory, Markov models, and Bayesian networks. These models primarily address situational elements generated by security tools, generating alert logs. While these models have demonstrated promising results in small- and medium-scale network applications, they still have limitations. For example, the basic probability distribution in evidence theory requires expert experience, resulting in inconsistent assessment results. Furthermore, Markov and Bayesian models, based on probability theory and knowledge reasoning, require extensive prior knowledge, resulting in high assessment costs and difficulty improving situation assessment efficiency. These limitations are becoming increasingly apparent in large-scale network environments. Neural networks, due to their advantages in solving complex problems through nonlinear mapping, are widely used in various fields and have therefore garnered significant attention in the field of situational awareness.

[0004] Xie Lixia et al. [1-2] First proposed to use BP neural network for network security situation assessment, and then used cuckoo optimization algorithm to optimize the weight of BP neural network to obtain an improved situation assessment model. Experimental results show that the improved method has faster convergence speed and better assessment effect than traditional BP neural network. [3] In view of the exponential growth of the dimension of traditional neural networks, which leads to an increase in computational complexity and is not suitable for large-scale complex networks, a quantitative assessment method for the network security status of wirelessly connected intelligent robot groups based on convolutional neural networks is proposed, with an accuracy rate of 95%; [4]Based on the reality that labeled situation data is difficult to obtain, in order to avoid the BP neural network training relying on labels, a situation assessment method based on deep autoencoder network is proposed. The situation data is trained in a semi-supervised learning manner to establish a situation assessment model. Experiments show that the root mean square error of this method is significantly smaller than that of BP neural network, but the evaluation process requires the participation of experts and lacks analysis of the situation assessment effect. [5-6] Taking network flow as the main situational element, autoencoder variants were applied twice to the field of network situational awareness. First, it was proposed to combine the variational autoencoder with the generative adversarial network to establish a threat testing model to evaluate the network security threat situation. However, the threat testing model is relatively complex and has high hardware requirements. The deep autoencoder model proposed subsequently performs binary and quinary classification of network anomaly types. The proposed model has high classification accuracy, but the data set used is too old and is not suitable for the current complex network environment.

[0005] Traditional traffic analysis reveals that all network events are reflected in traffic. Normal network traffic and abnormal network traffic have obvious differences in performance. Therefore, traffic analysis can be used to evaluate the network status. [7] ,Related research on intrusion detection shows that abnormal events rarely occur, ,so the distribution of normal samples and abnormal samples in real networks is ,very uneven. The supervised learning model first needs to label the ,network flow, which is not only time-consuming but also reduces the model ,efficiency.

[0006] To avoid the drawbacks of supervised models, this paper proposes an unsupervised network security situation assessment method using denoising autoencoder kernel density estimation. Autoencoders reduce the dimensionality of high-dimensional network situation data and extract latent features. However, since the autoencoder's output may simply be a copy of the input, supervised models lose their effectiveness. The high-dimensionality and nonlinearity of current network situation data contribute to the low accuracy of traditional network security situation assessment methods. Summary of the Invention

[0007] Aiming at the problem that the current network situation data has high dimensionality and nonlinear characteristics, which leads to the low accuracy of traditional network security situation assessment methods, the present invention discloses a

[0008] A network security situation assessment method based on denoising autoencoder kernel density estimation includes the following steps:

[0009] Obtain network flow data as situation data;

[0010] Preprocessing the situation data; dividing the preprocessed situation data into training set data and test set data according to a ratio;

[0011] Construct a denoising autoencoder network and a nuclear code estimation model based on the training set data;

[0012] The test set data is sequentially input into the denoising autoencoder network and the nuclear code estimation model to obtain the threat occurrence probability of the network flow data;

[0013] Based on the probability of threat occurrence of network flow data, a security assessment of the network situation is conducted to determine the level of network security situation.

[0014] Furthermore, the pre-processing of the situation data includes the following steps: removing repeated feature columns in the situation data and samples with infinite and null values ​​in the fields, and then normalizing the data.

[0015] Furthermore, the process of constructing a denoising autoencoder network and a kernel code estimation model based on the training set data is as follows: the training set data is sequentially input into the denoising autoencoder network to train the denoising autoencoder network and obtain the hidden layer features of the network flow data; the hidden layer features of the network flow data are input into the kernel density model to train the kernel density and construct the denoising autoencoder network and the kernel code estimation model. The specific process is as follows:

[0016] Step 1: Set the maximum number of training sessions;

[0017] Step 2: Initialize the parameters of the denoising autoencoder network;

[0018] Step 3: Input the training set data into the denoising autoencoder network and perform the following calculations on the input data in sequence:

[0019]

[0020]

[0021] y=g(w'h+b') (3)

[0022] x, h, and y represent input layer data, hidden layer data, and output data, respectively; q D represents a random distribution between [0,1], a is the noise factor; f and g represent the nonlinear excitation functions in the encoding and decoding process respectively; w and w' are weight parameters, b and b' are biases;

[0023] Step 4: Calculate the reconstruction error using the reconstruction error calculation formula;

[0024] Step 5: Minimize the reconstruction error; adjust the weight and bias parameters;

[0025] Step 6: Determine whether the training count value k is greater than the set maximum number of training times. If so, the training of the denoising self-programmed network is completed. Otherwise, k = k + 1, adjust the weight and bias parameters, and return to step 3.

[0026] Step 7: Input the training data into the trained denoising autoencoder network again and obtain the hidden layer feature data;

[0027] Step 8: Select the Gaussian function as the kernel function of the kernel density estimation model and set its window width;

[0028] Step 9: Use the hidden layer feature data as input to build a kernel density estimation model through the kernel density estimation formula to obtain the probability density distribution of the training data.

[0029] Furthermore, the process of sequentially inputting the test set data into the denoising autoencoder network and the nuclear code estimation model to obtain the threat occurrence probability of the network flow data is as follows:

[0030] Step 1: Test data Input into the trained denoising autoencoder network and obtain the hidden layer features of the test data

[0031] Step 2: Input the hidden layer features of the test data into the constructed kernel density estimation model to calculate the density value

[0032] Step 3: Pass Find outliers in the test data The smaller the outlier value, the more abnormal the existence;

[0033] Step 4: Invert the outlier value and normalize it to [0, 1]. The inverted and normalized value is used as the probability of threat occurrence.

[0034] Furthermore, the process of performing a security assessment on the network situation based on the probability of threat occurrence of the network flow data and determining the level of the network security situation is as follows:

[0035] Based on the probability of threat occurrence of network flow data, the confidentiality C, integrity I and availability A of network devices are scored and evaluated based on the common vulnerability scoring system to quantify the threat impact value;

[0036] Quantify the network security situation value based on the threat occurrence probability and threat impact value of network flow data;

[0037] Normalized network security situation value is divided into five levels: excellent, good, medium, poor, and critical.

[0038] Furthermore, the calculation formula of the threat impact value is as follows:

[0039]

[0040] Where: E i: threat impact value, i represents the network flow sequence number; C: confidentiality, I: integrity, A: availability.

[0041] A network security situation assessment system based on denoising autoencoder kernel density estimation, comprising:

[0042] Acquisition module: used to obtain network flow data as situation data;

[0043] Preprocessing module: used to preprocess the situation data; divide the preprocessed situation data into training set data, verification set data and test set data according to the proportion;

[0044] Training module: used to build a denoising autoencoder network and a nuclear code estimation model based on training set data;

[0045] Estimation module: used to input the test set data into the denoising autoencoder network and the nuclear code estimation model in sequence to obtain the threat probability of the network flow data;

[0046] Determination module: used to conduct security assessment of network situation based on the probability of threat occurrence of network flow data and determine the level of network security situation.

[0047] By adopting the above technical solution, the present invention provides a network security situation assessment method based on denoising autoencoder kernel density estimation. This method overcomes the limitation of traditional supervised feature learning-based quantitative network security situation assessment methods that rely on data labels for modeling. It utilizes the ability of denoising autoencoder networks to process redundant information and nonlinear feature learning to reduce the dimension of network situation data and extract latent situation features. At the same time, it combines the advantages of parameter-free estimation to propose kernel density estimation to estimate the density probability of latent features to obtain the probability of threat occurrence. This application uses network traffic as the basis for network situation assessment. This application proposes the use of denoising autoencoders to suppress the "copying" tendency of traditional autoencoders. This unsupervised network security situation assessment method combines denoising autoencoders with kernel density estimation to form a denoising autoencoder kernel density estimation model. The improved denoising autoencoder kernel density estimation model can improve sensitivity to abnormal network situations. This application quantifies the security situation value based on the threat occurrence probability to quantitatively assess the network security situation. Using a real network traffic database, the application conducts situation analysis and assessment of network environments with attack behaviors. The analysis shows that the proposed model improves the accuracy and recall rate by 3.51% and 5.99%, respectively, improving the accuracy of the assessment. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0049] Figure 1 is a flow chart of the present invention;

[0050] Figure 2 (a) is a process diagram of the noise reduction autoencoder kernel density estimation model of the present invention; (b) is a training and testing diagram of the noise reduction autoencoder kernel density estimation model of the present invention;

[0051] Figure 3 This is a structural diagram of the noise reduction autoencoding network of the present invention;

[0052] Figure 4 This is the ROC-AUC graph of the OCDAE-KDE model on the test set of the present invention;

[0053] Figure 5 This is the network security situation value result diagram of the Friday test set of the present invention. DETAILED DESCRIPTION

[0054] To make the technical solutions and advantages of the present invention more clear, the technical solutions in the embodiments of the present invention are clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention:

[0055] Figure 1 The flowchart of the present invention is a network security situation assessment method based on denoising autoencoder kernel density estimation, comprising the following steps:

[0056] Situation acquisition stage: obtain network flow data as situation data; preprocess the situation data; divide the preprocessed situation data into training set data, validation set data and test set data according to the proportion; the CICIDS-2017 dataset is selected as the research object in the situation acquisition stage of the denoising autoencoder network and nuclear code estimation model;

[0057] Situational understanding phase: Build a denoising autoencoder network and a nuclear code estimation model based on the training set data;

[0058] Situation Assessment Phase: The test set data is sequentially input into the denoising autoencoder network and the kernel code estimation model to obtain the threat probability of the network flow data. The denoising autoencoder network mainly processes redundant information and nonlinear feature learning on the network flow situation data to reduce the dimension and mine latent features in the hidden layer representation, removing redundant information. Then, combining the advantages of parameter-free estimation, the kernel density model is used to estimate the density probability of the hidden features, thereby deriving the threat probability.

[0059] Based on the probability of threat occurrence of network flow data, a security assessment of the network situation is conducted to determine the level of network security situation.

[0060] Furthermore, the preprocessing of the situation data includes the following steps: deleting duplicate feature columns and samples of fields with infinity and NaN values ​​in the situation data from the pandas library, and then performing normalization processing.

[0061] Furthermore, the process of constructing the denoising autoencoder network and the kernel code estimation model based on the training set data is as follows: Figure 2 (a) is a process diagram of the noise reduction autoencoder kernel density estimation model of the present invention; (b) is a training and testing diagram of the noise reduction autoencoder kernel density estimation model of the present invention; Figure 3 This is a structural diagram of the noise reduction autoencoding network of the present invention;

[0062] The process of constructing a denoising autoencoder network and a kernel code estimation model based on the training set data is as follows: the training set data is sequentially input into the denoising autoencoder network to train the denoising autoencoder network and obtain the hidden layer features of the network flow data; the hidden layer features of the network flow data are input into the kernel density model to train the kernel density and construct the denoising autoencoder network and the kernel code estimation model. The specific process is as follows:

[0063] Step 1: Set the maximum number of training sessions;

[0064] Step 2: Initialize the parameters of the denoising autoencoder network;

[0065] Step 3: Input the training data into the denoising autoencoder network and perform the following calculations on the input data in sequence:

[0066]

[0067]

[0068] y=g(w'h+b') (3)

[0069] Among them: x, h, y represent input layer data, hidden layer data, and output data respectively; q Drepresents a random distribution between [0,1], a is the noise factor; f and g represent the nonlinear excitation functions in the encoding and decoding process respectively; w and w' are weight parameters, b and b' are biases;

[0070] Step 4: Use the reconstruction error calculation formula to obtain the reconstruction error;

[0071] J DAE (θ)=∑L(x,y) (4)

[0072] Among them: J DAE (θ) represents the overall error, and L represents the reconstruction error of each sample;

[0073] Step 5: Use Adagrad optimizer to minimize the reconstruction error;

[0074] Step 6: Determine whether the training count value k is greater than the set maximum number of training times. If so, the training of the denoising self-programmed network is completed. Otherwise, k = k + 1, adjust the weight and bias parameters, and return to step 3.

[0075] Step 7: Select a Gaussian function as the kernel function of the kernel density estimation model and set the value of its window width z; the formula of the Gaussian function is as follows:

[0076]

[0077] Among them: K z (x) represents a Gaussian function; x represents a random variable;

[0078] Step 8: Use the hidden layer feature data as input to build a kernel density estimation model through the kernel density estimation formula to obtain the probability density distribution of the training data

[0079] The kernel density estimation formula is as follows:

[0080]

[0081] Where: h is the hidden layer data, i is the input sample number, and n is the total number of input samples;

[0082] During the testing phase, the test data is used to obtain the threat probability through the constructed denoising autoencoder kernel density estimation model. The specific steps are:

[0083] Step 1: Test data Input into the trained denoising autoencoder network and obtain the hidden layer features of the test data

[0084] Step 2: Input the hidden layer features of the test data into the constructed kernel density estimation model to calculate the density value

[0085] Step 3: Calculate the outlier value of the test data by calculating the outlier value of the hidden layer of the test data. The smaller the outlier value, the more likely it is to be abnormal, and vice versa, the more consistent it is with the distribution of normal samples.

[0086] The calculation formula of the outlier value of the hidden layer is as follows:

[0087]

[0088] Outliers in the hidden layer of the test data;

[0089] Step 4: Negate and normalize to [0, 1], and use the outliers in the hidden layer of the negated and normalized test data as the probability of threat occurrence

[0090] Furthermore, the process of performing a security assessment on the network situation based on the probability of threat occurrence of the network flow data and determining the level of the network security situation is as follows:

[0091] Based on the probability of threat occurrence of network flow data, the confidentiality C, integrity I and availability A of network devices are scored and evaluated based on the common vulnerability scoring system to quantify the threat impact value;

[0092] The calculation formula of the threat impact value is as follows:

[0093]

[0094] Where: E i : threat impact value, i represents the network flow sequence number; C: confidentiality, I: integrity, A: availability.

[0095] The scores are shown in Table 1:

[0096] Table 1 CIA Assessment Table

[0097]

[0098] Quantify the network security situation value based on the threat occurrence probability and threat impact value of network flow data;

[0099] The network security situation value comprehensively considers the probability of threat occurrence P i and threat impact value E i , let the network security status value S of the i-th network flow be i for:

[0100] S i =P i E i (9)

[0101] The normalized cybersecurity situation value is used to classify cybersecurity status based on the National Emergency Response Plan for Public Emergencies and the situation classification standards of the National Internet Emergency Response Center. The normalized cybersecurity situation value is divided into five intervals: [0.00, 0.20], (0.20, 0.40], (0.40, 0.60], (0.60, 0.80], and (0.80, 1.00]), corresponding to the five levels of cybersecurity status: excellent, good, moderate, poor, and critical.

[0102] A network security situation assessment system based on denoising autoencoder kernel density estimation, comprising:

[0103] Acquisition module: used to obtain network flow data as situation data;

[0104] Preprocessing module: used to preprocess the situation data; divide the preprocessed situation data into training set data, verification set data and test set data according to the proportion;

[0105] Training module: used to build a denoising autoencoder network and a nuclear code estimation model based on training set data;

[0106] Estimation module: used to input the test set data into the denoising autoencoder network and the nuclear code estimation model in sequence to obtain the threat probability of the network flow data;

[0107] Determination module: used to conduct security assessment of network situation based on the probability of threat occurrence of network flow data and determine the level of network security situation.

[0108] The computer environment used in the experiment was an Intel(R) Core(TM) i7-4790 CPU @ 3.60GHz, 8.00GB of RAM, the simulation language was Python 3.6, and TensorFlow 1.10, and the compilation environment was PyCharm Community Edition 2020.2.2 x64. The network traffic data used was from the CICIDS-2017 dataset. The abnormal data from Tuesday to Friday accounted for 3.2%, 4.9%, 0.49%, and 69.7% of the normal data, respectively. The sample distribution is very unbalanced. This paper selected 15% of the Monday data as the training set, and the Tuesday to Friday data as the test set. The data from Tuesday to Thursday was selected at 10 times the normal data ratio, and the data from Friday was selected at the original ratio. The final experimental data volume is shown in Table 2.

[0109] Table 2 Experimental dataset

[0110]

[0111] To test the true effectiveness of the proposed network security situation assessment method based on denoising auto-encoding kernel density estimation, after a large number of experiments, the parameters of the denoising auto-encoding network are selected as follows: both the input neurons and output neurons of the denoising auto-encoding network are 78, and the hidden layer neurons are determined according to to be 9, where m is the number of input neurons, the error function is the mean square error function, the optimizer is Adagrad, the activation functions of the encoder and decoder are both Sigmoid functions, the learning rate is 0.01, the number of iterations is 100, batch is 300, and the noise factor is 0.4. The kernel function of the kernel density estimation model is the Gaussian function, and its parameter bandwidth is determined according to

[0112] Example 1: Verification of the effectiveness of the denoising auto-encoding network and the kernel password estimation model

[0113] The test set is used to verify the effectiveness of the denoising auto-encoding network and the kernel password estimation model, and the ROC curve and AUC value are selected as indicators to evaluate the model's performance. The ROC curve is formed by connecting different points under different threshold settings. The ROC curve of an ideal model should be as close as possible to the upper left end. The AUC value represents the size of the area under the ROC curve, and the larger the value, the better the performance of the model. Figure 4 This is the ROC-AUC graph of the OCDAE-KDE model on the test set of this invention; the dotted line in the figure indicates AUC = 0.5, which means that the classification effect of the model at this value is just a random guess and is the boundary for judging whether the model is effective; a model with 0 < AUC < 0.5 indicates a very poor effect, and the prediction effect is worse than a random guess and is meaningless in real life; a model within the range of 0.5 < AUC < 1 is effective, and the performance of the model gets better as the AUC value increases. When AUC = 1, the model is a perfect classification model. It can be seen from the figure that the AUC values of the model proposed in this paper are all above 0.9 in the four-day test, indicating that the denoising auto-encoding kernel density estimation model not only has effectiveness but also excellent performance.

[0114] Example 2: Verification of the classification performance of the denoising auto-encoding network and the kernel password estimation model

[0115] The denoising autoencoder kernel density estimation model was compared with the autoencoder model (OCAE), the denoising autoencoder model (OCDAE), and the autoencoder kernel density estimation model (OCAE-KDE), using the accuracy, precision, recall, and F1-score evaluation metrics. TP stands for true positive, meaning the model classifies a sample as positive even though it is a true positive. FP stands for false positive, meaning the model classifies a sample as positive even though it is a negative. TN stands for true negative, meaning the model classifies a sample as negative even though it is a true negative. FN stands for false negative, meaning the model classifies a sample as negative even though it is a true positive. The detailed formula is shown in Equation (3-6).

[0116]

[0117]

[0118]

[0119]

[0120] Before calculating these four metrics, it's necessary to determine the anomaly score thresholds for each of the four models. OCDAE and OCAE use the average reconstruction error of the training set plus three times its standard deviation as the anomaly threshold. OCDAE-KDE and OCAE-KDE first sort the anomaly scores in the training set in ascending order and then select the value at the 0.5% percentile as the anomaly threshold. Table 3 compares the evaluation metric results. It shows that across the four days, this method achieved the highest accuracy and F1 score among the four models, improving by 3.51% and 5.99%, respectively, compared to OCAE.

[0121] Table 3 Comparison of evaluation index results of four models

[0122]

[0123]

[0124] Example 3: Situation Assessment Results Analysis

[0125] OCDAE-KDE and OCAE, OCDAE, and OCAE-KDE were used to calculate the network security situation value of the network flow during the attack period from Tuesday to Friday and visualize it for quantitative evaluation. The network security situation values ​​of the 300 test groups on Friday are as follows: Figure 5 As shown. Figure 5As can be seen from the results, peaks indicate abnormal situations and network threats. The OCAE model has a low baseline value and a relatively flat trend, with smaller fluctuations during attack moments. The OCDAE model, however, performs slightly better due to its "corruption" of input data, which suppresses replication. The OCDAE-KDE and OCAE-KDE models, combined with a kernel density model for probabilistic estimation of hidden layer features, are more capable of characterizing network threats. However, the baseline values ​​of the proposed method are higher than those of the other three models, indicating its strong sensitivity to abnormal network situations.

[0126] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.

[0127] [1]Xie Lixia,Wang Yachao,Yu Jinbo.Network security situationawareness based on neural network[J].Journal of Tsinghua University(Scienceand Technology),2013,53(12):1750-1760.

[0128] Xie Lixia, Wang Yachao, Yu Jinbo. Network security situation awareness based on neural network[J]. Journal of Tsinghua University: Science and Technology, 2013, 53(12):1750-1760.

[0129] [2]Xie Lixia,Wang Zhihua.Network security situation assessment methodbased on cuckoo search optimized back propagation neural network[J].Journalof Computer Applications,2017,37(7):1926-1930.

[0130] Xie Lixia, Wang Zhihua. Network security situation assessment method based on cuckoo search optimization BP neural network[J]. Computer Applications, 2017, 37(7): 1926-1930.

[0131] [3]HanWeiHong,Tian ZhiHong,Huang Zizhong,et al.QuantitativeAssessmentof Wireless Connected Intelligent Robot Swarms Network Security Situation[J].IEEE Access,2019,7(99):134293-134300.

[0132] [4] Zhang Yuchen, Zhang Renchuan, Liu Jing, Wang Yongwei. Network Security Situation Evaluaton Using Deep Auto-Encoders Network [J]. Computer Engineering and Applications, 2020, 56 (06): 92-98.

[0133] Zhang Yuchen, Zhang Renchuan, Liu Jing, Wang Yongwei. Network security situation assessment using deep autoencoder network[J]. Computer Engineering and Applications, 2020, 56(06):92-98.

[0134] [5] Yang Hongyu, Zeng Renyan, Wang Fengyan, et al. An Unsupervised Learning-Based Network Threat Situation Assessment Model for Internet of Things [J]. Security and Communication Networks, 2020, 2020 (9): 1-11.

[0135] [6]Yang Hongyu,Zeng Renyun.A deep learning network security situation assessment method[J].Journal ofXidianUniversity,2021,48(01):183-190.

[0136] [7]Lakkaraju K,Yurcik W,Lee A J.NVisionIP:NetFlow visualizations ofsystem state for security situational awareness[C] / / Workshop onVisualization&DataMining for Computer Security.2004.

Claims

1. A network security situation assessment method based on denoising autoencoder kernel density estimation, characterized in that: The following steps are involved: Obtain network flow data as situation data; Preprocessing the situation data; The pre-processed situation data is divided into training set data and test set data according to the proportion; Construct a denoising autoencoder network and kernel density estimation model based on the training set data; The test set data is sequentially input into the denoising autoencoder network and the kernel density estimation model to obtain the threat occurrence probability of the network flow data; Conduct security assessments on network situations based on the probability of threat occurrence in network flow data to determine the level of network security. The process of constructing a denoising autoencoder network and a kernel density estimation model based on the training set data is as follows: the training set data is sequentially input into the denoising autoencoder network to train the denoising autoencoder network and obtain the hidden layer features of the network flow data; the hidden layer features of the network flow data are input into the kernel density model to train the kernel density and construct the denoising autoencoder network and the kernel density estimation model. The specific process is as follows: Step 1: Set the maximum number of training sessions; Step 2: Initialize the parameters of the denoising autoencoder network; Step 3: Input the training set data into the denoising autoencoder network and perform the following calculations on the input data in sequence: (1) (2) (3) 、 、 Represent input layer data, hidden layer data, and output data respectively; express Random distribution between is the noise factor; and Respectively represent the nonlinear excitation functions in the encoding and decoding process; and is the weight parameter, and is bias; Step 4: Calculate the reconstruction error using the reconstruction error calculation formula; Step 5: Minimize the reconstruction error; adjust the weight and bias parameters; Step 6: Determine whether the training count value k is greater than the set maximum number of training times. If so, the training of the denoising self-programmed network is completed. Otherwise, k=k+1, and the weight and bias parameters will be adjusted, and return to step 3. Step 7: Input the training data into the trained denoising autoencoder network again and obtain the hidden layer feature data; Step 8: Select the Gaussian function as the kernel function of the kernel density estimation model and set its window width; Step 9: Use the hidden layer feature data as input to build a kernel density estimation model through the kernel density estimation formula to obtain the probability density distribution of the training data; The parameters of the denoising autoencoder network are as follows: the number of input neurons and output neurons of the denoising autoencoder network is 78, and the number of hidden layer neurons is 78. The number of neurons is determined to be 9, where m is the number of input neurons, the error function is the mean square error function, the optimizer is Adagrad, the activation function of the encoder and decoder are both Sigmoid functions, the learning rate is 0.01, the number of iterations is 100, the batch is 300, the noise factor is 0.4, and the kernel function of the kernel density estimation model is the Gaussian function. Its parameter window width is based on It is 2.1213.

2. The network security situation assessment method based on denoising autoencoder kernel density estimation according to claim 1 is characterized by: The pre-processing of the situation data comprises the following steps: removing repeated feature columns in the situation data and samples with infinite and null values ​​in the fields, and then normalizing the data.

3. The network security situation assessment method based on denoising autoencoder kernel density estimation according to claim 1 is characterized by: The process of inputting the test set data into the denoising autoencoder network and the kernel density estimation model in sequence to obtain the threat occurrence probability of the network flow data is as follows: Step 1: Test data Input into the trained denoising autoencoder network and obtain the hidden layer features of the test data ; Step 2: Input the hidden layer features of the test data into the constructed kernel density estimation model to calculate the density value ; Step 3: Pass Find outliers in the test data , the smaller the outlier value, the more abnormal it is; Step 4: Invert the outlier value and normalize it to [0, 1]. The inverted and normalized value is used as the probability of threat occurrence.

4. The network security situation assessment method based on denoising autoencoder kernel density estimation according to claim 1 is characterized in that: The process of performing a security assessment on the network situation based on the probability of threat occurrence of network flow data and determining the level of network security situation is as follows: Based on the probability of threat occurrence of network flow data, the confidentiality C, integrity I and availability A of network devices are scored and evaluated based on the common vulnerability scoring system to quantify the threat impact value; Quantify the network security situation value based on the threat occurrence probability and threat impact value of network flow data; Normalized network security situation value is divided into five levels: excellent, good, medium, poor, and critical.

5. The network security situation assessment method based on denoising autoencoder kernel density estimation according to claim 4 is characterized in that: The calculation formula of the threat impact value is as follows: (8) in: : Threat impact value, Indicates the network flow sequence number; C: confidentiality, I: integrity, A: availability.

6. A network security situation assessment system based on denoising autoencoder kernel density estimation, characterized in that: include: Acquisition module: used to obtain network flow data as situation data; Preprocessing module: used for preprocessing the situation data; The pre-processed situation data is divided into training set data and test set data according to the proportion; Training module: used to build a denoising autoencoder network and kernel density estimation model based on training set data; Estimation module: used to sequentially input the test set data into the denoising autoencoder network and the kernel density estimation model to obtain the threat probability of the network flow data; Determination module: used to conduct security assessment of network situation based on the probability of threat occurrence of network flow data and determine the level of network security situation; The process of constructing a denoising autoencoder network and a kernel density estimation model based on the training set data is as follows: the training set data is sequentially input into the denoising autoencoder network to train the denoising autoencoder network and obtain the hidden layer features of the network flow data; the hidden layer features of the network flow data are input into the kernel density model to train the kernel density and construct the denoising autoencoder network and the kernel density estimation model. The specific process is as follows: Step 1: Set the maximum number of training sessions; Step 2: Initialize the parameters of the denoising autoencoder network; Step 3: Input the training set data into the denoising autoencoder network and perform the following calculations on the input data in sequence: (1) (2) (3) 、 、 Represent input layer data, hidden layer data, and output data respectively; express Random distribution between is the noise factor; and Respectively represent the nonlinear excitation functions in the encoding and decoding process; and is the weight parameter, and is bias; Step 4: Calculate the reconstruction error using the reconstruction error calculation formula; Step 5: Minimize the reconstruction error; adjust the weight and bias parameters; Step 6: Determine whether the training count value k is greater than the set maximum number of training times. If so, the training of the denoising self-programmed network is completed. Otherwise, k=k+1, and the weight and bias parameters will be adjusted, and return to step 3. Step 7: Input the training data into the trained denoising autoencoder network again and obtain the hidden layer feature data; Step 8: Select the Gaussian function as the kernel function of the kernel density estimation model and set its window width; Step 9: Use the hidden layer feature data as input to build a kernel density estimation model through the kernel density estimation formula to obtain the probability density distribution of the training data; The parameters of the denoising autoencoder network are as follows: the number of input neurons and output neurons of the denoising autoencoder network is 78, and the number of hidden layer neurons is 78. The number of neurons is determined to be 9, where m is the number of input neurons, the error function is the mean square error function, the optimizer is Adagrad, the activation function of the encoder and decoder are both Sigmoid functions, the learning rate is 0.01, the number of iterations is 100, the batch is 300, the noise factor is 0.4, and the kernel function of the kernel density estimation model is the Gaussian function. Its parameter window width is based on It is 2.1213.

Citation Information

Patent Citations

  • Time sequence-based process data anomaly detection method

    CN110865625A

  • Power system operation risk assessment method based on weather zoning strategy

    CN112884601A