A Network Security Situation Assessment Method Based on Data Mining

Through a fully connected neural network based on negative curvature gradient descent and an extreme learning machine with quantum potential energy constraints, combined with the generative adversarial network of adaptive loss, data scarcity and feature extraction difficulties in the field of network security are solved, and a more efficient network security situation evaluation is achieved.

CN120074967BActive Publication Date: 2025-07-18CHENGDU ANZHUN NETWORK SECURITY TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510559535.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-07-18
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

The field of network security faces the problems of data scarcity and sample imbalance. Existing methods cannot effectively simulate complex attack scenarios. Traditional feature extraction methods have gradient vanishing and overfitting problems, and insufficient classification model accuracy and generalization capabilities.

Method used

A fully connected neural network based on negative curvature gradient descent is used for feature extraction, a classifier model is built in combination with an extreme learning organization based on quantum potential energy constraints, and a generative adversarial network with adaptive loss is used for data augmentation, and the learning rate and weight are dynamically adjusted to improve model training stability and classification accuracy.

Benefits of technology

It effectively solves the problems of data scarcity and sample imbalance, improves the generalization ability and classification accuracy of the model, can more accurately simulate complex network attacks, and significantly improves the accuracy and stability of network security situation evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120074967B_ABST
    Figure CN120074967B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of artificial intelligence technology, and discloses a network security situation assessment method based on data mining, including obtaining network security data; extracting features from the network security data by using a fully connected neural network based on negative curvature gradient descent to obtain feature data; constructing a classifier model by using an extreme learning machine based on quantum potential energy constraint, and training the classifier model by using the feature data; and performing network security situation assessment by using the trained classifier model. When dealing with variable network attack types, the present invention can provide more stable and accurate classification results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly relates to a network security situation assessment method based on data mining. Background Art

[0002] With the rapid development of information technology, the forms and means of cyberattacks have become increasingly complex and diverse, and traditional network security protection means and emergency response mechanisms are gradually difficult to cope with the increasingly complex network threats. Against this background, how to effectively evaluate and predict the network security situation has become an important issue in network security research and practice.

[0003] The Chinese invention patent with the publication number CN118114592B proposes an alternative modeling method for predicting saltwater intrusion dynamics under uncertain conditions, including the following steps: establishing a saltwater intrusion model based on physical processes and determining the value range of input parameters; obtaining input samples, importing them into the saltwater intrusion simulation model to obtain an output data set, and constructing an input-output data set; training three machine learning alternative models according to the input-output data set, and importing the input samples into the alternative models to obtain prediction data; combining the chloride concentration observation data with the prediction data of the three alternative models, and using the Bayesian model averaging algorithm to obtain the weights and variances of the models, and constructing an integrated machine learning alternative model of the numerical model. The present invention organically combines the Bayesian averaging algorithm and machine learning, quantifies the model uncertainty, constructs an integrated machine learning alternative model to improve the prediction performance, and proves the feasibility of integrated machine learning alternative modeling under uncertain conditions in predicting the migration dynamics of groundwater pollutants.

[0004] The Chinese invention patent with the publication number CN118551647A proposes a railway foreign object intrusion simulation method based on artificial intelligence generation and digital twin, belonging to the field of artificial intelligence, especially involving generative artificial intelligence technology and digital twin technology; the specific implementation process is as follows: establishing a railway foreign object intrusion scenario based on deep reinforcement learning technology, designing a Monte Carlo method according to the Markov hypothesis of the intruding foreign object, and constructing a railway foreign object intrusion generation model; based on the foreign object intrusion generation model, using the rainfall weather mathematical model to design a model for adding and removing the railway rainy day effect, and then combining the atmospheric scattering model with the depth of field of the railway scene to construct a railway weather environment model during foreign object intrusion generation, and constructing a railway foreign object intrusion simulation platform based on digital twin. The present invention improves the difficulties and pain points in the current actual work, such as the difficulty in collecting and experimenting with foreign object intrusion data and the limited data sources, and realizes the large-scale generation of available data.

[0005] The Chinese invention patent with the publication number CN118101274B proposes a method, device, equipment and medium for constructing a network intrusion detection model. The method for constructing the network intrusion detection model includes: obtaining a plurality of original data to form an original data set; wherein, the original data includes: numerical network traffic data, data type labels and annotation results; clustering and counting the original data with the same data type label to respectively obtain the quantity of the original data with the same data type label as the clustering and statistical results of each data type label; performing data augmentation operation on the original data set according to the clustering and statistical results of each data type label to obtain a target data set; training a preset model according to the target data set to obtain a network intrusion detection model. Through the technical solution of the present invention, the construction of the network intrusion detection model and the detection of network intrusion can be realized, and the accuracy of network intrusion detection work is improved.

[0006] The following problems still need to be further solved in the prior art:

[0007] 1. The field of network security often faces the problems of data scarcity and sample imbalance. Existing expansion methods mostly rely on simple sampling techniques or rule generation, and cannot effectively simulate complex attack scenarios, resulting in poor generalization ability of the model. Traditional generative adversarial networks often have unstable training when facing skewed data distributions, and the quality of the generated samples is difficult to guarantee.

[0008] 2. Traditional feature extraction methods usually adopt standard gradient descent algorithms, which may encounter problems such as gradient disappearance, gradient explosion or getting stuck in local optimal solutions, resulting in low training efficiency of neural networks and difficulty in effectively processing high-dimensional complex data.

[0009] 3. Traditional network security classification models have poor accuracy and generalization ability when facing complex network attack data. Although the extreme learning machine has a fast training speed, it lacks effective modeling of complex data relationships, which easily leads to insufficient classification accuracy and overfitting. Summary of the Invention

[0010] In view of the above deficiencies in the prior art, the present invention provides a network security situation assessment method based on data mining.

[0011] In order to achieve the above invention purpose, the technical solution adopted by the present invention is:

[0012] A network security situation assessment method based on data mining, comprising the following steps:

[0013] Obtain network security data;

[0014] Use a fully connected neural network based on negative curvature gradient descent to extract features from the network security data to obtain feature data;

[0015] A classifier model is constructed using an extreme learning machine based on quantum potential energy constraints, and the classifier model is trained using feature data;

[0016] The trained classifier model is used for network security situation assessment.

[0017] Furthermore, a fully connected neural network based on negative curvature gradient descent is used to extract features from network security data to obtain feature data, including:

[0018] Initialize the parameters of the neural network;

[0019] At each iteration, adjust the direction of gradient update to the negative curvature direction according to the standard gradient, and calculate the update amount of the corresponding neural network weights according to the negative curvature gradient;

[0020] Update the neural network parameters according to the calculated gradient information, and dynamically adjust the learning rate of the neural network;

[0021] Repeat the above steps until the preset iteration stop condition is met.

[0022] Furthermore, calculate the update amount of the corresponding neural network weights according to the negative curvature gradient, specifically:

[0023] ;

[0024] Where, represents the update amount of the neural network weights, represents the learning rate of the neural network, represents the coefficient for adjusting the influence of negative curvature, represents the loss function of the neural network, and represent the first derivative and the second derivative of the loss function respectively, represents the weights of the neural network; sign( ) is the sign function;

[0025] Perform path integral gradient estimation on all connection paths between neurons in each layer during forward propagation through Feynman path integral, specifically:

[0026] ;

[0027] Where, represents the network security data sample that changes along the potential path from the current state ; represents the current state of the network security data during the neural network processing; represents the differential of the path; represents the integral; It is the loss function of the neural network on the current path.

[0028] Furthermore, the learning rate of the neural network is dynamically adjusted, specifically as follows:

[0029] ;

[0030] where represents the learning rate of the neural network at the th iteration; represents the learning rate of the neural network at the th iteration; represents the learning rate decay coefficient, represents the current iteration number.

[0031] Furthermore, a classifier model is constructed using an extreme learning machine based on quantum potential constraint, and the classifier model is trained using feature data, including:

[0032] Initializing the weight matrix of the extreme learning machine based on quantum potential constraint;

[0033] Propagating the feature data through the weight matrix with quantum potential constraint to the hidden layer and calculating the hidden layer output;

[0034] Optimizing the weights from the hidden layer to the output layer using an adaptive orthogonal regularization strategy;

[0035] Calculating the quantum state correlation of the hidden layer output values and dynamically adjusting the output layer weights of the extreme learning machine;

[0036] Ending the model training after the forward propagation of all training data is completed.

[0037] Furthermore, initializing the weight matrix of the extreme learning machine based on quantum potential constraint, specifically as follows:

[0038] ;

[0039] where represents the number of neurons in the hidden layer of the extreme learning machine; represents the input data of the extreme learning machine; represents the imaginary unit; represents a random angle; represents the width parameter of the quantum potential; is the L2 norm;

[0040] The random angle is a dynamic distribution adjusted based on the statistical characteristics of network security data features, and the calculation method is expressed as:

[0041] ;

[0042] Among them, represents the error function, and respectively represent the mean and variance of the input data of the extreme learning machine.

[0043] Furthermore, the feature data is propagated to the hidden layer through the weight matrix with quantum potential constraints, and the output of the hidden layer is calculated. Specifically:

[0044] ;

[0045] Among them, represents the activation function of the extreme learning machine; represents the output of the hidden layer of the extreme learning machine; represents the bias term of the extreme learning machine;

[0046] The calculation method of the activation function of the extreme learning machine is expressed as:

[0047] ;

[0048] Among them, represents the input of the activation function of the extreme learning machine; represents the square of the modulus of the input of the activation function of the extreme learning machine; represents the parameter for controlling the activation intensity.

[0049] Furthermore, an adaptive orthogonal regularization strategy is used to optimize the weights from the hidden layer to the output layer. Specifically:

[0050] ;

[0051] Among them, represents the output layer weights of the extreme learning machine; represents the conjugate transpose of the output of the hidden layer of the extreme learning machine; represents the target output of the extreme learning machine; represents the regularization parameter of the extreme learning machine; represents the identity matrix;

[0052] The calculation method of the regularization parameter of the extreme learning machine is expressed as:

[0053] ;

[0054] Among them, represents the number of samples input to the extreme learning machine in the current batch; represents the regularization parameter baseline of the extreme learning machine; represents the -th output of the hidden layer of the extreme learning machine; Represents the average value of the outputs of all hidden layers of the extreme learning machine; Represents a set constant; Represents the coefficient for controlling the adjustment sensitivity; Represents the average correlation of the quantum states of the entire batch of data.

[0055] The calculation method of the average correlation of the quantum states of the entire batch of data is expressed as:

[0056] ;

[0057] Among them, Represents the quantum state after the th feature data input to the extreme learning machine is transformed by the hidden layer; Represents the quantum state after the th feature data input to the extreme learning machine is transformed by the hidden layer; Represents and the inner product between two quantum states.

[0058] Furthermore, after obtaining the network security data, a generative adversarial network based on adaptive loss is used to perform data augmentation on the network security data, including:

[0059] Initializing the parameters of the generator and the discriminator;

[0060] In the generation stage, the generator generates a batch of new network security data samples according to the current network parameters;

[0061] Using the discriminator to evaluate the generated network security data samples and the real network security data samples, and outputting the probability that each sample is a real sample;

[0062] In the adversarial training stage, based on the inherent complexity of the network security data and the stability of the generator during training, an adaptive loss function is calculated;

[0063] Using the dynamic game strategy to adjust the learning rates and update strategies of the generator and the discriminator, and dynamically adjusting the loss function according to the results of each round of confrontation;

[0064] Repeating the iteration until the preset stop iteration condition is satisfied.

[0065] The present invention has the following beneficial effects:

[0066] 1. The present invention solves the problems of insufficient data and sample skew through the generative adversarial network algorithm based on adaptive loss. The generated network security data not only increases in quantity but also improves in quality, can better simulate the complex patterns of actual attacks, and helps to improve the generalization ability and recognition accuracy of subsequent models.

[0067] 2. The present invention uses a negative curvature gradient descent optimization algorithm to significantly improve the training efficiency of neural networks in a high-dimensional data environment, overcome problems such as local optimal solutions and gradient vanishing in traditional gradient descent methods. Especially when dealing with complex network security data, it can more accurately extract effective features, thereby improving the accuracy of subsequent classification algorithms.

[0068] 3. The present invention uses an extreme learning machine classification algorithm based on quantum potential energy constraints to better capture complex non-linear relationships in network security data, avoid overfitting, and significantly improve the classification accuracy and the generalization ability of the model. Especially when dealing with variable network attack types, this method can provide more stable and accurate classification results. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] Figure 1 It is a schematic flow diagram of a network security situation assessment method based on data mining;

[0070] Figure 2 It is a schematic diagram of the architecture of a network security situation assessment platform based on data mining;

[0071] Figure 3 It is a schematic diagram for comparing the training stability of different GAN methods;

[0072] Figure 4 It is a schematic diagram for comparing the optimization trajectories of a neural network algorithm based on negative curvature gradient descent and the traditional gradient descent method in the parameter space;

[0073] Figure 5 It is a schematic diagram for comparing the distribution of path integral gradient contributions. DETAILED DESCRIPTION OF THE INVENTION

[0074] The following describes the specific embodiments of the present invention to facilitate those skilled in the art of the present technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions created using the concept of the present invention are within the scope of protection.

[0075] As Figure 1 shown, an embodiment of the present invention provides a network security situation assessment method based on data mining, including the following steps S1 to S4:

[0076] S1. Obtain network security data;

[0077] In an alternative embodiment of the present invention, as Figure 2As shown in the figure, in the network data collection layer of the platform, the collection of network security data comes from a variety of network security monitoring systems, such as intrusion detection systems, user operation logs, firewall logs, and network traffic analysis tools, etc.

[0078] The collection methods of network security data include real-time monitoring and historical record analysis. Through API interfaces and network packet capture tools, network security data sets are automatically collected and updated.

[0079] The storage format of network security data adopts the structured JSON format. In one embodiment, the attributes of network security data include:

[0080] Ra is the attack type code, such as DDoS, malware, etc.; Da is the importance rating of the target asset; Ca is the timestamp when the attack occurred; Pa is the network protocol type, such as TCP, UDP; Qa is the source IP address; Fa is the target IP address; Va is the source port number; Wa is the target port number; Na is the number of transmitted data packets; Za is the flag indicating whether the attack is successful.

[0081] It should be noted that this embodiment is only to illustrate a network security data format and type of the present invention. In actual applications, the attributes of network security data are usually more than 10, and the number of attributes of network security data may reach dozens or even hundreds.

[0082] Furthermore, the collected network security data is labeled. The labeling method of the present invention is manual labeling. In one embodiment, the labeling categories include:

[0083] Minor: Only an attempt at a threat, without causing an actual impact;

[0084] Medium: Has invaded part of the system but has not affected the core functions;

[0085] Severe: Damages the core functions of the system or causes a long-term service interruption;

[0086] Critical: Causes large-scale data leakage or long-term paralysis of the entire network.

[0087] In this embodiment, 5 examples of network security data are shown in Table 1.

[0088] Table 1

[0089]

[0090] It can be understood that in the task of the present invention, the collection, acquisition, labeling, and preprocessing of network security training data are time-consuming and laborious, and insufficient training samples are likely to lead to poor generalization ability of the model and affect the accuracy of the model at the same time.

[0091] In the sample enhancement layer of the platform, network security data augmentation is achieved through a generative model. The present invention adopts a generative adversarial network algorithm based on adaptive loss for sample generation, thereby realizing network security data augmentation. Traditional generative adversarial networks often have difficulty generating relatively complex samples when facing skewed network security data distributions. Based on the inherent complexity of network security data and the stability of the generator during training, an adaptive loss function is calculated to promote the fast and stable learning of the generative model, enabling the network security data generated each time to better simulate complex data distributions.

[0092] Specifically, the training process of the generative adversarial network algorithm based on adaptive loss is as follows:

[0093] S101. Initialize the parameters of the generator and the discriminator. Let the generator be , and the discriminator be . Set the weight of the generator as , and the weight of the discriminator as. In a certain situation, the initialization method is expressed as:

[0094]

[0095]

[0096] Among them, is a normal distribution with a mean of 0 and a standard deviation of the identity matrix; obeys a specific distribution; is the identity matrix; is a normal distribution.

[0097] S102. In the generation stage, the generator generates a batch of new network security data samples according to the current network parameters. These samples initially attempt to mimic the distribution characteristics of real network security data, expressed as:

[0098]

[0099] Among them, is a latent vector sampled from the prior distribution; is the generator function, is the th iteration of the generator weight; is at the th iteration, the sample generated by the generator; represents the weight of the feature transformation; represents element-wise multiplication.

[0100] In one embodiment, the quality of the generated samples is improved through an adaptive feature transformation strategy, and the learning ability of the model for different network security data is enhanced. Based on the dynamic weighted sum and adjustment of the input features of the generator, feature weight learning is achieved. The calculation method of the weights of the feature transformation is expressed as:

[0101]

[0102] Wherein, is the Sigmoid activation function; is the parameter matrix of weight learning, specifically the weight parameter matrix of the generator.

[0103] S103. The discriminator evaluates the generated network security data samples and the real network security data samples, and outputs the probability that each sample is a real sample, which is expressed as:

[0104]

[0105] Wherein, means that the generated samples and the real samples are combined into a batch for input; is the discriminator function, is the discriminator weight at the th iteration; is the real sample; is the discrimination result of the discriminator in the th iteration.

[0106] S104. In the adversarial training stage, based on the inherent complexity of the network security data and the stability of the generator during training, an adaptive loss function is calculated to promote the fast and stable learning of the generation model. The calculation method is expressed as:

[0107]

[0108] Wherein, is the reconstruction loss function, specifically the reconstruction loss function for calculating the generated network security data and the real network security data, which is obtained by calculating the L2 norm; is the complexity-dependent loss function; the calculation method of the adjustment coefficient at the th iteration is expressed as ;

[0109] Furthermore, the calculation method of the complexity-dependent loss function is expressed as:

[0110]

[0111] Wherein, is the complexity scoring function; represents the expectation; To conform to a specific distribution; For the distribution of real network security data; For the noise distribution.

[0112] Furthermore, the complexity scoring function evaluates the complexity of the input network security data, specifically evaluating the complexity or rarity of the abnormal patterns included in the generated network security data. Let the input of the complexity scoring function be , then the calculation method of the complexity scoring function is expressed as:

[0113]

[0114] Among them, Is the variance calculation function, used to measure the irregularity and prediction difficulty of network security data; Represents the normalization parameter, preset by humans, such as set to 0.1, Is the input of the complexity scoring function, corresponding to Or In the calculation formula of the complexity-dependent loss function.

[0115] Furthermore, the adjustment coefficient is adjusted according to the number of iterations, used to adjust the weight between the standard adversarial loss and the loss based on the current training complexity. The calculation method is expressed as:

[0116]

[0117] Among them, Is the parameter that controls the steepness of the curve, Is the preset iteration cycle. When the training reaches a certain cycle, the complexity loss starts to increase, ensuring that as the training progresses, the model gradually increases its attention to complex network security data; Is the current number of iterations. Preferably, Is set to 2.

[0118] S105. Use the dynamic game strategy to adjust the learning rates and update strategies of the generator and discriminator, and dynamically adjust the loss function according to the results of each round of confrontation, giving priority to enhancing the ability to generate scarce category network security data, expressed as:

[0119]

[0120]

[0121] Among them, Is the weight of the generator at the -th iteration; Is the weight of the discriminator at the -th iteration; is the learning rate of the generator; is the learning rate of the discriminator; is the loss function of the generator; is the loss function of the discriminator; represents the gradient with respect to the weights of the generator; represents the weights with respect to the parameters of the discriminator.

[0122] Furthermore, the loss function of the generator combines the entropy of the generated network security data for comprehensive calculation to more comprehensively evaluate the quality of the generated network security data and increase the diversity of the generated samples, expressed as:

[0123]

[0124] where, represents the entropy function, is the adjustment coefficient that controls the influence of the entropy term; is the adaptive loss function.

[0125] Furthermore, the loss function of the discriminator uses a regularization term to prevent overfitting, expressed as:

[0126]

[0127] where, represents the true sample label (1 for true sample, 0 for generated sample); is the regularization coefficient for weight decay to reduce the model complexity and the risk of overfitting; represents the square of the L2 norm of the discriminator weights.

[0128] During the adversarial training process of the generator and the discriminator, the dynamic game strategy dynamically adjusts the adjustment coefficient that controls the influence of the entropy term and the regularization coefficient so that the generated network security data better simulates the real data distribution each time, expressed as:

[0129]

[0130]

[0131] where, is the learning rate adjustment factor for controlling the dynamic adjustment of the learning rate; is the partial derivative symbol; is the th iteration's adjustment coefficient that controls the influence of the entropy term, is the th iteration's regularization coefficient; is the th iteration's adjustment coefficient that controls the influence of the entropy term, is the The regularization coefficient of the next iteration. Preferably, it is set to 0.01.

[0132] S106. Repeat the above steps iteratively until the preset stop iteration condition is met, which indicates that the model training is completed. In one embodiment, the preset stop iteration condition is to reach the preset maximum number of iterations. Preferably, the preset maximum number of iterations is set to 1000 times.

[0133] After the network security data augmentation model training is completed, use the trained network security data augmentation model to increase the number of samples. In one embodiment, assume that the original collected samples are 800, and the network security data augmentation model generates 200 samples through augmentation. Then the augmented network security data set contains 1000 samples.

[0134] As Figure 3 shown, by comparing the training process stability of different generative adversarial network methods, the effectiveness of the dynamic game strategy is verified. The experiment compares the adaptive loss algorithm proposed in the present invention with the classical generative adversarial network (original GAN), the least squares generative adversarial network (LSGAN), and the generative adversarial network with gradient penalty (WGAN), and observes the trend of the generator loss value changing with the number of iterations. The experimental results show that the fluctuation amplitude of the loss curve of the method of the present invention is significantly smaller than that of other methods. In particular, there is no severe oscillation or divergence phenomenon in the middle and late stages of training, indicating that the dynamically adjusted weight coefficient in the adaptive loss function can automatically balance the contributions of the reconstruction loss and the complexity loss according to the model training state, avoiding the problem of single optimization direction caused by fixed loss weights in traditional methods, thus effectively suppressing the mode collapse phenomenon and enabling the model to converge continuously and stably.

[0135] S2. Use a fully connected neural network based on negative curvature gradient descent to extract features from the network security data to obtain feature data;

[0136] In an optional embodiment of the present invention, at the feature extraction layer of the platform, a neural network algorithm based on negative curvature gradient descent is used as the feature extraction model to extract features from the augmented network security data.

[0137] The present invention uses a 5-layer fully connected neural network for feature extraction. In the prior art, some solutions use neural networks for feature extraction. In some neural network structures, problems such as gradient disappearance, gradient explosion, or getting stuck in local optimal solutions may be encountered, affecting the training stability and model performance.

[0138] In the present invention, negative curvature gradient descent accelerates the convergence of gradients during training by leveraging the negative curvature direction. Especially in complex, high-dimensional data environments, it can significantly improve the efficiency and accuracy of the model when processing non-linear network security data.

[0139] Specifically, the training process of the neural network algorithm based on negative curvature gradient descent is as follows:

[0140] S201. Initialize the parameters of the neural network. In one embodiment, the initialization method is expressed as:

[0141]

[0142]

[0143] where represents the initial weight matrix of the neural network, represents the initial bias vector of the neural network, and represent the dimensions of the input layer and output layer of the neural network respectively; is the zero vector function; is the function to take random values; is the weight of the neural network; is the bias of the neural network.

[0144] S202. In some cases, the traditional gradient descent direction may not be optimal. In the present invention, at each iteration, first calculate the standard gradient, and then determine whether the direction of gradient update needs to be adjusted to the negative curvature direction. The calculation method of the adjustment of the negative curvature gradient corresponding to the update amount of the neural network weight is expressed as:

[0145]

[0146] where is the update amount of the neural network weight, is the learning rate of the neural network, is the coefficient to adjust the influence of negative curvature, is the loss function of the neural network, and sign( ) is the sign function; and represent the first-order and second-order derivatives of the loss function respectively. By using the sign function of the second-order derivative to adjust the gradient descent direction, as a means of adjusting the negative curvature gradient, it is convenient to cross the saddle point during parameter update.

[0147] S203. For each training batch, not only calculate the gradient of the current batch, but also consider the gradient contributions of all possible paths from the current state to possible future states through Feynman path integration. Moreover, integrating over potential state transitions takes into account the contributions of all possible paths of network security data changes to the current gradient, in order to more comprehensively estimate the gradient and the impact of these transitions on the current parameter update. The path integral gradient estimation method is expressed as:

[0148]

[0149] In the formula, represents the network security data sample that changes along the potential path from the current state ; is the current state of the network security data during the neural network processing; represents the path of the network security data features during forward propagation, corresponding to the connection paths between neurons in each layer; represents the differentiation with respect to the path; represents the integration; is the loss function of the neural network on the current path.

[0150] S204. Update the neural network parameters according to the calculated gradient information. Moreover, the learning rate of the neural network is dynamically adjusted to address the problems of vanishing gradients or exploding gradients during the training process, and the learning rate is slowly decreased as the training progresses to improve the training stability and promote convergence. The adjustment method is expressed as:

[0151]

[0152] In the formula, is the learning rate of the neural network at the -th iteration; is the learning rate of the neural network at the -th iteration; is the learning rate decay coefficient, is the current iteration number. Preferably, is set to 0.95.

[0153] Furthermore, the update method of the weight parameters of the neural network is expressed as:

[0154]

[0155] In the formula, represents the weight of the neural network at the -th iteration, represents the weight of the neural network at the -th iteration.

[0156] S205. Repeat the above steps iteratively until the preset iteration stop condition is met, which indicates that the model training is completed. In one embodiment, the preset iteration stop condition is to reach the preset maximum number of iterations. Preferably, the preset maximum number of iterations is set to 1000 times.

[0157] S3. Construct a classifier model using an extreme learning machine based on quantum potential energy constraint, and use the feature data to train the classifier model;

[0158] In an alternative embodiment of the present invention, at the classification layer of the platform, an extreme learning machine classification algorithm based on quantum potential energy constraint is used as the classifier model.

[0159] Although the traditional extreme learning machine algorithm has advantages in training speed, when dealing with network security data with complex relationships, it is prone to problems such as insufficient classification accuracy and insufficient generalization ability.

[0160] The quantum potential energy constraint strategy is inspired by the potential energy model in quantum mechanics. It uses the wave function characteristics of quantum states to constrain the characteristics of network security training data of the extreme learning machine, enhances the model's perception ability of the internal relationship of network security data, prevents overfitting, and improves the classification accuracy and generalization ability.

[0161] Specifically, the training process of the extreme learning machine classification algorithm based on quantum potential energy constraint is as follows:

[0162] S301. Considering the influence of quantum potential energy to ensure that the preliminary non-linear transformation of network security data can reflect its quantum characteristics, initialize the weight matrix of the extreme learning machine, expressed as:

[0163]

[0164] In the formula, represents the number of neurons in the hidden layer of the extreme learning machine; represents the input data of the extreme learning machine, and the input data of the extreme learning machine is the feature vector obtained by feature extraction; represents the imaginary unit; is a random angle; represents the width parameter of the quantum potential energy, which controls the spread degree of the quantum state; is the L2 norm.

[0165] Furthermore, considering quantum randomness, the random angle is a dynamic distribution adjusted based on the statistical characteristics of network security data features, and the calculation method is expressed as:

[0166]

[0167] In the formula, is the error function, and are the mean and variance of the input data of the extreme learning machine respectively, which are used to adjust the distribution of random angles to adapt to different characteristics of network security data.

[0168] S302. Calculate the output of the hidden layer of the extreme learning machine. The network security data after feature extraction is propagated to the hidden layer through the weight matrix with quantum potential constraint, and its output is calculated, expressed as:

[0169]

[0170] In the formula, is the activation function of the extreme learning machine; is the output of the hidden layer of the extreme learning machine; is the bias term of the extreme learning machine, and the same quantum initialization is adopted as the weight of the extreme learning machine.

[0171] Furthermore, the activation function of the extreme learning machine is a quantized activation function, which uses the quantum interference effect to simulate the non-linear mapping change of the network security data after feature extraction, and can effectively increase the non-linear expression ability of the model for features. The calculation method is expressed as:

[0172]

[0173] In the formula, is the input of the activation function of the extreme learning machine; represents the square of the modulus of the input of the activation function of the extreme learning machine; is the parameter controlling the activation intensity.

[0174] S303. Optimize the weight from the hidden layer to the output layer by using the adaptive orthogonal regularization strategy. Through the regularization process, the weight matrix is adaptively adjusted to optimize the output performance of the algorithm. The regularization method of the output layer weight of the extreme learning machine is expressed as:

[0175]

[0176] In the formula, is the output layer weight of the extreme learning machine; represents the conjugate transpose of the output of the hidden layer of the extreme learning machine; is the target output of the extreme learning machine, representing the true label vector of the sample; is the regularization parameter of the extreme learning machine, which is adaptively adjusted through the change of the quantum state; represents the identity matrix.

[0177] Furthermore, the regularization parameter of the extreme learning machine can be dynamically adjusted according to the data of the current batch, better adapting to the characteristics of network security data after feature extraction and preventing overfitting. The calculation method is expressed as:

[0178]

[0179] In the formula, is the number of samples input to the extreme learning machine in the current batch; is the baseline of the regularization parameter of the extreme learning machine; is the output of the th hidden layer of the extreme learning machine; is the average value of all hidden layer outputs of the extreme learning machine; is a small constant to avoid the denominator being zero; is the coefficient controlling the adjustment sensitivity; represents the average correlation of the quantum states of the entire batch of data. The higher its value, the closer the internal connection between the data. Preferably, is set to 0.001, is set to 0.01.

[0180] S304. Adopt a dynamic adjustment mechanism based on the quantum state correlation to improve the adaptability and accuracy of the algorithm in dealing with complex network security environments. By analyzing the quantum state correlation of the hidden layer output, dynamically adjust the output layer weights of the extreme learning machine, enabling the algorithm to adaptively adjust the learning strategy based on the internal connection and changes of the network security data after feature extraction. The calculation method of the average correlation of the quantum states of the entire batch of data is expressed as:

[0181]

[0182] In the formula, represents the quantum state after the th data input to the extreme learning machine is transformed by the hidden layer; represents the quantum state after the th data input to the extreme learning machine is transformed by the hidden layer; represents and the inner product between two quantum states, used to measure their similarity.

[0183] S305. After the forward propagation of all training data is completed, the model training stage ends. At this time, the model can accurately classify effectively according to the input data features. The final output of the model is expressed as:

[0184]

[0185] In the formula, It is the model output of the extreme learning machine, corresponding to the prediction probability vector of the extreme learning machine for the samples. The category corresponding to the element with the largest value in this vector is taken as the prediction category of the extreme learning machine for the samples.

[0186] S5. Use the trained classifier model to conduct network security situation assessment.

[0187] In an alternative embodiment of the present invention, at the network data application layer of the platform, the trained model is used to process new samples and evaluate the network security situation. In one embodiment, the original network security data collected by the network data collection layer is input into the feature extraction layer and undergoes feature processing using the trained feature extraction model. Further, the processed features are input into the classification layer and classified using the trained classifier model, thereby obtaining the classification result. In this embodiment, the classification categories include:

[0188] Minor: Only a threat attempt, without causing actual impact;

[0189] Medium: Has invaded some systems but has not affected the core functions;

[0190] Severe: Causes damage to the core functions of the system or long-term service interruption;

[0191] Critical: Results in large-scale data leakage or long-term network paralysis across the board.

[0192] As Figure 4 shown, by comparing the optimization trajectories of the neural network algorithm based on negative curvature gradient descent and the traditional gradient descent method in the parameter space, the global optimization ability of the new algorithm on complex non-convex loss surfaces is verified, and it is observed whether the algorithm can effectively avoid local minima and saddle points. The experimental results show that the loss surface formed by this technology is smoother and has good convergence, while the traditional method shows multiple scattered local minimum regions in the same parameter space, intuitively demonstrating the optimization effect of the new algorithm on the parameter update path by adjusting in the negative curvature direction, and significantly improving the ability to cross non-ideal critical points.

[0193] As Figure 5 shown, by analyzing the advantages of the path integral gradient estimation method through a polar plot, the experimental comparison of the contribution distributions of the traditional single-batch gradient calculation and the new path integral gradient estimation in different propagation directions is carried out, aiming to verify the comprehensive consideration effect on potential data change paths. The experimental results show that the gradient distribution of this technology presents a more uniform radial pattern, with smaller differences in contribution intensities in each direction, while the traditional method has obvious directional preferences and intensity fluctuations, proving that the path integral method effectively improves the comprehensiveness and stability of gradient estimation by integrating multi-path gradient information, and enhances the adaptability of the model to data distribution changes.

[0194] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce a means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or in one block or multiple blocks.

[0195] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction means that implements the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or in one block or multiple blocks.

[0196] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or in one block or multiple blocks.

[0197] In the present invention, specific embodiments are used to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

[0198] Those of ordinary skill in the art will realize that the embodiments described herein are for helping readers understand the principles of the present invention, and it should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations that do not depart from the essence of the present invention according to the technical revelations disclosed in the present invention, and these deformations and combinations are still within the protection scope of the present invention.

Claims

1. A network security situation assessment method based on data mining, characterized in that, Including the following steps: Obtain network security data; Use a fully connected neural network based on negative curvature gradient descent to extract features from the network security data to obtain feature data; Use an extreme learning machine based on quantum potential energy constraint to construct a classifier model, and use the feature data to train the classifier model. The training of the classifier model includes initializing the weight matrix of the extreme learning machine based on quantum potential energy constraint; Propagate the feature data through the weight matrix with quantum potential energy constraint to the hidden layer and calculate the output of the hidden layer; Use an adaptive orthogonal regularization strategy to optimize the weights from the hidden layer to the output layer; Calculate the quantum state correlation of the output values of the hidden layer and dynamically adjust the output layer weights of the extreme learning machine; End the model training after the forward propagation of all training data is completed; Use the trained classifier model to evaluate the network security situation.

2. The network security situation assessment method based on data mining according to claim 1, characterized in that Use a fully connected neural network based on negative curvature gradient descent to extract features from the network security data to obtain feature data, including: initializing the parameters of the neural network; At each iteration, according to the standard gradient, adjust the direction of gradient update to the negative curvature direction, and calculate the update amount of the corresponding neural network weights according to the negative curvature gradient; Update the neural network parameters according to the calculated gradient information and dynamically adjust the learning rate of the neural network; Repeat the iteration until the preset stop iteration condition is met.

3. A network security situation assessment method based on data mining according to claim 2, characterized in that, Calculate the update amount of the corresponding neural network weights according to the negative curvature gradient, specifically: Among them, ΔW p represents the update amount of the neural network weights, η p represents the learning rate of the neural network, β p represents the coefficient for adjusting the influence of negative curvature, L(W p , X p ) represents the loss function of the neural network, and respectively represent the first-order derivative and the second-order derivative of the loss function, W p represents the weights of the neural network; Perform path integral gradient estimation on all connection paths between neurons in each layer during forward propagation through Feynman path integral, specifically: Among them, X p (τ) represents a network security data sample that changes along the potential path τ from the current state X p ; X p represents the current state of network security data during the neural network processing; dτ represents the differential of the path; ∫ represents the integral; L(W p , X p (τ)) is the loss function of the neural network on the current path.

4. A network security situation assessment method based on data mining according to claim 2, characterized in that, Dynamically adjust the learning rate of the neural network, specifically: Among them, represents the learning rate of the neural network at the (t + 1)-th iteration; represents the learning rate of the neural network at the t-th iteration; δ p represents the learning rate decay coefficient, and int(t) represents the current iteration number.

5. A network security situation assessment method based on data mining according to claim 1, characterized in that, Initialize the weight matrix of the extreme learning machine based on quantum potential energy constraint, specifically: Among them, N u represents the number of neurons in the hidden layer of the extreme learning machine; x u represents the input data of the extreme learning machine; i * represents the imaginary unit; θ u represents the random angle; represents the width parameter of the quantum potential; |||| is the L2 norm; The random angle is a dynamic distribution adjusted according to the statistical characteristics of the network security data features, and the calculation method is expressed as: where erf() represents the error function, and represent the mean and variance of the input data of the extreme learning machine, respectively.

6. A network security situation assessment method based on data mining according to claim 5, characterized in that, Propagate the feature data through the weight matrix with quantum potential energy constraint to the hidden layer and calculate the output of the hidden layer, specifically: H u = f u (W u · x u + b u ); Among them, f u () represents the activation function of the extreme learning machine; H u represents the output of the hidden layer of the extreme learning machine; b u represents the bias term of the extreme learning machine; The calculation method of the activation function of the extreme learning machine is expressed as: Among them, z us represents the input of the activation function of the extreme learning machine; |z us | 2 represents the square of the modulus of the input of the activation function of the extreme learning machine; represents the parameter for controlling the activation intensity.

7. A network security situation assessment method based on data mining according to claim 6, characterized in that, Use an adaptive orthogonal regularization strategy to optimize the weights from the hidden layer to the output layer, specifically: Among them, β u represents the output layer weights of the extreme learning machine; represents the conjugate transpose of the output of the hidden layer of the extreme learning machine; Y represents the target output of the extreme learning machine; λ u represents the regularization parameter of the extreme learning machine; I represents the identity matrix; The calculation method of the regularization parameter of the extreme learning machine is expressed as: Among them, N u represents the number of samples input to the extreme learning machine in the current batch; λ u0 represents the regularization parameter baseline of the extreme learning machine; h uk represents the k-th output of the hidden layer of the extreme learning machine; represents the average value of all hidden layer outputs of the extreme learning machine; ∈ ua represents a set constant; γ u represents the coefficient for controlling the adjustment sensitivity; C u represents the average correlation of the quantum states of the entire batch of data; The calculation method of the average correlation of the quantum states of the entire batch of data is expressed as: Among them, <ψ ui | represents the quantum state after the conversion of the i-th feature data input to the extreme learning machine through the hidden layer; |ψ uj > represents the quantum state after the conversion of the j-th feature data input to the extreme learning machine through the hidden layer; <ψ ui |ψ uj > represents <ψ ui | and |ψ uj > the inner product between the two quantum states.

8. A network security situation assessment method based on data mining according to claim 1, characterized in that, After obtaining the network security data, use a generative adversarial network based on adaptive loss to perform data augmentation on the network security data, including: initializing the parameters of the generator and the discriminator; In the generation stage, the generator generates a batch of new network security data samples according to the current network parameters; Use the discriminator to evaluate the generated network security data samples and the real network security data samples, and output the probability of each sample being a real sample; In the adversarial training stage, calculate the adaptive loss function based on the inherent complexity of the network security data and the stability of the generator during training; Use a dynamic game strategy to adjust the learning rates and update strategies of the generator and the discriminator, and dynamically adjust the loss function according to the results of each round of confrontation; Repeat the iteration until the preset stop iteration condition is met.

Citation Information

Patent Citations

  • Method, device, equipment and medium for constructing network intrusion detection model

    CN118101274B

  • An alternative modeling approach for predicting saltwater intrusion dynamics under uncertainty

    CN118114592B

  • Railway foreign matter invasion simulation method based on artificial intelligence generation and digital twinning

    CN118551647A

  • Network security data analysis method based on big data

    CN118944919A