Network security situation assessment method based on data mining

By applying adaptive loss generative adversarial networks, negative curvature gradient descent optimization algorithms and quantum potential energy constraints in the field of network security, the problems of data scarcity and sample imbalance are solved, the generalization ability and recognition accuracy of the model are improved, and the training efficiency and classification accuracy are significantly improved.

CN120074967AActive Publication Date: 2025-05-30CHENGDU ANZHUN NETWORK SECURITY TECH CO LTD

Patent Information

Application Number
CN202510559535.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-05-30
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

The existing technology faces the problems of data scarcity and sample imbalance in the field of network security, which leads to poor generalization capabilities of the model. Traditional methods are inefficient in training when processing high-dimensional complex data, and are prone to problems such as gradient vanishing, gradient explosion or local optimal solutions.

Method used

Data augmentation network algorithm based on adaptive loss is used to augment data, feature extraction is performed through negative curvature gradient descent optimization algorithm, and classifier models are built using extreme learning institutions based on quantum potential energy constraints.

Benefits of technology

It improves the quality and quantity of network security data, enhances the generalization ability and recognition accuracy of the model, significantly improves the training efficiency of neural networks in high-dimensional data environments, avoids overfitting, and improves the classification accuracy and generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120074967A_ABST
    Figure CN120074967A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and discloses a network security situation assessment method based on data mining, which comprises the following steps: acquiring network security data; performing feature extraction on the network security data by adopting a full-connection neural network based on negative curvature gradient descent to obtain feature data; building a classifier model by using an extreme learning machine based on quantum potential energy constraint, and training the classifier model by using the feature data; and performing network security situation assessment by using the trained classifier model. When variable network attack types are processed, a more stable and accurate classification result can be provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly relates to a network security situation assessment method based on data mining. Background Art

[0002] With the rapid development of information technology, the forms and means of cyberattacks have become increasingly complex and diverse, and traditional network security protection means and emergency response mechanisms are gradually difficult to cope with the increasingly complex network threats. In this context, how to effectively evaluate and predict the network security situation has become an important issue in network security research and practice.

[0003] The Chinese invention patent with the publication number CN118114592B proposes an alternative modeling method for predicting the dynamic of saltwater intrusion under uncertain conditions, including the following steps: establishing a saltwater intrusion model based on physical processes and determining the value range of input parameters; obtaining input samples, importing them into the saltwater intrusion simulation model to obtain an output data set, and constructing an input-output data set; training three machine learning alternative models according to the input-output data set, and importing the input samples into the alternative models to obtain prediction data; combining the chloride concentration observation data with the prediction data of the three alternative models, and using the Bayesian model averaging algorithm to obtain the weights and variances of the models, and constructing an integrated machine learning alternative model of the numerical model. The present invention organically combines the Bayesian averaging algorithm and machine learning, quantifies the model uncertainty, constructs an integrated machine learning alternative model to improve the prediction performance, and proves the feasibility of integrated machine learning alternative modeling under uncertain conditions in predicting the dynamic migration of groundwater pollutants.

[0004] The Chinese invention patent with the publication number CN118551647A proposes a railway foreign object intrusion simulation method based on artificial intelligence generation and digital twin, belonging to the field of artificial intelligence, especially involving generative artificial intelligence technology and digital twin technology; the specific implementation process is as follows: establishing a railway foreign object intrusion scenario based on deep reinforcement learning technology, designing a Monte Carlo method according to the Markov hypothesis of the intrusion foreign object, and constructing a railway foreign object intrusion generation model; based on the foreign object intrusion generation model, using the rainfall weather mathematical model to design a model for adding and removing the railway rainy day effect, and then combining the atmospheric scattering model with the depth of field of the railway scene to construct a railway weather environment model during foreign object intrusion generation, and constructing a railway foreign object intrusion simulation platform based on digital twin. The present invention improves the difficulties and pain points in the current actual work, such as the difficulty in collecting and experimenting with foreign object intrusion data and the limited data sources, and realizes the large-scale generation of available data.

[0005] The Chinese invention patent with the publication number CN118101274B proposes a method, device, equipment and medium for constructing a network intrusion detection model. The method for constructing the network intrusion detection model includes: obtaining a plurality of original data to form an original data set; wherein, the original data includes: numerical network traffic data, data type labels and annotation results; clustering and counting the original data with the same data type label, and respectively obtaining the number of the original data with the same data type label as the clustering and statistical results of each data type label; performing data augmentation operations on the original data set according to the clustering and statistical results of each data type label to obtain a target data set; training a preset model according to the target data set to obtain a network intrusion detection model. Through the technical solution of the present invention, the construction of the network intrusion detection model and the detection of network intrusion can be realized, and the accuracy of network intrusion detection work is improved.

[0006] The following problems still need to be further solved in the prior art:

[0007] 1. The field of network security often faces the problems of data scarcity and sample imbalance. Existing expansion methods mostly rely on simple sampling techniques or rule generation, and cannot effectively simulate complex attack scenarios, resulting in poor generalization ability of the model. Traditional generative adversarial networks often have unstable training when facing skewed data distributions, and the quality of the generated samples is difficult to guarantee.

[0008] 2. Traditional feature extraction methods usually adopt standard gradient descent algorithms, which may encounter problems such as gradient disappearance, gradient explosion or falling into local optimal solutions, resulting in low training efficiency of neural networks and difficulty in effectively processing high-dimensional complex data.

[0009] 3. Traditional network security classification models have poor accuracy and generalization ability when facing complex network attack data. Although the extreme learning machine has a fast training speed, it lacks effective modeling of complex data relationships, which easily leads to insufficient classification accuracy and overfitting. Summary of the Invention

[0010] In view of the above deficiencies in the prior art, the present invention provides a network security situation assessment method based on data mining.

[0011] In order to achieve the above invention purpose, the technical solution adopted by the present invention is:

[0012] A network security situation assessment method based on data mining, including the following steps:

[0013] Obtain network security data;

[0014] Use a fully connected neural network based on negative curvature gradient descent to extract features from the network security data to obtain feature data;

[0015] A classifier model is constructed using an extreme learning machine based on quantum potential energy constraints, and the classifier model is trained using feature data;

[0016] The trained classifier model is used for network security situation assessment.

[0017] Furthermore, a fully connected neural network based on negative curvature gradient descent is used to extract features from network security data to obtain feature data, including:

[0018] Initialize the parameters of the neural network;

[0019] At each iteration, adjust the direction of gradient update to the negative curvature direction according to the standard gradient, and calculate the update amount of the corresponding neural network weights according to the negative curvature gradient;

[0020] Update the neural network parameters according to the calculated gradient information, and dynamically adjust the learning rate of the neural network;

[0021] Repeat the above steps until the preset iteration stop condition is met.

[0022] Furthermore, the update amount of the corresponding neural network weights is calculated according to the negative curvature gradient, specifically:

[0023] ;

[0024] Where, represents the update amount of the neural network weights, represents the learning rate of the neural network, represents the coefficient for adjusting the influence of negative curvature, represents the loss function of the neural network, and respectively represent the first derivative and the second derivative of the loss function, represents the weights of the neural network; sign( ) is the sign function;

[0025] Path integral gradient estimation is performed on all connection paths between neurons in each layer during forward propagation through the Feynman path integral, specifically:

[0026] ;

[0027] Where, represents the network security data sample that changes along the potential path from the current state ; represents the current state of the network security data during the neural network processing; represents the differential of the path; represents the integral; is the loss function of the neural network on the current path.

[0028] Furthermore, the learning rate of the neural network is dynamically adjusted, specifically as follows:

[0029] ;

[0030] where represents the learning rate of the neural network at the -th iteration; represents the learning rate of the neural network at the -th iteration; represents the learning rate decay coefficient, represents the current iteration number.

[0031] Furthermore, a classifier model is constructed using an extreme learning machine based on quantum potential constraint, and the classifier model is trained using feature data, including:

[0032] Initializing the weight matrix of the extreme learning machine based on quantum potential constraint;

[0033] Propagating the feature data through the weight matrix with quantum potential constraint to the hidden layer and calculating the hidden layer output;

[0034] Optimizing the weights from the hidden layer to the output layer using an adaptive orthogonal regularization strategy;

[0035] Calculating the quantum state correlation of the hidden layer output value and dynamically adjusting the output layer weights of the extreme learning machine;

[0036] Ending the model training after the forward propagation of all training data is completed.

[0037] Furthermore, initializing the weight matrix of the extreme learning machine based on quantum potential constraint, specifically as follows:

[0038] ;

[0039] where represents the number of neurons in the hidden layer of the extreme learning machine; represents the input data of the extreme learning machine; represents the imaginary unit; represents a random angle; represents the width parameter of the quantum potential; is the L2 norm;

[0040] The random angle is a dynamic distribution adjusted based on the statistical characteristics of network security data features, and the calculation method is expressed as:

[0041] ;

[0042] Among them, represents the error function, and respectively represent the mean and variance of the input data of the extreme learning machine.

[0043] Furthermore, the feature data is propagated to the hidden layer through the weight matrix with quantum potential constraints, and the output of the hidden layer is calculated. Specifically:

[0044] ;

[0045] Among them, represents the activation function of the extreme learning machine; represents the output of the hidden layer of the extreme learning machine; represents the bias term of the extreme learning machine;

[0046] The calculation method of the activation function of the extreme learning machine is expressed as:

[0047] ;

[0048] Among them, represents the input of the activation function of the extreme learning machine; represents the square of the modulus of the input of the activation function of the extreme learning machine; represents the parameter for controlling the activation intensity.

[0049] Furthermore, the weights from the hidden layer to the output layer are optimized using an adaptive orthogonal regularization strategy. Specifically:

[0050] ;

[0051] Among them, represents the output layer weights of the extreme learning machine; represents the conjugate transpose of the output of the hidden layer of the extreme learning machine; represents the target output of the extreme learning machine; represents the regularization parameter of the extreme learning machine; represents the identity matrix;

[0052] The calculation method of the regularization parameter of the extreme learning machine is expressed as:

[0053] ;

[0054] Among them, represents the number of samples input to the extreme learning machine in the current batch; represents the regularization parameter baseline of the extreme learning machine; represents the -th output of the hidden layer of the extreme learning machine; Represents the average value of the outputs of all hidden layers of the extreme learning machine; Represents a set constant; Represents the coefficient for controlling the adjustment sensitivity; Represents the average correlation of the quantum states of the entire batch of data.

[0055] The calculation method of the average correlation of the quantum states of the entire batch of data is expressed as:

[0056] ;

[0057] Among them, Represents the quantum state after the th feature data input to the extreme learning machine is transformed by the hidden layer; Represents the quantum state after the th feature data input to the extreme learning machine is transformed by the hidden layer; Represents and the inner product between two quantum states.

[0058] Furthermore, after obtaining the network security data, a generative adversarial network based on adaptive loss is used to perform data augmentation on the network security data, including:

[0059] Initializing the parameters of the generator and the discriminator;

[0060] In the generation stage, the generator generates a batch of new network security data samples according to the current network parameters;

[0061] Using the discriminator to evaluate the generated network security data samples and the real network security data samples, and outputting the probability that each sample is a real sample;

[0062] In the adversarial training stage, based on the inherent complexity of the network security data and the stability of the generator during training, an adaptive loss function is calculated;

[0063] Using the dynamic game strategy to adjust the learning rates and update strategies of the generator and the discriminator, and dynamically adjusting the loss function according to the results of each round of confrontation;

[0064] Repeating the iteration until the preset stop iteration condition is satisfied.

[0065] The present invention has the following beneficial effects:

[0066] 1. Through the generative adversarial network algorithm based on adaptive loss, the present invention solves the problems of insufficient data and sample skew. The generated network security data not only increases in quantity but also improves in quality, can better simulate the complex patterns of actual attacks, and helps to improve the generalization ability and recognition accuracy of subsequent models.

[0067] 2. The present invention uses a negative curvature gradient descent optimization algorithm to significantly improve the training efficiency of neural networks in a high-dimensional data environment, overcoming problems such as local optimal solutions and gradient vanishing in traditional gradient descent methods. Especially when dealing with complex network security data, it can more accurately extract effective features, thereby improving the accuracy of subsequent classification algorithms.

[0068] 3. The present invention uses an extreme learning machine classification algorithm based on quantum potential energy constraints to better capture complex non-linear relationships in network security data, avoid overfitting, and significantly improve the classification accuracy and the generalization ability of the model. Especially when dealing with diverse types of network attacks, this method can provide more stable and accurate classification results. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] Figure 1 It is a schematic flow diagram of a network security situation assessment method based on data mining;

[0070] Figure 2 It is a schematic architecture diagram of a network security situation assessment platform based on data mining;

[0071] Figure 3 It is a schematic diagram for comparing the training stability of different GAN methods;

[0072] Figure 4 It is a schematic diagram for comparing the optimization trajectories of a neural network algorithm based on negative curvature gradient descent and the traditional gradient descent method in the parameter space;

[0073] Figure 5 It is a schematic diagram for comparing the path integral gradient contribution distributions. DETAILED DESCRIPTION OF THE INVENTION

[0074] The following describes the specific embodiments of the present invention to facilitate those skilled in the art of the present technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.

[0075] As Figure 1 shown, the embodiments of the present invention provide a network security situation assessment method based on data mining, including the following steps S1 to S4:

[0076] S1. Obtain network security data;

[0077] In an alternative embodiment of the present invention, as Figure 2As shown in the figure, at the network data collection layer of the platform, the collection of network security data comes from a variety of network security monitoring systems, such as intrusion detection systems, user operation logs, firewall logs, and network traffic analysis tools, etc.

[0078] The collection methods of network security data include real-time monitoring and historical record analysis. Through API interfaces and network packet capture tools, network security data sets are automatically collected and updated.

[0079] The storage format of network security data adopts the structured JSON format. In one embodiment, the attributes of network security data include:

[0080] Ra is the attack type code, such as DDoS, malware, etc.; Da is the importance rating of the target asset; Ca is the timestamp when the attack occurred; Pa is the network protocol type, such as TCP, UDP; Qa is the source IP address; Fa is the target IP address; Va is the source port number; Wa is the target port number; Na is the number of transmitted data packets; Za is the flag indicating whether the attack is successful or not.

[0081] It should be noted that this embodiment is only to illustrate a network security data format and type of the present invention. In actual applications, the attributes of network security data are usually more than 10, and the number of attributes of network security data may reach dozens or even hundreds.

[0082] Furthermore, the collected network security data is labeled. The labeling method of the present invention is manual labeling. In one embodiment, the labeling categories include:

[0083] Minor: Only a threat attempt, without actual impact;

[0084] Medium: Has invaded part of the system but has not affected the core functions;

[0085] Severe: Damages the core functions of the system or causes a long-term service interruption;

[0086] Critical: Causes large-scale data leakage or long-term paralysis of the entire network.

[0087] In this embodiment, 5 examples of network security data are shown in Table 1.

[0088] Table 1

[0089] It can be understood that in the task of the present invention, the collection, acquisition, labeling, and preprocessing of network security training data are time-consuming and laborious, and insufficient training samples are likely to lead to poor generalization ability of the model and affect the accuracy of the model at the same time.

[0090] In the sample enhancement layer of the platform, network security data augmentation is achieved through a generative model. The present invention uses an algorithm of generative adversarial network based on adaptive loss for sample generation, thereby realizing network security data augmentation. Traditional generative adversarial networks often have difficulty generating relatively complex samples when facing skewed network security data distributions. Based on the inherent complexity of network security data and the stability of the generator during training, an adaptive loss function is calculated to promote the fast and stable learning of the generative model, enabling the network security data generated each time to better simulate complex data distributions.

[0091] Specifically, the training process of the generative adversarial network algorithm based on adaptive loss is as follows:

[0092] S101. Initialize the parameters of the generator and the discriminator. Let the generator be , and the discriminator be . Set the weight of the generator as , and the weight of the discriminator as. In a certain situation, the initialization method is expressed as:

[0093]

[0094]

[0095] where is a normal distribution with a mean of 0 and a standard deviation of the identity matrix; obeys a specific distribution; is the identity matrix; is a normal distribution.

[0096] S102. In the generation stage, the generator generates a batch of new network security data samples according to the current network parameters. These samples initially attempt to imitate the distribution characteristics of real network security data, expressed as:

[0097]

[0098] where is a latent vector sampled from the prior distribution; is the generator function, is the th iteration of the generator weight; is at the th iteration, the sample generated by the generator; represents the weight of the feature transformation; represents element-wise multiplication.

[0099] In one embodiment, the quality of the generated samples is improved through an adaptive feature transformation strategy, and the learning ability of the model for different network security data is strengthened. Based on the dynamic weighted sum and adjustment of the input features of the generator, feature weight learning is achieved. The calculation method of the weights of the feature transformation is expressed as:

[0100]

[0101] Wherein, is the Sigmoid activation function; is the parameter matrix of weight learning, specifically the weight parameter matrix of the generator.

[0102] S103. The discriminator evaluates the generated network security data samples and the real network security data samples, and outputs the probability that each sample is a real sample, which is expressed as:

[0103]

[0104] Wherein, means that the generated samples and the real samples are combined into one batch for input; is the discriminator function, is the discriminator weight at the th iteration; is the real sample; is the discrimination result of the discriminator in the th iteration.

[0105] S104. In the adversarial training stage, based on the inherent complexity of the network security data and the stability of the generator during training, an adaptive loss function is calculated to promote the fast and stable learning of the generation model. The calculation method is expressed as:

[0106]

[0107] Wherein, is the reconstruction loss function, specifically the reconstruction loss function for calculating the generated network security data and the real network security data, which is obtained by calculating the L2 norm; is the complexity-dependent loss function; the calculation method of the adjustment coefficient at the th iteration is expressed as ;

[0108] Furthermore, the calculation method of the complexity-dependent loss function is expressed as:

[0109]

[0110] Wherein, is the complexity scoring function; represents the expectation; To conform to a specific distribution; For the distribution of real network security data; For the noise distribution.

[0111] Furthermore, the complexity scoring function evaluates the complexity of the input network security data, specifically evaluating the complexity or rarity of the abnormal patterns contained in the generated network security data. Let the input of the complexity scoring function be , then the calculation method of the complexity scoring function is expressed as:

[0112]

[0113] Among them, Is the variance calculation function, used to measure the irregularity and prediction difficulty of network security data; Represents the normalization parameter, preset by humans, such as set to 0.1, Is the input of the complexity scoring function, corresponding to Or .

[0114] Furthermore, the adjustment coefficient is adjusted according to the number of iterations, used to adjust the weight between the standard adversarial loss and the loss based on the current training complexity. The calculation method is expressed as:

[0115]

[0116] Among them, Is the parameter that controls the steepness of the curve, Is the preset iteration period. When the training reaches a certain period, the complexity loss starts to increase, ensuring that as the training progresses, the model gradually increases its attention to complex network security data; Is the current number of iterations. Preferably, Is set to 2.

[0117] S105. Use the dynamic game strategy to adjust the learning rates and update strategies of the generator and discriminator, and dynamically adjust the loss function according to the results of each round of confrontation, giving priority to improving the ability to generate scarce category network security data, expressed as:

[0118]

[0119]

[0120] Among them, Is the weight of the generator at the th iteration; Is the weight of the discriminator at the th iteration; is the learning rate of the generator; is the learning rate of the discriminator; is the loss function of the generator; is the loss function of the discriminator; represents the gradient with respect to the weights of the generator; represents the weights with respect to the parameters of the discriminator.

[0121] Furthermore, the loss function of the generator is comprehensively calculated in combination with the entropy of the generated network security data to more comprehensively evaluate the quality of the generated network security data and increase the diversity of the generated samples, expressed as:

[0122]

[0123] Among them, represents the entropy function, is the adjustment coefficient that controls the influence of the entropy term; is the adaptive loss function.

[0124] Furthermore, the loss function of the discriminator uses a regularization term to prevent overfitting, expressed as:

[0125]

[0126] Among them, represents the true sample label (1 for true sample, 0 for generated sample); is the regularization coefficient for weight decay to reduce the model complexity and the risk of overfitting; represents the square of the L2 norm of the discriminator weights.

[0127] During the adversarial training process of the generator and the discriminator, the dynamic game strategy dynamically adjusts the adjustment coefficient and the regularization coefficient that control the influence of the entropy term so that the network security data generated each time better simulates the real data distribution, expressed as:

[0128]

[0129]

[0130] Among them, is the learning rate adjustment factor for controlling the dynamic adjustment of the learning rate; is the partial derivative symbol; is the th iteration's adjustment coefficient that controls the influence of the entropy term, is the th iteration's regularization coefficient; is the th iteration's adjustment coefficient that controls the influence of the entropy term, is the The regularization coefficient for the next iteration. Preferably, it is set to 0.01.

[0131] S106. Repeat the above steps iteratively until the preset iteration stop condition is met, which indicates that the model training is completed. In one embodiment, the preset iteration stop condition is to reach the preset maximum number of iterations. Preferably, the preset maximum number of iterations is set to 1000 times.

[0132] After the network security data augmentation model training is completed, the trained network security data augmentation model is used to increase the number of samples. In one embodiment, assume that the original collected samples are 800, and the network security data augmentation model generates 200 samples through augmentation. Then the augmented network security data set contains 1000 samples.

[0133] As Figure 3 shown, by comparing the training process stability of different generative adversarial network methods, the effectiveness of the dynamic game strategy is verified. The experiment compares the adaptive loss algorithm proposed in the present invention with the classical generative adversarial network (original GAN), the least squares generative adversarial network (LSGAN), and the generative adversarial network with gradient penalty (WGAN), and observes the trend of the generator loss value changing with the number of iterations. The experimental results show that the fluctuation amplitude of the loss curve of the method of the present invention is significantly smaller than that of other methods, especially there is no severe oscillation or divergence phenomenon in the middle and late stages of training, indicating that the dynamically adjusted weight coefficient in the adaptive loss function can automatically balance the contributions of the reconstruction loss and the complexity loss according to the model training state, avoiding the problem of single optimization direction caused by fixed loss weights in traditional methods, thus effectively suppressing the mode collapse phenomenon and enabling the model to converge continuously and stably.

[0134] S2. Use a fully connected neural network based on negative curvature gradient descent to extract features from the network security data to obtain feature data;

[0135] In an alternative embodiment of the present invention, at the feature extraction layer of the platform, a neural network algorithm based on negative curvature gradient descent is used as the feature extraction model to extract features from the augmented network security data.

[0136] The present invention uses a 5-layer fully connected neural network for feature extraction. In the prior art, some solutions use neural networks for feature extraction. In some neural network structures, problems such as gradient disappearance, gradient explosion, or getting stuck in local optimal solutions may be encountered, affecting the training stability and model performance.

[0137] In the present invention, negative curvature gradient descent accelerates the convergence of gradients during training by leveraging the negative curvature direction. Especially in complex, high-dimensional data environments, it can significantly improve the efficiency and accuracy of the model when processing non-linear network security data.

[0138] Specifically, the training process of the neural network algorithm based on negative curvature gradient descent is as follows:

[0139] S201. Initialize the parameters of the neural network. In one embodiment, the initialization method is expressed as:

[0140]

[0141]

[0142] In the formula, represents the initial weight matrix of the neural network, represents the initial bias vector of the neural network, and represent the dimensions of the input layer and output layer of the neural network respectively; is the zero vector function; is the function to take random values; is the weight of the neural network; is the bias of the neural network.

[0143] S202. In some cases, the traditional gradient descent direction may not be optimal. In the present invention, in each iteration, first calculate the standard gradient, and then determine whether the direction of gradient update needs to be adjusted to the negative curvature direction. The calculation method of the adjustment of the negative curvature gradient corresponding to the update amount of the neural network weight is expressed as:

[0144]

[0145] In the formula, is the update amount of the neural network weight, is the learning rate of the neural network, is the coefficient to adjust the influence of negative curvature, is the loss function of the neural network, and sign( ) is the sign function; and represent the first-order and second-order derivatives of the loss function respectively. By using the sign function of the second-order derivative to adjust the direction of gradient descent, as a means of adjusting the negative curvature gradient, it is convenient to cross the saddle point during parameter update.

[0146] S203. For each training batch, not only calculate the gradient of the current batch, but also consider the gradient contributions of all possible paths from the current state to possible future states through Feynman path integration. Moreover, integrating over potential state transitions takes into account the contributions of all possible paths of network security data changes to the current gradient, in order to more comprehensively estimate the gradient and estimate the impact of these transitions on the current parameter update. The path integral gradient estimation method is expressed as:

[0147]

[0148] In the formula, represents the network security data sample that changes along the potential path from the current state ; is the current state of the network security data during the neural network processing; represents the path of the network security data features during forward propagation, corresponding to the connection paths between neurons in each layer; represents the differentiation with respect to the path; represents the integration; is the loss function of the neural network on the current path.

[0149] S204. Update the neural network parameters according to the calculated gradient information. Moreover, the learning rate of the neural network dynamically adjusts to cope with the problems of gradient disappearance or explosion during the training process, and the learning rate is slowly decreased as the training progresses to improve the training stability and promote convergence. The adjustment method is expressed as:

[0150]

[0151] In the formula, is the learning rate of the neural network at the -th iteration; is the learning rate of the neural network at the -th iteration; is the learning rate decay coefficient, is the current iteration number. Preferably, is set to 0.95.

[0152] Furthermore, the update method of the weight parameter of the neural network is expressed as:

[0153]

[0154] In the formula, represents the weight of the neural network at the -th iteration, represents the weight of the neural network at the -th iteration.

[0155] S205. Repeat the above steps iteratively until the preset iteration stop condition is met, indicating that the model training is completed. In one embodiment, the preset iteration stop condition is to reach the preset maximum number of iterations. Preferably, the preset maximum number of iterations is set to 1000 times.

[0156] S3. Construct a classifier model using an extreme learning machine based on quantum potential energy constraint, and train the classifier model using the feature data.

[0157] In an alternative embodiment of the present invention, in the classification layer of the platform, an extreme learning machine classification algorithm based on quantum potential energy constraint is used as the classifier model.

[0158] Although the traditional extreme learning machine algorithm has an advantage in training speed, when dealing with network security data with complex relationships, it is prone to problems such as insufficient classification accuracy and insufficient generalization ability.

[0159] The quantum potential energy constraint strategy is inspired by the potential energy model in quantum mechanics. It uses the wave function characteristics of quantum states to constrain the characteristics of network security training data of the extreme learning machine, enhances the model's perception ability of the internal relationships of network security data, prevents overfitting, and improves the classification accuracy and generalization ability.

[0160] Specifically, the training process of the extreme learning machine classification algorithm based on quantum potential energy constraint is as follows:

[0161] S301. Considering the influence of quantum potential energy to ensure that the preliminary non-linear transformation of network security data can reflect its quantum characteristics, initialize the weight matrix of the extreme learning machine, expressed as:

[0162]

[0163] In the formula, represents the number of neurons in the hidden layer of the extreme learning machine; represents the input data of the extreme learning machine, and the input data of the extreme learning machine is the feature vector obtained by feature extraction; represents the imaginary unit; is a random angle; represents the width parameter of the quantum potential energy, controlling the dispersion degree of the quantum state; is the L2 norm.

[0164] Furthermore, considering quantum randomness, the random angle is a dynamic distribution adjusted based on the statistical characteristics of network security data features, and the calculation method is expressed as:

[0165]

[0166] In the formula, is the error function, and are the mean and variance of the input data of the extreme learning machine respectively, which are used to adjust the distribution of the random angles to adapt to the different characteristics of network security data.

[0167] S302. Calculate the output of the hidden layer of the extreme learning machine. The network security data after feature extraction is propagated to the hidden layer through the weight matrix with quantum potential constraint, and its output is calculated, expressed as:

[0168]

[0169] In the formula, is the activation function of the extreme learning machine; is the output of the hidden layer of the extreme learning machine; is the bias term of the extreme learning machine, and the same quantum initialization is adopted as the weight of the extreme learning machine.

[0170] Furthermore, the activation function of the extreme learning machine is a quantized activation function, which uses the quantum interference effect to simulate the nonlinear mapping change of the network security data after feature extraction, and can effectively increase the nonlinear expression ability of the model for features. The calculation method is expressed as:

[0171]

[0172] In the formula, is the input of the activation function of the extreme learning machine; represents the square of the modulus of the input of the activation function of the extreme learning machine; is the parameter controlling the activation intensity.

[0173] S303. Optimize the weight from the hidden layer to the output layer by using the adaptive orthogonal regularization strategy. Through the regularization process, the weight matrix is adaptively adjusted to optimize the output performance of the algorithm. The regularization method of the output layer weight of the extreme learning machine is expressed as:

[0174]

[0175] In the formula, is the output layer weight of the extreme learning machine; represents the conjugate transpose of the output of the hidden layer of the extreme learning machine; is the target output of the extreme learning machine, representing the true label vector of the sample; is the regularization parameter of the extreme learning machine, which is adaptively adjusted through the change of the quantum state; represents the identity matrix.

[0176] Furthermore, the regularization parameter of the extreme learning machine can be dynamically adjusted according to the data of the current batch, better adapting to the characteristics of network security data after feature extraction and preventing overfitting. The calculation method is expressed as:

[0177]

[0178] In the formula, is the number of samples input to the extreme learning machine in the current batch; is the baseline of the regularization parameter of the extreme learning machine; is the output of the th hidden layer of the extreme learning machine; is the average value of all hidden layer outputs of the extreme learning machine; is a small constant to avoid the denominator being zero; is the coefficient to control the adjustment sensitivity; represents the average correlation of the quantum states of the entire batch of data. The higher its value, the closer the internal connection between the data. Preferably, is set to 0.001, is set to 0.01.

[0179] S304. Adopt a dynamic adjustment mechanism based on the correlation of quantum states to improve the adaptability and accuracy of the algorithm in dealing with complex network security environments. By analyzing the correlation of quantum states of the hidden layer outputs, dynamically adjust the output layer weights of the extreme learning machine, enabling the algorithm to adaptively adjust the learning strategy based on the internal connection and changes of the network security data after feature extraction. The calculation method of the average correlation of the quantum states of the entire batch of data is expressed as:

[0180]

[0181] In the formula, represents the quantum state after the th data input to the extreme learning machine is transformed by the hidden layer; represents the quantum state after the th data input to the extreme learning machine is transformed by the hidden layer; represents and the inner product between two quantum states, used to measure their similarity.

[0182] S305. After the forward propagation of all training data is completed, the model training stage ends. At this time, the model can accurately classify according to the input data features. The final output of the model is expressed as:

[0183]

[0184] In the formula, It is the model output of the extreme learning machine, corresponding to the prediction probability vector of the extreme learning machine for the samples. The category corresponding to the element with the largest value in this vector is taken as the prediction category of the extreme learning machine for the samples.

[0185] S5. Use the trained classifier model to conduct network security situation assessment.

[0186] In an optional embodiment of the present invention, at the network data application layer of the platform, the trained model is used to process new samples and evaluate the network security situation. In one embodiment, the original network security data collected by the network data collection layer is input into the feature extraction layer and undergoes feature processing in the trained feature extraction model. Further, the processed features are input into the classification layer and classified in the trained classifier model to obtain the classification result. In this embodiment, the classification categories include:

[0187] Slight: Only a threat attempt, without causing actual impact;

[0188] Medium: Has invaded part of the system but has not affected the core functions;

[0189] Severe: Has damaged the core functions of the system or caused a long-term service interruption;

[0190] Critical: Has led to large-scale data leakage or long-term paralysis of the entire network.

[0191] As Figure 4 shown, by comparing the optimization trajectories of the neural network algorithm based on negative curvature gradient descent and the traditional gradient descent method in the parameter space, the global optimization ability of the new algorithm on complex non-convex loss surfaces is verified, and it is observed whether the algorithm can effectively avoid local minima and saddle points. The experimental results show that the loss surface formed by this technology is smoother and has good convergence, while the traditional method shows multiple scattered local minimum regions in the same parameter space, intuitively demonstrating the optimization effect of the new algorithm on the parameter update path by adjusting in the negative curvature direction, significantly improving the ability to cross non-ideal critical points.

[0192] As Figure 5 shown, by analyzing the advantages of the path integral gradient estimation method through a polar plot, the experimental comparison of the contribution distributions of the traditional single-batch gradient calculation and the new path integral gradient estimation in different propagation directions aims to verify the comprehensive consideration effect on potential data change paths. The experimental results show that the gradient distribution of this technology presents a more uniform radial pattern, with smaller differences in the contribution intensities in each direction, while the traditional method has obvious directional preferences and intensity fluctuations, proving that the path integral method effectively improves the comprehensiveness and stability of gradient estimation by integrating multi-path gradient information, enhancing the adaptability of the model to data distribution changes.

[0193] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combinations of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0194] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0195] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are performed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0196] Specific embodiments are applied in the present invention to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

[0197] Those of ordinary skill in the art will realize that the embodiments described herein are for helping readers understand the principles of the present invention, and it should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations that do not depart from the essence of the present invention based on the technical revelations disclosed in the present invention, and these deformations and combinations are still within the protection scope of the present invention.

Claims

1. A network security situation assessment method based on data mining, characterized in that: The following steps are involved: Obtain cybersecurity data; A fully connected neural network based on negative curvature gradient descent is used to extract features from network security data to obtain feature data; The classifier model is constructed by using the extreme learning machine based on quantum potential energy constraints, and the classifier model is trained using feature data; Use the trained classifier model to conduct network security situation assessment.

2. A network security situation assessment method based on data mining according to claim 1, characterized in that: A fully connected neural network based on negative curvature gradient descent is used to extract features from network security data to obtain feature data, including: Initialize the parameters of the neural network; At each iteration, the direction of the gradient update is adjusted to the negative curvature direction according to the standard gradient, and the update amount of the corresponding neural network weight is calculated according to the negative curvature gradient; Update the neural network parameters based on the calculated gradient information and dynamically adjust the learning rate of the neural network; Repeat the iteration until the preset stop iteration condition is met.

3. A network security situation assessment method based on data mining according to claim 2, characterized in that: The update amount of the corresponding neural network weight is calculated according to the negative curvature gradient, specifically: ; in, represents the update amount of the neural network weights, represents the learning rate of the neural network, Represents the coefficient for adjusting the negative curvature effect, represents the loss function of the neural network, and They represent the first and second order derivatives of the loss function respectively, represents the weight of the neural network; The Feynman path integral is used to estimate the path integral gradient of all connection paths between neurons in each layer during forward propagation, specifically: ; in, Indicates that from the current state Along potential paths Changing cybersecurity data samples; Represents the current state of cybersecurity data during neural network processing; represents the differentiation of the path; represents integral; is the loss function of the neural network in the current path.

4. A network security situation assessment method based on data mining according to claim 2, characterized in that: Dynamically adjust the learning rate of the neural network, specifically: ; in, Indicates The learning rate of the neural network for each iteration; Indicates The learning rate of the neural network for each iteration; represents the learning rate decay coefficient, Indicates the current iteration number.

5. A network security situation assessment method based on data mining according to claim 1, characterized in that: The classifier model is constructed using an extreme learning machine based on quantum potential energy constraints, and the classifier model is trained using feature data, including: Initialize the weight matrix of the extreme learning machine based on quantum potential constraints; The feature data is propagated to the hidden layer through the weight matrix with quantum potential energy constraints, and the hidden layer output is calculated; Optimize the weights from hidden layers to output layers using an adaptive orthogonal regularization strategy; Calculate the quantum state correlation of the hidden layer output value and dynamically adjust the output layer weight of the extreme learning machine; Model training ends after the forward propagation of all training data is completed.

6. A network security situation assessment method based on data mining according to claim 5, characterized in that: The weight matrix of the extreme learning machine is initialized based on the quantum potential energy constraint, specifically: ; in, represents the number of neurons in the hidden layer of the extreme learning machine; Represents the input data of the extreme learning machine; represents an imaginary unit; represents a random angle; represents the width parameter of the quantum potential; is the L2 norm; The random angle is a dynamic distribution adjusted based on the statistical characteristics of network security data features, and the calculation method is expressed as: ; in, represents the error function, and They represent the mean and variance of the input data of the extreme learning machine respectively.

7. A network security situation assessment method based on data mining according to claim 6, characterized in that: The feature data is propagated to the hidden layer through the weight matrix with quantum potential energy constraints, and the hidden layer output is calculated, specifically: ; in, represents the activation function of the extreme learning machine; represents the hidden layer output of the extreme learning machine; represents the bias term of the extreme learning machine; The calculation method of the activation function of the extreme learning machine is expressed as: ; in, Represents the input of the activation function of the extreme learning machine; represents the square of the modulus of the input to the activation function of the extreme learning machine; Represents the parameter that controls the strength of the activation.

8. A network security situation assessment method based on data mining according to claim 7, characterized in that: The adaptive orthogonal regularization strategy is used to optimize the weights from the hidden layer to the output layer, specifically: ; in, represents the output layer weight of the extreme learning machine; represents the conjugate transpose of the hidden layer output of the extreme learning machine; represents the target output of the extreme learning machine; represents the regularization parameter of the extreme learning machine; represents the identity matrix; The calculation method of the regularization parameter of the extreme learning machine is expressed as: ; in, Indicates the number of samples input to the extreme learning machine in the current batch; represents the regularization parameter baseline of the extreme learning machine; represents the hidden layer of the extreme learning machine Outputs; Represents the average value of all hidden layer outputs of the extreme learning machine; Indicates setting constant; The coefficient representing the control adjustment sensitivity; Represents the average correlation of the quantum states of the entire batch of data; The calculation method of the average correlation of the quantum state of the entire batch of data is expressed as: ; in, represents the first The quantum state of feature data after conversion by the hidden layer; represents the first The quantum state of feature data after conversion by the hidden layer; express and The inner product between two quantum states.

9. A network security situation assessment method based on data mining according to claim 1, characterized in that: After obtaining the network security data, a generative adversarial network based on adaptive loss is used to expand the network security data, including: Initialize the parameters of the generator and discriminator; In the generation phase, the generator generates a batch of new network security data samples based on the current network parameters; Use the discriminator to evaluate the generated network security data samples and the real network security data samples, and output the probability that each sample is a real sample; In the adversarial training phase, an adaptive loss function is calculated based on the inherent complexity of cybersecurity data and the stability of the generator during training. Use dynamic game strategies to adjust the learning rate and update strategy of the generator and discriminator, and dynamically adjust the loss function according to the results of each round of confrontation; Repeat the iteration until the preset stop iteration condition is met.

Citation Information

Patent Citations

  • Method, device, equipment and medium for constructing network intrusion detection model

    CN118101274B

  • An alternative modeling approach for predicting saltwater intrusion dynamics under uncertainty

    CN118114592B

  • Railway foreign matter invasion simulation method based on artificial intelligence generation and digital twinning

    CN118551647A

  • Network security data analysis method based on big data

    CN118944919A

  • Network security situation element extraction method and system based on hybrid deep learning

    CN119030767A

Cited By

  • Power transmission and transformation foundation flood vulnerability assessment method and system based on optimized neural network

    CN121981557A