Network security risk assessment method based on big data
By constructing a network security risk assessment method based on big data, the problems of data processing bottlenecks and rigid model in the existing technology are solved, efficient and accurate risk assessment and rapid response are achieved, adapting to the evolution of network technology, reducing false alarm rates and optimizing system operation and maintenance costs.
Patent Information
- Application Number
- CN202510658565.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-07-25
AI Technical Summary
The existing technology is difficult to deal with massive heterogeneous data and new attacks, and there are problems such as data processing bottlenecks, rigid model, lag in collaborative defense and inefficient knowledge iteration, especially in hybrid cloud environments and smart city Internet of Things.
A network security risk assessment method based on big data is adopted, and a dynamic risk assessment model is constructed by collecting multi-source heterogeneous data, cleaning and normalizing processing in real time, and a hybrid model of correlation analysis algorithm, random forest classifier and long-term and short-term memory network is used to calculate threat value, and risk levels and protection suggestions are output through dynamic weight allocation and self-learning optimization mechanism.
It improves data processing efficiency and accuracy, enhances the accuracy and robustness of risk assessment, reduces false alarm rate, improves the practicality and response speed of the system, and reduces operation and maintenance costs.
Smart Images

Figure CN120378198A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of network security, and particularly to a network security risk assessment method based on big data. Background Art
[0002] In recent years, network attacks have become intelligent, large-scale, and targeted, challenging the traditional network security risk assessment mechanism. Existing technologies are difficult to cope with massive heterogeneous data and new types of attacks. For example, it is difficult to locate Kubernetes API attacks in a hybrid cloud environment, and the misjudgment rate of attacks on smart meters in the Internet of Things of smart cities is high. The existing technologies have the following disadvantages: data processing bottlenecks, low efficiency of real-time cleaning and feature alignment; model rigidity, static risk assessment models cannot adapt to technological evolution; collaborative defense lag, delay in risk assessment and security device policy execution; inefficient knowledge iteration, lack of self-learning mechanism, and manual intervention is required for rule base updates. Therefore, a network security risk assessment method based on big data is proposed. Summary of the Invention
[0003] The purpose of the present invention is to provide a network security risk assessment method based on big data to solve the problems in the existing technology mentioned in the above background art, namely: data processing bottlenecks, low efficiency of real-time cleaning and feature alignment; model rigidity, static risk assessment models cannot adapt to technological evolution; collaborative defense lag, delay in risk assessment and security device policy execution; inefficient knowledge iteration, lack of self-learning mechanism, and manual intervention is required for rule base updates.
[0004] To achieve the above purpose, the present invention provides the following technical solution: A network security risk assessment method based on big data, including the following steps:
[0005] Step 1: Collect multi-source heterogeneous network security data;
[0006] Step 2: Perform real-time cleaning and normalization processing on the data to generate standardized feature vectors;
[0007] Step 3: Build a dynamic risk assessment model, and calculate the threat value of network nodes through a hybrid model that combines an association analysis algorithm, a random forest classifier, and a long short-term memory network;
[0008] Step 4: Optimize the risk assessment results based on the dynamic weight allocation of threat values, and output the risk level and protection suggestions;
[0009] Step 5: By setting a self-learning optimization mechanism, when the defense strategy is triggered, record the actual protection effect data, and update the model parameters through a reinforcement learning algorithm.
[0010] Preferably, in the above Step 2, the specific steps of the data cleaning and normalization processing include:
[0011] Perform time series segmentation on the data using a sliding time window, and the window length satisfies the formula:
[0012]
[0013] where λ is the threat event occurrence rate, ∈ is the preset error threshold, and N total is the total amount of data, and T min is the minimum window time;
[0014] Calculate the sliding time window length T w for time series segmentation of network security data to balance the threat event capture granularity and computational efficiency;
[0015] Perform one-hot encoding on discrete features and Z-score normalization on continuous features.
[0016] Preferably, in the above step three, the calculation of the threat value in the dynamic risk assessment model satisfies the following formula:
[0017]
[0018] where w i is the dynamic weight of the i-th type of feature, f(x i ) is the classification probability output by RFC, α is the time series correction coefficient, and LSTM(x t-k:t ) is the predicted outlier of historical time series data;
[0019] Calculate the threat score S threat and dynamically evaluate the threat level by combining the static feature classification probability and the time series data prediction result.
[0020] Preferably, the adjustment method of the above-mentioned dynamic weight w i includes:
[0021] Based on the gradient descent optimization algorithm with real-time threat intelligence feedback, the objective function is:
[0022]
[0023] where is the predicted threat value of the j-th sample, is the actual attack event loss value of the j-th sample, β is the regularization parameter used to control the sparsity of the weight w, and w is the weight vector representing the weights of various types of features;
[0024] Update the weight every time interval Δt and constrain w i = 1.
[0025] Preferably, in the fourth step above, the risk level division adopts the fuzzy comprehensive evaluation method, and the risk threshold interval is defined as:
[0026]
[0027] Among them, S threat is the threat score, and θ1 and θ2 are dynamically adjusted according to historical attack data.
[0028] Preferably, the above real-time risk assessment is achieved in the following way:
[0029] By deploying a lightweight risk assessment agent, synchronize data to the central server periodically;
[0030] When a mutation of S is detected threat trigger abnormal traffic mirroring and analysis, generate defense rules and issue them to the firewall.
[0031] Preferably, the above method integrates a visual interaction interface and supports the following functions:
[0032] Display the threat traceability path in the form of a Sankey diagram;
[0033] Dynamically mark vulnerable nodes in the network topology based on the risk heat map;
[0034] Provide an interface for generating automated repair scripts.
[0035] Preferably, the training process of the above long short-term memory network includes:
[0036] The input layer is designed as a three-dimensional tensor X∈R b×t×d , where b is the batch size, t is the time step, and d is the feature dimension;
[0037] The loss function adopts weighted cross-entropy:
[0038]
[0039] Among them, c is the total number of attack types, γ C is the class imbalance correction factor, y c is the one-hot encoding of the true label, indicating whether the sample belongs to class c, and p c is the probability of class c predicted by the model.
[0040] Preferably, in the first step above, the multi-source heterogeneous network security data includes network traffic logs, system vulnerability information, user behavior data, and external threat intelligence.
[0041] Preferably, in the fifth step above, set the reward function:
[0042]
[0043] Among them, the number of blocked attacks is the number of attacks successfully identified and blocked by the model, the number of false alarms is the number of normal behaviors misidentified as attacks by the model, the response delay is the time from the model detecting an attack to taking blocking measures, and η is the delay penalty coefficient.
[0044] Compared with the prior art by adopting the above technical solutions, the present invention has the following technical effects:
[0045] 1. The present invention can efficiently process massive data, overcomes the real-time and accuracy problems existing in traditional methods when processing multi-source data, and the application of technologies such as sliding time window and Z-score normalization further improves the efficiency and accuracy of data processing.
[0046] 2. The present invention constructs a hybrid model integrating association analysis algorithm, random forest classifier and long short-term memory network to calculate the threat value of network nodes. By the hybrid model, the advantages of different algorithms can be fully utilized to improve the accuracy and robustness of risk assessment. At the same time, the application of the dynamic weight allocation algorithm makes the risk assessment result more in line with the actual network situation and reduces the false alarm rate.
[0047] 3. The present invention adopts the fuzzy comprehensive evaluation method to divide the risk level, and dynamically adjusts the risk threshold interval according to historical attack data, making the risk assessment result more accurate and reliable. At the same time, the design of the real-time risk assessment module and the visual interactive interface further improves the practicability of the system and the user experience.
[0048] 4. The present invention realizes the continuous optimization of the model through the self-learning optimization mechanism, improves the response speed and accuracy of the system, and also reduces the operation and maintenance cost of the system. Brief Description of the Drawings
[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0050] Figure 1 It is a schematic diagram of the method of the present invention. Detailed Embodiments
[0051] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0052] Embodiment
[0053] In the prior art, there are the following disadvantages: data processing bottleneck, low efficiency of real-time cleaning and feature alignment; model rigidity, static risk assessment model cannot adapt to technological evolution; collaborative defense lag, delay in risk assessment and security device policy execution; low knowledge iteration efficiency, lack of self-learning mechanism, and manual intervention is required for rule base update.
[0054] Please refer to Figure 1 , the present invention provides a technical solution: a network security risk assessment method based on big data, including the following steps:
[0055] Step 1: Collect multi-source heterogeneous network security data, where the multi-source heterogeneous network security data includes network traffic logs, system vulnerability information, user behavior data, and external threat intelligence;
[0056] The content of the network traffic log is the packet information in the network, including source IP, destination IP, port number, protocol type, timestamp, packet size, etc.; the content of the system vulnerability information is the vulnerability information existing in the system, including vulnerability ID, description, severity level, affected range, repair suggestions, etc.; the content of the user behavior data is the operation behavior of the user, including login time, accessed files or resources, executed commands, permission changes, etc.; the content of the external threat intelligence is the threat information globally, including malicious IP addresses, domain names, attack gang activities, emerging vulnerabilities, attack techniques, etc.
[0057] Step 2: Perform real-time cleaning and normalization processing on the data to generate a standardized feature vector;
[0058] The specific steps of the data cleaning and normalization processing include:
[0059] Perform time series segmentation on the data using a sliding time window, and the window length satisfies the formula:
[0060]
[0061] where λ is the threat event occurrence rate, the average number of threat events occurring per unit time; ∈ is the preset error threshold, used to control the accuracy requirement of threat event detection; N total is the total amount of data, such as the total number of network traffic logs or the total data volume; T minis the minimum window time; to prevent the window from being too small, which may lead to waste of computing resources or instability of the model;
[0062] Calculate the length T of the sliding time window w , which is used for time series segmentation of network security data to balance the threat event capture granularity and computing efficiency;
[0063] Perform one-hot encoding on discrete features and perform Z-score normalization on continuous features;
[0064] By efficiently processing massive data, the real-time and accuracy problems existing in traditional methods when dealing with multi-source data are overcome. The application of technologies such as sliding time window and Z-score normalization further improves the efficiency and accuracy of data processing.
[0065] Step 3: Build a dynamic risk assessment model. Calculate the threat value of network nodes through a hybrid model that combines association analysis algorithm, random forest classifier, and long short-term memory network;
[0066] The calculation of the threat value in the dynamic risk assessment model satisfies the following formula:
[0067]
[0068] where w i is the dynamic weight of the i-th type of feature, reflecting the contribution of different features to the threat score; f(x i ) is the classification probability output by RFC, α is the time series correction coefficient, which is used to balance the influence of static features and time series features; LSTM(x t-k:t ) is the predicted outlier of historical time series data;
[0069] Calculate the threat score S threat , and combine the static feature classification probability and the time series data prediction result to dynamically evaluate the threat level.
[0070] The adjustment method of the dynamic weight w i includes:
[0071] The gradient descent optimization algorithm based on real-time threat intelligence feedback, and the objective function is:
[0072]
[0073] where is the predicted threat value of the j-th sample, is the actual attack event loss value of the j-th sample, β is the regularization parameter, which is used to control the sparsity of the weight w, and w is the weight vector, representing the weights of various features;
[0074] Update the weight every time interval Δt and constrain wi = 1;
[0075] The training process of the long short-term memory network includes:
[0076] The input layer is designed as a three-dimensional tensor X ∈ R b×t×d , where b is the batch size, t is the time step, and d is the feature dimension;
[0077] The loss function uses weighted cross-entropy:
[0078]
[0079] where c is the total number of attack types, γ C is the class imbalance correction factor, y c is the one-hot encoding of the true label, indicating whether the sample belongs to class c, and p c is the probability of class c predicted by the model;
[0080] By constructing a hybrid model that combines the fusion association analysis algorithm, random forest classifier, and long short-term memory network, it is used to calculate the threat value of network nodes. Through the hybrid model, the advantages of different algorithms can be fully utilized to improve the accuracy and robustness of risk assessment. At the same time, the application of the dynamic weight allocation algorithm makes the risk assessment result more in line with the actual network situation and reduces the false alarm rate.
[0081] Step 4: Based on the dynamic weight allocation of the threat value, optimize the risk assessment result, and output the risk level and protection suggestions;
[0082] The risk level is divided using the fuzzy comprehensive evaluation method, and the risk threshold interval is defined as:
[0083]
[0084] where S threat is the threat score, indicating the threat level of the current event; θ1 and θ2 are dynamically adjusted according to historical attack data; they are used to divide the boundaries of low, medium, and high risks;
[0085] The thresholds θ1 and θ2 are dynamically adjusted according to historical attack data, and the specific method is as follows:
[0086] Data collection: Collect historical attack events and their corresponding threat scores S threat ; Record the actual impact of each event.
[0087] Data analysis: Conduct statistical analysis on historical data to determine the distribution characteristics of low, medium, and high risks; for example, θ1 and θ2 can be determined through clustering analysis or the quantile method.
[0088] Threshold Update: Dynamically adjust θ1 and θ2 according to the analysis results to adapt to the latest threat patterns; for example: θ1 can be set as the upper limit of historical low-risk events; θ2 can be set as the upper limit of historical medium-risk events.
[0089] Real-time risk assessment is achieved through the following methods:
[0090] Deploy lightweight risk assessment agents to synchronize data to the central server periodically;
[0091] When detecting S threat mutation, trigger abnormal traffic mirroring and analysis, generate defense rules and distribute them to the firewall;
[0092] By adopting the fuzzy comprehensive evaluation method for risk level classification and dynamically adjusting the risk threshold range according to historical attack data, the risk assessment results are made more accurate and reliable. At the same time, the design of the real-time risk assessment module further improves the practicality and user experience of the system.
[0093] Step Five: By setting up a self-learning optimization mechanism, when the defense strategy is triggered, record the actual protection effect data and update the model parameters through the reinforcement learning algorithm;
[0094] Set up a reward function:
[0095]
[0096] Among them, the number of blocked attacks is the number of attacks successfully identified and blocked by the model, the number of false alarms is the number of normal behaviors misidentified as attacks by the model, the response delay is the time from the model detecting an attack to taking blocking measures, and η is the delay penalty coefficient; through the self-learning optimization mechanism, continuous optimization of the model is achieved, the response speed and accuracy of the system are improved, and the operation and maintenance cost of the system is also reduced.
[0097] This method integrates a visual interaction interface, which supports the following functions: displaying the threat tracing path in the form of a Sankey diagram; dynamically marking vulnerable nodes in the network topology based on the risk heat map; providing an interface for generating automated repair scripts; through the integrated visual interaction interface, the threat score S threat and the risk level classification rules can be presented to users in an intuitive and easy-to-operate manner. The interface design should focus on real-time, interactivity, and scalability, and at the same time combine the backend data processing capabilities to achieve efficient threat assessment and management.
[0098] This method is deployed in a distributed computing framework, including:
[0099] Implementing high-concurrency data collection using Apache Kafka; Using Apache Kafka can achieve high-concurrency and high-throughput data collection, effectively meet the real-time processing requirements of large-scale data streams, and ensure the real-time and accuracy of data;
[0100] Performing real-time feature calculation based on Spark Streaming; Performing real-time feature calculation based on Spark Streaming can achieve fast processing and feature extraction of data streams, providing strong data support for subsequent risk assessment;
[0101] Dynamically expanding risk assessment nodes through the Kubernetes container cluster; Dynamically expanding risk assessment nodes through the Kubernetes container cluster can flexibly adjust computing resources according to actual needs, improve the scalability and stability of the system, and ensure good performance in high-concurrency scenarios.
[0102] In summary, the present invention can efficiently process massive data, overcome the real-time and accuracy problems existing in traditional methods when dealing with multi-source data, and the application of technologies such as sliding time windows and Z-score normalization further improves the efficiency and accuracy of data processing; By constructing a hybrid model integrating association analysis algorithms, random forest classifiers, and long short-term memory networks to calculate the threat values of network nodes, the hybrid model can make full use of the advantages of different algorithms to improve the accuracy and robustness of risk assessment. At the same time, the application of the dynamic weight allocation algorithm makes the risk assessment results more in line with the actual network situation and reduces the false alarm rate; The fuzzy comprehensive evaluation method is used for risk level division, and the risk threshold interval is dynamically adjusted according to historical attack data, making the risk assessment results more accurate and reliable. At the same time, the design of the real-time risk assessment module and the visual interaction interface further improves the practicality and user experience of the system; Through the self-learning optimization mechanism, the continuous optimization of the model is achieved, the response speed and accuracy of the system are improved, and the operation and maintenance costs of the system are also reduced.
[0103] Those skilled in the art can understand that the features recited in the various embodiments and / or claims of the present invention can be combined or combined in various ways, even if such combinations or combinations are not explicitly recited in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features recited in the various embodiments and / or claims of the present invention can be combined and combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.
Claims
1. A network security risk assessment method based on big data, characterized in that: It includes the following steps: Step 1: Collect multi-source heterogeneous network security data; Step 2: Perform real-time cleaning and normalization processing on the data to generate standardized feature vectors; Step 3: Build a dynamic risk assessment model, and calculate the threat value of network nodes through a hybrid model that combines association analysis algorithms, random forest classifiers, and long short-term memory networks; Step 4: Optimize the risk assessment results based on the dynamic weight allocation of threat values, and output the risk level and protection suggestions; Step 5: By setting up a self-learning optimization mechanism, when the defense strategy is triggered, record the actual protection effect data, and update the model parameters through reinforcement learning algorithms.
2. The network security risk assessment method based on big data according to claim 1, wherein: In Step 2, the specific steps of the data cleaning and normalization processing include: Use a sliding time window to perform time series segmentation on the data, and the window length satisfies the formula: Among them, λ is the threat event occurrence rate, ∈ is the preset error threshold, N total is the total amount of data, T min is the minimum window time; Calculate the sliding time window length T w , which is used for time series segmentation of network security data to balance the capture granularity of threat events and computational efficiency; Perform one-hot encoding on discrete features, and perform Z-score normalization processing on continuous features.
3. A method for network security risk assessment based on big data according to claim 1, characterized in that: In Step 3, the calculation of the threat value in the dynamic risk assessment model satisfies the following formula: Among them, w i is the dynamic weight of the i-th type of feature, f(x i ) is the classification probability output by the RFC, α is the time series correction coefficient, and LSTM(x t-k:t ) is the predicted outlier of the historical time series data; Calculate the threat score S threat , and dynamically evaluate the threat level by combining the static feature classification probability and the time series data prediction result.
4. The network security risk assessment method based on big data according to claim 3, characterized in that: The adjustment method of the dynamic weight w i includes: Based on the gradient descent optimization algorithm with real-time threat intelligence feedback, the objective function is: Among them, is the predicted threat value of the j-th sample; is the actual attack event loss value of the j-th sample, β is the regularization parameter used to control the sparsity of the weight w, and w is the weight vector representing the weights of various features; Update the weights at every time interval Δt and constrain w i = 1.
5. A method for network security risk assessment based on big data according to claim 1, characterized in that: In Step 4, the risk level is divided using the fuzzy comprehensive evaluation method, and the defined risk threshold interval is: Among them, S threat is the threat score, and θ1 and θ2 are dynamically adjusted according to historical attack data.
6. The method for network security risk assessment based on big data according to claim 3, characterized in that: The real-time risk assessment is achieved through the following methods: By deploying lightweight risk assessment agents, synchronize data to the central server periodically; When S is detected threat a mutation occurs, triggering abnormal traffic mirroring and analysis, generating defense rules and sending them to the firewall.
7. A method for network security risk assessment based on big data according to claim 1, characterized in that: The method integrates a visual interaction interface, which supports the following functions: Display the threat traceability path in the form of a Sankey diagram; Dynamically mark vulnerable nodes in the network topology based on the risk heat map; Provide an interface for generating automated repair scripts.
8. A method for network security risk assessment based on big data according to claim 1, characterized in that: The training process of the long short-term memory network includes: The input layer is designed as a three-dimensional tensor X ∈ R b×t×d , where b is the batch size, t is the time step, and d is the feature dimension; The loss function uses weighted cross-entropy: Among them, c is the total number of attack types, and γ C is the class imbalance correction factor, and y c is the one-hot encoding of the true label, indicating whether the sample belongs to class c, and p c is the probability of class c predicted by the model.
9. A method for network security risk assessment based on big data according to claim 1, characterized in that: In Step 1, the multi-source heterogeneous network security data includes network traffic logs, system vulnerability information, user behavior data, and external threat intelligence.
10. A method for network security risk assessment based on big data according to claim 1, characterized in that: In Step 5, set the reward function: Among them, the number of blocked attacks is the number of attacks successfully identified and blocked by the model, the number of false alarms is the number of normal behaviors misidentified as attacks by the model, the response delay is the time from the model detecting an attack to taking blocking measures, and η is the delay penalty coefficient.
Citation Information
Cited By
Risk assessment method and system for information protection
CN120768704A
Method and system for risk assessment of information protection
CN120768704B
Automatic disposal method for security risk of supercomputing system
CN120910851A
Real-time network security threat early warning analysis method
CN121098621A
Server security risk identification method and system
CN121125263A