Network security detection method based on big data
By combining distributed sensor networks and adaptive learning models with deep neural networks and reinforcement learning algorithms, the shortcomings of existing network security detection systems in millisecond-level response are addressed, enabling rapid identification and defense against network attacks and improving the real-time performance and reliability of network security.
Patent Information
- Application Number
- CN202511155711.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-18
- Publication Date
- 2025-11-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing network security detection systems are unable to complete data analysis and response within milliseconds, resulting in the failure to detect network attacks in a timely manner and posing security risks.
Real-time data acquisition is achieved using a distributed sensor network, combined with incremental update processing, multi-level preprocessing, and an adaptive learning model. Attack features are extracted using deep neural networks and reinforcement learning algorithms to generate targeted defense strategies, and defense measures are implemented within milliseconds through a distributed security response system.
It enables rapid identification and millisecond-level response to network attacks, improving the real-time performance and reliability of network security protection and reducing system security risks.
Smart Images

Figure CN120979729A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of big data, and particularly to a network security detection method based on big data. BACKGROUND
[0002] The existing network security detection method mainly relies on traditional firewalls, intrusion detection systems and intrusion prevention systems. These methods analyze network traffic based on predefined rules or features to detect known attack patterns or viruses. However, with the continuous development of attack technology, network attack means is becoming more and more complex, and traditional security detection methods are facing great challenges. Especially in the face of advanced persistent threats, existing systems often have difficulty in effectively identifying and defending. In order to improve the accuracy and response speed of detection, more and more researchers begin to explore network security detection technology based on big data analysis.
[0003] Using big data analysis technology for network security detection can identify and predict potential security threats in real time through real-time monitoring and analysis of large-scale data. Compared with traditional methods, big data analysis can not only handle larger scale data, but also more accurately find security vulnerabilities in complex network environments. Through machine learning, data mining and other technologies, potential patterns of network security events can be mined from historical data, improving the ability of security event prediction and response. At present, many network security products have begun to try to combine big data technology with traditional security detection systems to improve their intelligent level and ability to cope with complex network environments. Big data technology can discover previously unnoticed security risks by deep learning on a large amount of network traffic, log files and other data.
[0004] Although big data has made some progress in network security detection, the existing detection methods based on big data still have some problems. The defect is that the existing technology cannot effectively solve the real-time problem in data processing. With the explosive growth of data, existing systems face a big bottleneck in real-time data stream processing. Network attacks often occur suddenly and covertly, and existing detection systems are difficult to complete data analysis and response within milliseconds. This leads to the security protection system failing to discover network attacks in time, causing great security risks to enterprises and users. SUMMARY
[0005] In view of the defects of the prior art, the present application provides a network security detection method based on big data, which solves the technical problem that the existing detection system is difficult to complete data analysis and response within milliseconds, leading to the problem that the security protection system fails to discover network attacks in time.
[0006] To achieve the above purpose, the present application is realized by the following technical scheme: a network security detection method based on big data, comprising: S1. Real-time collection of data streams in network environment, the data streams are obtained through a distributed sensor network, an incremental update processing method is adopted to ensure the timeliness and effectiveness of the data; S2. Multi-level preprocessing of the collected data streams, the multi-level preprocessing includes data denoising, feature selection and dimension reduction, the feature selection is filtered according to the correlation between historical data and real-time data through an automated algorithm, the most relevant features to network security threats are extracted to ensure the efficiency of the data in subsequent analysis; S3. Real-time analysis of the processed data streams based on an adaptive learning model, the adaptive learning model combines deep neural networks and reinforcement learning algorithms, automatically extracts complex attack features, dynamically adjusts learning parameters, and optimizes itself according to real-time data feedback to ensure rapid adaptation to new attacks; S4. After identifying potential attack behavior, automatically generate targeted defense strategies, execute the strategies through a distributed security response system, the distributed security response system adjusts defense measures in real time according to different security threat levels and attack types, and ensures millisecond-level response during defense execution, the multi-level preprocessing and real-time analysis allow second-level response, reducing security risks to the system; S5. Continuously optimize the adaptive learning model through a system feedback mechanism.
[0007] Preferably, the data streams include network traffic, user behavior, server logs and external attack source information.
[0008] Preferably, the automated algorithm used for filtering is a correlation weighted feature selection algorithm, which first calculates the correlation weight of each feature through the following steps, the formula is: wherein, represents the weight of feature , is the total number of time steps of the data set, is the value of the th feature at time step , is the value of the target variable at time step , represents the Pearson correlation coefficient of feature and target variable at time step , and are the variances of feature and target variable at time step , respectively.
[0009] Preferably, the adaptive learning model stores historical experience through an experience replay mechanism and uses a target network to stabilize Q-value updates. The formula for the adaptive learning model is: For parameters of the online network, Here are the parameters of the target network, and D is the empirical replay buffer. For at any time Observed environmental conditions For at any time The actions taken To perform the action The instant reward obtained afterward This refers to the state at the next moment after the action is performed. This is a discount factor used to balance current returns with future returns. To perform expectation calculations on samples in experience replay, ensuring the stability and efficiency of the training process.
[0010] Preferably, the defense strategy includes: traffic rate limiting, traffic filtering, network isolation, real-time virus scanning and isolation, behavior analysis, and dynamic access control.
[0011] Preferably, the distributed security response system includes multiple nodes, each node sharing real-time security threat information through a collaborative learning mechanism and automatically generating targeted defense strategies. The core architecture of the distributed security response system includes: Threat perception module: Collects network traffic, system logs and other security event data from various nodes in real time through sensors or monitoring systems; Attack identification module: Uses machine learning to analyze the collected data stream and identify potential attack patterns; Defense strategy generation module: Generates corresponding defense strategies based on attack patterns, and selects the optimal strategy using reinforcement learning or game theory models; Distributed response module: Distributes the defense strategy to each node and executes the corresponding defense operations.
[0012] Preferably, the system feedback mechanism is a positive feedback mechanism, specifically: in: For defensive strategies The weight, It is an adjustment factor that determines the degree to which a successful defense affects the strategy weights. It is an immediate reward, representing the reward for a successful defense strategy against an attack.
[0013] Preferably, the system feedback mechanism dynamically adjusts the feature selection and decision process of the model based on attack detection results, historical data and real-time monitoring data, and improves the detection ability of complex and hidden attacks.
[0014] The application provides a network security detection method based on big data. The network security detection method based on big data realizes real-time and efficient collection and cleaning of multi-source data streams in a network environment through the combination of a distributed sensor network and incremental update processing. The introduction of multi-level preprocessing and correlation weighted feature selection algorithms not only greatly reduces data redundancy and noise interference, but also automatically extracts key features highly related to security threats, ensuring the accuracy and computational efficiency of subsequent analysis. The adaptive learning model based on deep neural networks and reinforcement learning algorithms can dynamically adjust learning parameters and update stable Q values through experience replay and target networks, quickly respond to new and hidden attacks, and realize accurate identification and self-optimization of various complex attack features.
[0015] For the identified potential attack behavior, the system automatically generates and issues targeted defense strategies, and through a distributed security response system, implements multiple defense measures such as traffic restriction, isolation, scanning and dynamic access control within milliseconds, significantly improving the overall protection speed and reliability. The cooperative learning mechanism enables nodes to share threat information and jointly optimize strategies, enhancing the scalability and robustness of the system; the positive feedback mechanism continuously optimizes the model based on real-time monitoring and historical data, further improving the detection and defense capabilities of complex and hidden threats, ensuring the long-term stability and intelligent evolution of the network security protection system. BRIEF DESCRIPTION OF DRAWINGS
[0016] Figure 1 is a flowchart illustrating the implementation of the application. DETAILED DESCRIPTION
[0017] The technical solutions in the embodiments of the application will be described clearly and completely below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, rather than all the embodiments of the application. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of the application.
[0018] As shown in Figure 1 The embodiments of the application provide a network security detection method based on big data, which comprises: S1. Real-time collection of data streams in a network environment, the data streams are obtained through a distributed sensor network, and an incremental update processing method is used to ensure the timeliness and effectiveness of the data. The data streams include network traffic, user behavior, server logs and external attack source information.
[0019] S2. Multi-level preprocessing is performed on the collected data stream, including data denoising, feature selection, and dimension reduction. Feature selection is performed by an automated algorithm based on the correlation between historical data and real-time data to extract the most relevant features related to network security threats, ensuring efficient data analysis in subsequent analysis.
[0020] The automated algorithm used for screening is a correlation-weighted feature selection algorithm. This algorithm first calculates the correlation weight of each feature using the following steps, with the formula being: where, represents the weight of feature , is the total number of time steps in the data set, is the value of the th feature at time step , is the value of the target variable at time step , represents the Pearson correlation coefficient between feature and the target variable at time step , and are the variances of feature and the target variable at time step ,
[0021] The specific implementation is as follows: For example, assume we have collected data at 5 time points: Network latency (ms): 30, 40, 35, 45, 50.
[0022] Number of concurrent connections: 100, 80, 120, 90, 110.
[0023] Security threat label (0 for no threat, 1 for threat): 0, 1, 0, 1, 1.
[0024] Calculate the statistics of each indicator: The average network latency is 40 ms, with a standard deviation of approximately 7.07 ms.
[0025] The average number of concurrent connections is 100, with a standard deviation of approximately 14.14.
[0026] The average threat label is 0.6, with a standard deviation of approximately 0.49.
[0027] Calculate the correlation: The correlation between network latency and threat labeling is about 0.866.
[0028] The correlation between concurrent connection number and threat labeling is about -0.577.
[0029] Obtain feature weights: The weight of network latency is 0.866 after taking the absolute value.
[0030] The weight of concurrent connection number is 0.577 after taking the absolute value.
[0031] Screening according to the weight threshold value 0.6: Network latency (0.866 >= 0.6) is retained.
[0032] Concurrent connection number (0.577 < 0.6) is eliminated.
[0033] Through the above steps, without using formulas and tables, the correlation weighted feature selection is completed by directly substituting the data, and finally only the network latency feature most related to security threats is retained.
[0034] S3. Real-time analysis of the processed data stream based on an adaptive learning model, which combines deep neural networks and reinforcement learning algorithms to automatically extract complex attack features, dynamically adjust learning parameters, and self-optimize based on real-time data feedback to ensure rapid adaptation to new attacks.
[0035] The adaptive learning model stores historical experience through an experience replay mechanism and uses a target network to stabilize Q value updates. The formula of the adaptive learning model is: is the parameter of the online network, is the parameter of the target network, D is the experience replay buffer, is the state of the environment observed at time is the action taken at time is the state of the environment observed at time is the action taken at time is the immediate reward obtained after executing action is the state of the environment observed at time is the discount factor used to balance current and future rewards, is the expected operation on samples in the experience replay, ensuring the stability and efficiency of the training process.
[0036] Specific implementation is as follows: Preparation: The discount factor is 0.9.
[0037] The experience replay buffer holds multiple interactions, and we’ll pick three at random to demonstrate.
[0038] First experience: Interaction content: State A, perform “isolate traffic” action, receive immediate reward 1, enter state B.
[0039] The target network computes the maximum Q-value for state B to be 2.0.
[0040] So the target value for this update is 1 plus 0.9 times 2.0, which is 2.8.
[0041] In the online network, the current Q-value for state A and this action is 2.5.
[0042] The difference is 0.3, and after the network parameters are updated, this Q-value rises to approximately 2.65.
[0043] Second experience: Interaction content: State C, perform “throttle” action, receive reward 0, enter state D.
[0044] The target network computes the maximum Q-value for state D to be 1.5.
[0045] The target value for this update is 0 plus 0.9 times 1.5, which is 1.35.
[0046] The current Q-value in the online network is 1.2.
[0047] The difference is 0.15, and after the update, this Q-value becomes approximately 1.28.
[0048] Third experience: Interaction content: State E, perform “allow traffic” action, receive negative reward –1, enter state F.
[0049] The target network computes the maximum Q-value for state F to be 0.5.
[0050] The target value for this update is –1 plus 0.9 times 0.5, which is –0.55.
[0051] The current Q-value in the online network is –0.4.
[0052] The difference is –0.15, and after the update, this Q-value drops to approximately –0.44.
[0053] Through the above three examples with specific data, it can be seen that the model first samples from the experience cache, then calculates the maximum Q value of the next state using the target network, combines it with the immediate reward to form the learning target for this time, then compares it with the current Q value of the online network, and finally adjusts the parameters of the online network to make the Q value converge gradually and remain stable, so as to quickly adapt to new attacks.
[0054] S4. After identifying potential attack behavior, automatically generate targeted defense strategies, and execute the strategies through a distributed security response system. The distributed security response system adjusts defense measures in real time according to different security threat levels and attack types, and ensures millisecond-level response during defense execution. Multi-level preprocessing and real-time analysis allow for second-level response, reducing security risks to the system. Defense strategies include: traffic rate limiting, traffic filtering, network isolation, real-time virus scanning and isolation, behavior analysis, and dynamic access control.
[0055] The distributed security response system includes multiple nodes, each of which shares real-time security threat information through a cooperative learning mechanism and automatically generates targeted defense strategies. The core architecture of the distributed security response system includes: Threat perception module: Collect network traffic, system logs, and other security event data from each node in real time through sensors or monitoring systems; Attack identification module: Analyze the collected data stream using machine learning-based methods to identify potential attack patterns; Defense strategy generation module: Generate corresponding defense strategies based on attack patterns, and use reinforcement learning or game theory models to select the optimal strategy; Distributed response module: Distribute defense strategies to each node and execute corresponding defense operations.
[0056] S5. Continuously optimize the adaptive learning model through a system feedback mechanism. The system feedback mechanism is a positive feedback mechanism, specifically: Where: is the weight of the defense strategy is the adjustment factor, which determines the degree of influence of successful defense on the strategy weight is the immediate reward, representing the reward brought by the successful defense of the attack.
[0057] The system feedback mechanism dynamically adjusts the feature selection and decision-making process of the model based on attack detection results, historical data, and real-time monitoring data, improving the detection capability of complex and hidden attacks.
[0058] While embodiments of the application have been shown and described, it is to be understood that the embodiments described are merely exemplary of the principles and application of the present application. Numerous modifications and adaptions can be effected without departing from the spirit and scope of the present application, which is not limited to the exact construction and arrangement described. It is intended, therefore, to cover all modifications and adaptions that fall within the scope of the claims and their equivalents.
Claims
1. A network security detection method based on big data, characterized in that, include: S1. Real-time acquisition of data streams in the network environment, wherein the data streams are acquired through a distributed sensor network; S2. Perform multi-level preprocessing on the collected data stream. The multi-level preprocessing includes data denoising, feature selection, and dimensionality reduction. The feature selection is performed by an automated algorithm to filter based on the correlation between historical data and real-time data, and to extract the features most relevant to network security threats. S3. Real-time analysis of the processed data stream is performed based on an adaptive learning model, which combines deep neural networks and reinforcement learning algorithms to self-optimize based on feedback from real-time data. S4. After identifying potential attack behaviors, a targeted defense strategy is automatically generated and executed through a distributed security response system. The distributed security response system adjusts defense measures in real time according to different security threat levels and attack types, and ensures millisecond-level response during the defense execution phase. The multi-layered preprocessing and real-time analysis allow for second-level response. S5. Continuously optimize the adaptive learning model through system feedback mechanisms.
2. The network security detection method based on big data according to claim 1, characterized in that: The data stream includes network traffic, user behavior, server logs, and information about external attack sources.
3. The network security detection method based on big data according to claim 1, characterized in that: The automated algorithm used for screening is a relevance-weighted feature selection algorithm, and the formula is: in, Representation of features The weight, It is the total number of time steps in the dataset. It is the first Each feature at time step The value of time, The target variable at time step The value of time, Representation of features and target variable At time step Pearson correlation coefficient at time and Features and target variable At time step Variance over time.
4. The network security detection method based on big data according to claim 1, characterized in that: The adaptive learning model stores historical experience through an experience replay mechanism and uses a target network to stabilize Q-value updates. The formula for the adaptive learning model is: For parameters of the online network, Here are the parameters of the target network, and D is the empirical replay buffer. For at any time Observed environmental conditions For at any time The actions taken To perform the action The instant reward obtained afterward This refers to the state at the next moment after the action is performed. This is a discount factor used to balance current returns with future returns. This is for the expectation operation of samples in the experience replay.
5. The network security detection method based on big data according to claim 1, characterized in that: The defense strategies include: traffic rate limiting, traffic filtering, network isolation, real-time virus scanning and isolation, behavioral analysis, and dynamic access control.
6. The network security detection method based on big data according to claim 5, characterized in that: The distributed security response system comprises multiple nodes, and its core architecture includes: Threat perception module: Collects network traffic, system logs and other security event data from various nodes in real time through sensors or monitoring systems; Attack identification module: Uses machine learning to analyze the collected data stream and identify potential attack patterns; Defense strategy generation module: Generates corresponding defense strategies based on attack patterns, and selects the optimal strategy using reinforcement learning or game theory models; Distributed response module: Distributes the defense strategy to each node and executes the corresponding defense operations.
7. The network security detection method based on big data according to claim 1, characterized in that: The system feedback mechanism is a positive feedback mechanism, specifically: in: For defensive strategies The weight, It is an adjustment factor. It's an instant reward.
8. The network security detection method based on big data according to claim 1, characterized in that: The system feedback mechanism dynamically adjusts the model's feature selection and decision-making process based on attack detection results, historical data, and real-time monitoring data, thereby improving the detection capability for complex and covert attacks.