Network security situation awareness method based on JS divergence
Through multi-source data acquisition, JS divergence abnormal detection and network behavior completion model, combined with multi-dimensional threat classification algorithm, data acquisition, detection and storage problems in network security situation awareness are solved, and comprehensive, accurate judgment and efficient management of network security situations are achieved.
Patent Information
- Application Number
- CN202510701734.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-08-05
AI Technical Summary
The existing network security technology has shortcomings in data collection, abnormal detection, data processing and threat classification, resulting in insufficient comprehensive and accurate judgment of network security situations, and lack of flexibility and security in data storage and management.
The multi-source data acquisition module is used to obtain network security parameters, and the detection threshold is dynamically adjusted using an abnormality detection model based on JS divergence, data completion is performed in combination with the network behavior completion model, and detailed storage management is carried out through a multi-dimensional threat classification algorithm.
It realizes a comprehensive and accurate judgment of the network security situation, reduces the rate of missed and false alarms, ensures the integrity and security of data, and improves data storage and management efficiency.
Smart Images

Figure CN120434017A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security technology, and in particular to a network security situation awareness method based on JS divergence. Background Art
[0002] In today's digital age, networks have become deeply integrated into every aspect of social life. Business operations, government management, and everyday personal activities all rely heavily on the security and stability of the network environment. With the continuous expansion of networks and the increasing complexity of application scenarios, network security faces unprecedented challenges. With the continuous emergence of various new cyberattack methods, traditional network security protection strategies are gradually exposing their shortcomings.
[0003] From a data collection perspective, previous methods often employed limited methods, making it difficult to comprehensively capture various security parameters within a network environment. These methods focused solely on a few key indicators, neglecting comprehensive data collection on traffic characteristics, protocol types, access logs, and attack behavior. This resulted in a lack of comprehensive and accurate assessment of network security trends. For example, some simple monitoring tools only captured basic network traffic values and were unable to deeply analyze protocol anomalies within the traffic, making it difficult to detect potential security threats in a timely manner.
[0004] Traditional anomaly detection methods are mostly based on fixed thresholds or simple rule matching. This approach cannot adapt to dynamic changes in the network environment and is prone to missed or false positives if network traffic patterns or attack methods change. For example, if a sudden increase in network traffic does not indicate an attack, a detection system based on fixed thresholds may falsely issue an alert. Furthermore, in the face of new attack methods, a system may be unable to identify the true threat due to its lack of ability to learn new behavior patterns, rendering network security defenses ineffective.
[0005] Existing methods for addressing missing and conflicting network data are also inadequate. Missing data leads to incomplete information, hindering accurate assessments of network security posture; conflicting data can interfere with analytical results and bias security decisions. Some data repair algorithms simply use averages or historical data for infill, failing to fully consider network topology and time series characteristics. This makes it difficult to accurately complete data and restore the true network security status.
[0006] Traditional technologies also have significant shortcomings in threat classification and data storage. Threat classification is not detailed or scientific enough to effectively differentiate between different types and severity of threats, hindering security managers from quickly identifying and addressing critical issues. Data storage methods also lack flexibility and security, failing to properly store and encrypt data based on its importance and sensitivity, creating the risk of data leaks. For example, critical enterprise network security data may be stored in the same area as general data without strict access control. This makes sensitive information vulnerable to theft in the event of an attack. Summary of the Invention
[0007] The purpose of the present invention is to provide a network security situation awareness method based on JS divergence to solve the problems raised in the above background technology.
[0008] To achieve the above object, the present invention provides the following technical solution: a network security situation awareness method based on JS divergence, the method comprising:
[0009] Synchronously acquiring security parameters of the target network environment through a multi-source data acquisition module, wherein the multi-source data acquisition module includes a monitoring unit for at least one network security element, and the security parameters include traffic characteristics, protocol types, access logs, and attack behaviors;
[0010] Inputting the security parameters into a preset anomaly detection model based on JS divergence to identify abnormal data distribution in the security parameters to obtain preprocessed security data, wherein the anomaly detection model dynamically adjusts the detection threshold based on the distribution difference of historical security data;
[0011] Inputting the pre-processed security data into a preset network behavior completion model to perform pattern completion on missing or conflicting data and generate a continuous and unified security situation dataset, wherein the network behavior completion model determines the completion weights based on the topological association characteristics and temporal behavior characteristics of the target network;
[0012] The security situation data set is classified into threat levels using a preset multi-dimensional threat classification algorithm and stored to form integrated perception data.
[0013] Preferably, the step of constructing the anomaly detection model includes:
[0014] Acquire a historical security data set, wherein each data in the historical security data set is annotated with an anomaly type and distribution difference;
[0015] Divide the training subsets based on the anomaly type and distribution difference, each training subset corresponding to an attack scenario; use the training subsets to train the initial detection model in parallel until the anomaly recognition accuracy of the initial detection model for each attack scenario is greater than or equal to a preset first threshold, and then stop training to obtain an intermediate detection model;
[0016] The historical safety data set is input into the intermediate detection model, and it is verified whether the anomaly recognition result output by the intermediate detection model meets the preset error range; if so, the intermediate detection model is determined as the anomaly detection model.
[0017] Preferably, synchronously acquiring the security parameters of the target network environment through the multi-source data acquisition module includes collecting monitoring data of one of the security elements in the following manner:
[0018] Establishing a communication connection with a target monitoring unit, wherein the target monitoring unit is deployed at a key node of the target network environment;
[0019] Periodically reading the real-time data stream of the target monitoring unit according to a preset sampling frequency, and marking a collection timestamp based on the timing characteristics of the real-time data stream;
[0020] According to the logical topology relationship of the target network, the real-time data streams of different nodes at the same timestamp are logically aligned to form an associated security parameter set.
[0021] Preferably, inputting the security parameter into a preset anomaly detection model based on JS divergence includes:
[0022] Extracting a distribution difference segment from the security parameter, wherein the distribution difference segment is a data segment in which a JS divergence calculation result exceeds a preset difference threshold within a continuous time window;
[0023] Generate an anomaly assessment index based on the duration and divergence amplitude of the distribution difference segment;
[0024] The corresponding detection algorithm is dynamically selected according to the anomaly assessment index, wherein the local density detection algorithm is used for short-term high-amplitude differences, and the sliding window detection algorithm is used for long-term low-amplitude differences.
[0025] Preferably, the method further comprises:
[0026] After identifying the abnormal data distribution, performing a data consistency check on the pre-processed security data;
[0027] If the verification finds that the data conflict rate exceeds the preset second threshold, the network behavior completion model is triggered to perform priority completion on the conflict data, wherein the high-priority conflict data is a data segment with a continuous conflict duration exceeding the preset duration.
[0028] Preferably, the network behavior completion model includes the following completion steps:
[0029] Constructing a logical topology model based on the distribution of key nodes of the target network, wherein each topology node corresponds to the logical position of a monitoring unit;
[0030] Calculate the logical completion weight based on the security parameter correlation of adjacent topological nodes;
[0031] Combined with the temporal behavior trend of the security parameters, the missing or conflicting topological nodes are completed in a spatiotemporal joint manner.
[0032] Preferably, the method further comprises:
[0033] After the completion is completed, the completion result is verified for logical consistency, where the verification method includes comparing the deviation between the completed data and the actual data of the adjacent nodes;
[0034] If the deviation exceeds a preset third threshold, the logic completion weight is readjusted and completion is iterated until the deviation is less than the third threshold.
[0035] Preferably, classifying the security situation data set by threat level using a preset multi-dimensional threat classification algorithm includes:
[0036] Classify the first-level threat labels according to the security parameter type, wherein the first-level threat labels include traffic anomaly, protocol violation, and attack behavior;
[0037] Under each level of threat label, the secondary threat sub-labels are further divided based on the severity of the threat;
[0038] The classified security data is stored in different partitions of the distributed storage system according to the label level.
[0039] Preferably, the method further comprises:
[0040] Configure encryption keys for threat tags based on pre-set data security levels;
[0041] Upon receiving a data access request, verify whether the key provided by the requester matches the encryption key of the target threat tag;
[0042] If a match is found, the data access interface for the corresponding threat tag is opened.
[0043] Preferably, the method further comprises optimizing the network behavior completion model in the following manner:
[0044] Counting the distribution of completion errors of the security situation dataset at different time periods;
[0045] Determining parameter adjustments for the completion model based on the error distribution, wherein a logic completion weight is increased during a high error period and a time completion weight is increased during a low error period;
[0046] The network behavior completion model is iteratively optimized based on the parameter adjustment amount until a completion error rate is less than a preset fourth threshold.
[0047] Compared with the prior art, the present invention has the following beneficial effects:
[0048] The network security situation awareness method based on JS divergence proposed in the present invention shows significant advantages in many aspects and provides a more reliable and efficient solution for network security protection. In the data collection stage, the multi-source data collection module covers a variety of network security element monitoring units, and can simultaneously obtain rich security parameters such as traffic characteristics, protocol types, access logs and attack behaviors. This comprehensive data collection method is like building a three-dimensional information perception system for network security situation analysis. For example, in a complex enterprise network environment, it is possible to not only monitor the size and direction of network traffic, but also conduct in-depth analysis of whether the protocol type is compliant, as well as abnormal access behaviors in access logs. This greatly enriches the data source for network security analysis, makes the judgment of network security situation more accurate and comprehensive, and avoids misjudgment of security risks due to missing data.
[0049] The anomaly detection model built on JS divergence dynamically adjusts detection thresholds based on the distribution differences of historical security data, exhibiting strong adaptive capabilities. In real-world network environments, network traffic and attack methods are constantly changing, and traditional fixed-threshold detection methods are prone to missed alerts or false positives. However, the anomaly detection model of this invention can learn and adapt to these changes in real time. Through in-depth analysis of historical data, it accurately captures the distribution differences between normal and abnormal data. For example, in response to new DDoS attacks, the model can dynamically adjust detection thresholds based on unusual changes in traffic distribution during an attack, accurately identifying attack behaviors. This effectively improves anomaly detection accuracy, reduces missed alerts and false positives, and provides a more reliable first line of defense for network security.
[0050] The network behavior completion model uses the topological correlation characteristics and temporal behavior characteristics of the target network to determine the completion weights and accurately complete the missing or conflicting data. This feature ensures the integrity and accuracy of the security situation data set. During network data transmission, data missing and conflict problems often occur due to network failures, equipment failures, and other reasons. Traditional data completion methods are often ineffective and cannot fully consider the actual situation of the network. The completion model of the present invention can reasonably infer the value of missing or conflicting data by analyzing the correlation between nodes in the network topology and the trend of data changes over time. For example, in a campus network with a complex topology, when the data of a node is missing, the model can accurately complete the missing data based on the data of adjacent nodes and the traffic change pattern in the time series, providing complete and reliable data support for subsequent security situation analysis.
[0051] The multi-dimensional threat classification algorithm, combined with a distributed storage system, enables efficient classification, storage, and management of security data. First-level threat tags are classified based on security parameter type, and second-level threat subtags are further subdivided based on threat severity. This detailed classification method enables security managers to quickly locate and identify threats of different types and severity. At the same time, the classified data is stored in different partitions of the distributed storage system, and encryption keys are configured according to the data security level, greatly improving data security and access management efficiency. For example, when a security incident occurs, managers can quickly obtain relevant data from the corresponding partition, quickly determine the nature and severity of the threat based on the threat tag, and take targeted countermeasures in a timely manner. Furthermore, a strict encryption key verification mechanism effectively prevents the risk of data leakage and ensures the confidentiality and integrity of network security data. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 This is a working principle diagram of the network security situation awareness method based on JS divergence according to the present invention;
[0053] Figure 2 Flowchart for inputting anomaly detection models into security parameters;
[0054] Figure 3 Flowchart for data consistency verification and completion;
[0055] Figure 4 Flowchart of the multi-dimensional threat classification algorithm. DETAILED DESCRIPTION
[0056] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0057] See also Figures 1-4 The present invention provides a network security situation awareness method based on JS divergence, and the specific implementation steps are as follows:
[0058] The security parameters of the target network environment are synchronously acquired through a multi-source data acquisition module. The multi-source data acquisition module contains a monitoring unit for at least one network security element, and the collected security parameters cover aspects such as traffic characteristics, protocol types, access logs, and attack behaviors. The various monitoring units in the multi-source data acquisition module work together to comprehensively monitor each key node and link of the network to ensure that the acquired data fully and accurately reflects the operating status of the network. For example, in an enterprise network, the traffic characteristic monitoring unit can monitor the size of data traffic between different areas in the network, the changing trend of traffic, etc. in real time; the protocol type monitoring unit is responsible for identifying the various protocols being used in the network and determining whether there are any abnormal protocol usage; the access log monitoring unit records the access behavior of users in the network, including information such as access time and accessed resources; and the attack behavior monitoring unit attempts to detect whether there are signs of attack behavior in the network through specific algorithms and rules.
[0059] The acquired security parameters are fed into a pre-set JS divergence-based anomaly detection model. This anomaly detection model can identify abnormal data distributions within the security parameters, thereby generating pre-processed security data. The anomaly detection model dynamically adjusts detection thresholds based on the distribution differences of historical security data. This allows the model to better adapt to changes in the network environment and improve detection accuracy. In real-world applications, the network environment is constantly changing. For example, network traffic can fluctuate significantly over time, and new network applications and protocols may emerge continuously. The anomaly detection model analyzes historical security data to learn the distribution patterns of data under normal circumstances. When the distribution of current data differs from that of historical data by more than a certain threshold, it identifies abnormal data. Furthermore, the model dynamically adjusts the detection threshold based on real-time changes in the network environment to avoid false positives and missed detections.
[0060] The pre-processed security data is fed into a pre-set network behavior completion model. This model completes patterns in missing or conflicting data and generates a continuous and unified security situation dataset. The network behavior completion model determines completion weights based on the topological correlation and temporal behavior characteristics of the target network. In a network, nodes have certain topological relationships, and network behavior exhibits a certain degree of continuity and regularity over time. The network behavior completion model leverages these characteristics to calculate the degree of correlation between nodes and the temporal trends of data. This determines the weights of various factors when completing missing or conflicting data, thereby achieving more accurate completion.
[0061] Using a pre-defined multidimensional threat classification algorithm, security situation data sets are classified and stored by threat level, generating integrated perception data. This algorithm analyzes and categorizes security data from multiple perspectives, differentiating threats by type and severity to facilitate subsequent management and processing. This classified security data is stored in a dedicated storage system for easy access and utilization, providing network security managers with intuitive and comprehensive information on the network security situation.
[0062] The present invention will be further described below in conjunction with Examples 1 to 6:
[0063] Example 1:
[0064] When building an anomaly detection model, we first obtain a historical security dataset, in which each data point is annotated with the anomaly type and distribution difference. This annotation information is obtained through analysis and research of a large number of historical network security incidents and is highly accurate and reliable. For example, a large internet service provider's network collected network security data for the past year, including normal network traffic data and network data during different types of attacks. This data is analyzed in detail to determine the anomaly type corresponding to each data point, such as a DDoS attack or SQL injection attack, and the degree of difference between the data and the normal data distribution is calculated.
[0065] Training subsets are divided based on anomaly type and distribution differences, with each subset corresponding to a specific attack scenario. This allows for separate model training for different attack scenarios, improving the model's detection capabilities. For example, using the DDoS attack scenario as an example, all DDoS attack-related data is filtered from historical security datasets and assembled into a training subset. The data in this training subset exhibits typical DDoS attack characteristics, such as large volumes of traffic requests and specific traffic patterns.
[0066] The initial detection model is trained in parallel using the training subsets until the initial detection model's anomaly recognition accuracy for each attack scenario is greater than or equal to the preset first threshold. The training is stopped to obtain an intermediate detection model. During the training process, the use of parallel computing can greatly improve the training efficiency. For example, multiple high-performance servers are used to train different training subsets at the same time, and each server is responsible for the training task of one attack scenario. JS divergence (Jensen-Shannon Divergence) is an indicator that measures the difference between two probability distributions. In this model, it is used to judge the degree of difference between the current data distribution and the normal data distribution. Assume that the probability distribution of normal data is , the probability distribution of the current data to be detected is , the calculation formula of JS divergence is:
[0067]
[0068] in, , Represents KL divergence (Kullback-Leibler Divergence), which is calculated as follows:
[0069]
[0070] In the formula, and Represents the probability distribution and In the By continuously adjusting the parameters of the initial detection model, the model's anomaly recognition accuracy for each attack scenario is continuously improved. When the preset first threshold (such as 95%) is reached, the model training effect is considered good, and training is stopped to obtain the intermediate detection model.
[0071] Input the historical security dataset into the intermediate detection model and verify whether the anomaly identification results output by the intermediate detection model meet the preset error range. If so, the intermediate detection model is determined to be the anomaly detection model. The preset error range is set based on actual application requirements and experience; for example, the allowable false positive rate does not exceed 5%. After inputting the historical security dataset into the intermediate detection model, the model outputs the anomaly identification results for each piece of data. These results are compared with the anomaly types marked in the historical security dataset to calculate the false positive rate. If the false positive rate is within the preset error range, it indicates that the intermediate detection model can accurately identify anomaly data and can be determined as the final anomaly detection model. If the false positive rate exceeds the preset error range, the model parameters need to be readjusted and training continues.
[0072] Example 2:
[0073] When using the multi-source data acquisition module to synchronously acquire security parameters of the target network environment, taking the monitoring data of one security factor as an example, a communication connection with the target monitoring unit is first established. The target monitoring unit is deployed at key nodes in the target network environment, which are crucial for the normal operation and security monitoring of the network. For example, in a financial institution's network, the nodes where core servers, firewalls, and other devices are located are key nodes. Target monitoring units deployed at these locations can obtain critical data in real time.
[0074] The real-time data stream of the target monitoring unit is read periodically at the preset sampling frequency, and the collection timestamp is marked based on the timing characteristics of the real-time data stream. The preset sampling frequency is determined according to the scale of the network, the speed of data change, and the actual application requirements. For example, for networks with faster changes in network traffic, a higher sampling frequency can be set, such as 10 sampling times per second; for networks with relatively slow data changes, the sampling frequency can be appropriately reduced, such as 1 sampling every 10 seconds. The marking of the collection timestamp can accurately record the collection time of the data, which facilitates subsequent time series analysis of the data. For example, if a set of network traffic data is collected at 10:00:00 on October 1, 2024, the corresponding timestamp will be marked for this set of data.
[0075] Based on the target network's logical topology, real-time data streams from different nodes at the same timestamp are logically aligned to form a set of associated security parameters. The target network's logical topology describes the connection methods and data flow between nodes in the network. By analyzing logical topology relationships, it's possible to determine which nodes' data are associated. For example, in an enterprise campus network, data exchange occurs between network nodes in the office area and those in the data center. Based on the logical topology relationships, the real-time data streams from the office area nodes and the data center nodes at the same timestamp can be logically aligned and merged into a set of associated security parameters. This set contains security parameter information for multiple nodes at the same moment, providing a more comprehensive data foundation for subsequent data analysis.
[0076] Example 3:
[0077] When inputting security parameters into the preset anomaly detection model based on JS divergence, the distribution difference fragments in the security parameters are first extracted. The distribution difference fragment refers to the data segment in which the JS divergence calculation result exceeds the preset difference threshold within the continuous time window. The size of the continuous time window and the preset difference threshold are determined based on the historical data of the network and the actual security requirements. For example, the continuous time window is set to 10 seconds and the preset difference threshold is 0.5. In actual calculations, 10 seconds is used as a time window, and this window is continuously slid to calculate the JS divergence of the data in each window. Assume that the probability distribution of normal data is , the probability distribution of the data in the current window is , calculated according to the JS divergence calculation formula introduced above If the calculated result is greater than 0.5, then the 10-second data segment is a distribution difference segment.
[0078] Anomaly assessment indicators are generated based on the duration and divergence amplitude of the distribution difference segment. The duration reflects the duration of the anomaly, and the divergence amplitude reflects the severity of the anomaly. For example, if a distribution difference segment lasts for 30 seconds and its divergence amplitude is 0.8, then a comprehensive anomaly assessment indicator can be generated based on these two values. The specific generation method can be a weighted summation method, where the weight of the duration is , the weight of the divergence amplitude is , anomaly evaluation indicators The calculation formula is:
[0079]
[0080] in, and The value of is determined through experiments and analysis according to the actual situation, for example , .
[0081] The corresponding detection algorithm is dynamically selected based on the anomaly assessment metric. For short-term high-amplitude differences, the local density detection algorithm is used, while for long-term low-amplitude differences, the sliding window detection algorithm is used. For short-term high-amplitude differences, where the distribution difference segment is short-lived but has a high dispersion amplitude, the local density detection algorithm can more accurately detect anomalies. The core concept of the local density detection algorithm is to calculate the density around a data point. The density of anomalies is typically significantly different from that of normal points. For long-term low-amplitude differences, the sliding window detection algorithm is more suitable. The sliding window detection algorithm continuously slides a fixed-size window, performing statistical analysis on the data within the window, and determining whether anomalies exist. For example, when monitoring network traffic, if a brief traffic spike (a short-term high-amplitude difference) is detected, the local density detection algorithm can quickly locate the possible source of the anomaly. For a slow increase in traffic over a period of time (a long-term low-amplitude difference), the sliding window detection algorithm can better track this trend and promptly identify potential security threats.
[0082] After identifying the abnormal data distribution, the pre-processed security data is subjected to data consistency verification. The purpose of data consistency verification is to ensure the accuracy and integrity of the data. For example, check whether there are contradictions in the fields in the data, whether the value range of the data is reasonable, etc. If the verification finds that the data conflict rate exceeds the preset second threshold, the network behavior completion model is triggered to prioritize the conflict data, where high-priority conflict data is a data segment with a continuous conflict duration exceeding the preset duration. The preset second threshold and the preset duration are set according to the stability of the network and the data quality requirements. For example, the preset second threshold is 10%, and the preset duration is 60 seconds. If the data conflict rate exceeds 10%, and there are data segments with a continuous conflict duration exceeding 60 seconds, these data segments will be treated as high-priority conflict data and will be processed by the network behavior completion model first.
[0083] Example 4:
[0084] When completing the network behavior completion model, it first constructs a logical topology model based on the distribution of key nodes in the target network. Each topology node corresponds to the logical location of a monitoring unit. When constructing the logical topology model, it is necessary to consider the physical connections and data transmission relationships between nodes in the network. For example, in a city's traffic monitoring network, the locations of each surveillance camera are key nodes. Based on the distribution of these nodes and the data transmission paths between them, a logical topology model is constructed. Each topology node in the model represents the logical location of a surveillance camera, providing a clear representation of the network structure.
[0085] The logical completion weight is calculated based on the correlation of security parameters of adjacent topological nodes. There is usually a certain correlation between adjacent topological nodes. For example, the change of network traffic of one node may affect the traffic of adjacent nodes. By analyzing the correlation between the security parameters of adjacent nodes, the logical completion weight can be determined. and nodes The safety parameters are and , the correlation between them can be measured by Pearson correlation coefficient To measure, the calculation formula is:
[0086]
[0087] in, is the number of data samples, and Node and nodes No. Sample values, and Node and nodes The sample mean of . Logical completion weight According to the Pearson correlation coefficient Make adjustments, such as The higher the correlation, the greater the logical completion weight, which means that when completing data, the data of the adjacent node is more important to the completion of the current node data.
[0088] Combined with the temporal behavior trend of security parameters, the missing or conflicting topological nodes are completed in a spatiotemporal manner. Security parameters usually have a certain trend of change over time. For example, network traffic fluctuates regularly at different times of the day. By analyzing the temporal behavior trend of security parameters, the value of missing or conflicting data can be better predicted. Taking network traffic as an example, a time series model such as the ARIMA model (Autoregressive Integrated Moving Average Model) can be established based on historical traffic data. Assume for The network traffic value at time t, the general form of the ARIMA model is:
[0089]
[0090] in, is the autoregressive order, is the moving average order, and are model parameters, is a white noise sequence. This model can predict future traffic values based on past traffic data. Combined with logical completion weights, it can perform spatiotemporal joint completion of data for missing or conflicting topological nodes. For example, when traffic data for a node is missing, a possible value is first predicted based on the time series model. This predicted value is then adjusted based on the logical completion weights of adjacent nodes to achieve a more accurate completion result.
[0091] After the completion is completed, the completion result is verified for logical consistency. The verification method includes comparing the deviation between the completed data and the actual data of the neighboring nodes. If the deviation exceeds the preset third threshold, the logical completion weight is readjusted and the completion is iterated until the deviation is less than the third threshold. The preset third threshold is set according to the accuracy requirements of the network data, for example, it is set to 5%. If the deviation between the completed data and the actual data of the neighboring nodes exceeds 5%, it means that the completion result may be inaccurate, and it is necessary to readjust the logical completion weight and perform the completion operation again until the deviation is less than 5%. Through this iterative method, the accuracy of the completion result can be continuously improved.
[0092] Example 5:
[0093] When using a pre-defined multi-dimensional threat classification algorithm to categorize the security situation dataset into threat levels, first-level threat labels are assigned based on security parameter types. These labels include traffic anomaly, protocol violation, and attack behavior. This classification is based on common types of network security threats. For example, the traffic anomaly category primarily targets unusual fluctuations, increases, or decreases in network traffic; the protocol violation category identifies the use of illegal or unusual protocols; and the attack behavior category encompasses various known attack methods, such as DDoS attacks and hacker intrusions.
[0094] Each threat level is further divided into sub-level threat sub-tags based on threat severity. For example, traffic anomaly can be categorized into three sub-level threat sub-tags: mild traffic anomaly, moderate traffic anomaly, and severe traffic anomaly. This classification can be based on factors such as the severity and duration of the traffic anomaly, as well as its impact on network performance. For example, if network traffic exceeds 10%-30% of normal traffic for a short duration, with minimal impact on network performance, it is classified as mild traffic anomaly. If traffic exceeds 30%-50% of normal traffic for a sustained period, with a certain impact on network performance, it is classified as moderate traffic anomaly. If traffic exceeds 50% or more of normal traffic for a prolonged period, with a significant impact on network performance, it is classified as severe traffic anomaly.
[0095] The classified security data is stored in different partitions of the distributed storage system according to the label level. Distributed storage systems have the advantages of high reliability and high scalability. When storing security data, the data is stored in different partitions according to the first-level threat label and the second-level threat sub-label. For example, all traffic anomaly data is stored in a large partition. Within this partition, mild traffic anomaly data, moderate traffic anomaly data, and severe traffic anomaly data are stored in different sub-partitions according to the second-level threat sub-label. This storage method facilitates data management and query. When you need to query threat data of a certain type and severity, you can quickly locate the corresponding storage location.
[0096] The encryption key of the threat tag is configured according to the preset data security level. The data security level is determined by the importance and sensitivity of the data. For example, for security data related to the core business of the enterprise, it is set to a high security level; for general network operation data, it is set to a low security level. Different security levels correspond to different encryption keys. When a data access request is received, it is verified whether the key provided by the requester matches the encryption key of the target threat tag. If it matches, the data access interface corresponding to the threat tag is opened. This ensures that only authorized users can access the corresponding security data, thereby improving data security. For example, a company's network security manager has a high-security key. When he requests access to severe traffic anomaly data, the system will verify whether his key matches the encryption key of the severe traffic anomaly tag. If it matches, access is allowed, otherwise access is denied.
[0097] Example 6:
[0098] When optimizing the network behavior completion model, we first statistically analyze the completion error distribution of the security situation dataset at different time periods. The completion error can be measured by calculating the difference between the completed data and the actual real-world data. For example, when completing network traffic data, we can calculate the absolute or relative error between the completed and actual traffic values. The network's operating status may vary at different times, and the performance of the completion model may also vary. For example, during daytime on weekdays, when network traffic is high and complex, the completion model may face greater challenges, and the completion error may be relatively large. Meanwhile, during periods like late night when network traffic is low and relatively stable, the completion error may be smaller. By statistically analyzing the completion error at different time periods, we can understand the model's performance under different network conditions.
[0099] Based on the completion error distribution, identify high-error periods and corresponding key topological nodes. High-error periods are when completion errors are significantly higher than other periods, while key topological nodes are nodes with large completion errors during these periods. For example, between 3:00 PM and 5:00 PM each day, when network traffic is at its peak, if a core switch node experiences consistently high completion errors, then this period is considered a high-error period, and this core switch node is considered a key topological node. Identifying high-error periods and key topological nodes facilitates targeted model optimization.
[0100] Collect security parameter samples from key topological nodes during periods of high error incidence, perform feature engineering on these samples, and extract features that are strongly correlated with completion errors. Feature engineering includes operations such as data cleaning, feature selection, and feature transformation. For example, network traffic data from key topological nodes may contain some noise data that needs to be cleaned and removed. Through methods such as correlation analysis, select features that are highly correlated with completion errors, such as the rate of change of traffic and peak traffic. Transform some features, such as performing a logarithmic transformation on traffic data, to make its distribution more consistent with the model's requirements. Extracting features that are strongly correlated with completion errors can help the model better understand the causes of completion errors, thereby improving optimization results.
[0101] The extracted features are used to conduct supervised learning training on the network behavior completion model. During the training process, the features are used as input and the completion error is used as output. By continuously adjusting the model parameters, the model can better predict and reduce the completion error. Suppose the input feature vector is , the completion error is The output of the model is , you can use Mean Squared Error (MSE) as the loss function, and the calculation formula is:
[0102]
[0103] in, is the sample size, For the The true completion error of samples is For the model By minimizing the mean square error, the model parameters are continuously updated to improve the performance of the model.
[0104] After training is complete, the optimized network behavior completion model is validated. This validation involves testing in a real-world network environment and comparing the completion errors before and after optimization. For example, a time period different from the training data can be selected, and the security parameters of key topological nodes during that time period are input into the pre- and post-optimization models, respectively, to calculate the completion errors. If the completion error of the optimized model is significantly reduced, the optimization is effective. If the completion error does not significantly improve, it may be necessary to readjust the training parameters or further analyze the characteristics before re-optimization.
[0105] To ensure that the optimized model maintains good performance in different network scenarios, multi-scenario testing is also necessary. This can be done by simulating different network loads, network topologies, and attack scenarios. For example, a high-load network environment can be simulated to increase network traffic pressure and observe the model's completion error in this situation. Simulating different attack scenarios, such as DDoS attacks and port scan attacks, can also be used to verify the model's ability to accurately complete data under attack. Multi-scenario testing allows for a comprehensive assessment of the model's robustness and adaptability, ensuring its effective operation in a variety of complex network environments.
[0106] During the actual application of the model, a continuous monitoring mechanism must be established. The model's completion error and performance indicators should be monitored in real time. If the completion error exceeds a preset threshold, the model should be retrained and optimized promptly. The preset threshold can be determined based on the actual network's security requirements and performance standards. For example, when the completion error exceeds 15%, the model retraining process is triggered. Through continuous monitoring and optimization, the network behavior completion model can always accurately complete network security data, providing reliable data support for network security situational awareness.
[0107] Furthermore, with the continuous development of network technology and the evolving nature of cybersecurity threats, the model needs to be regularly updated and upgraded. New cybersecurity data should be collected and the model retrained to adapt to new network environments and security threats. For example, a large-scale model update could be conducted annually, incorporating new data collected over the past year into the training set and retraining the model to improve its accuracy and adaptability. At the same time, the latest research findings in the cybersecurity field should be monitored, and new algorithms and technologies should be applied to the model to continuously improve its performance and competitiveness.
[0108] In terms of data management, data used for model training and validation requires strict management and protection. A data backup mechanism should be established to regularly back up data to prevent data loss. Furthermore, data should be encrypted to ensure security and privacy. For example, symmetric encryption algorithms can be used to encrypt data, ensuring that only authorized personnel can decrypt and use the data. Secure transmission protocols, such as SSL / TLS, should also be used to prevent data theft or tampering during transmission.
[0109] During model deployment and maintenance, integration with existing network security systems must also be considered. Ensure that the optimized network behavior completion model seamlessly integrates with existing intrusion detection systems, firewalls, and other security devices, enabling data sharing and collaboration. For example, complete security data can be transmitted to the intrusion detection system in real time to help it more accurately detect network attacks; combine the model's analysis results with firewall policy configurations to achieve automated security protection. By integrating with existing network security systems, the model can fully leverage its capabilities and improve the security protection capabilities of the entire network.
[0110] Finally, to facilitate model operation and management for network security administrators, a corresponding user interface needs to be developed. This user interface should be intuitive and concise, allowing administrators to easily view information such as the model's operating status, complete errors, and perform model configuration and parameter adjustments. For example, a web-based user interface could be developed, allowing administrators to access the interface through a browser, monitor the model's operation in real time, and perform necessary operations. Detailed operation guides and help documentation should also be provided to reduce operational complexity and improve work efficiency.
[0111] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.
[0112] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A network security situation awareness method based on JS divergence, characterized in that: include: Synchronously acquiring security parameters of the target network environment through a multi-source data acquisition module, wherein the multi-source data acquisition module includes a monitoring unit for at least one network security element, and the security parameters include traffic characteristics, protocol types, access logs, and attack behaviors; Inputting the security parameters into a preset anomaly detection model based on JS divergence to identify abnormal data distribution in the security parameters to obtain preprocessed security data, wherein the anomaly detection model dynamically adjusts the detection threshold based on the distribution difference of historical security data; Inputting the pre-processed security data into a preset network behavior completion model to perform pattern completion on missing data or process conflicting data, and generate a continuous and unified security situation dataset, wherein the network behavior completion model determines the completion weight based on the topological association characteristics and temporal behavior characteristics of the target network; The security situation data set is classified into threat levels using a preset multi-dimensional threat classification algorithm and stored to form integrated perception data.
2. The network security situation awareness method based on JS divergence according to claim 1 is characterized in that: The steps of constructing the anomaly detection model include: Acquire a historical security data set, wherein each data in the historical security data set is annotated with an anomaly type and distribution difference; Divide the training subsets based on the anomaly type and distribution difference, each training subset corresponding to an attack scenario; use the training subsets to train the initial detection model in parallel until the anomaly recognition accuracy of the initial detection model for each attack scenario is greater than or equal to a preset first threshold, and then stop training to obtain an intermediate detection model; The historical safety data set is input into the intermediate detection model, and it is verified whether the anomaly recognition result output by the intermediate detection model meets the preset error range; if so, the intermediate detection model is determined as the anomaly detection model.
3. The network security situation awareness method based on JS divergence according to claim 1 is characterized in that: Synchronously acquiring the security parameters of the target network environment through the multi-source data acquisition module includes collecting monitoring data of one of the security elements in the following manner: Establishing a communication connection with a target monitoring unit, wherein the target monitoring unit is deployed at a key node of the target network environment; Periodically reading the real-time data stream of the target monitoring unit according to a preset sampling frequency, and marking a collection timestamp based on the timing characteristics of the real-time data stream; According to the logical topology relationship of the target network, the real-time data streams of different nodes at the same timestamp are logically aligned to form an associated security parameter set.
4. The network security situation awareness method based on JS divergence according to claim 1 is characterized in that: Inputting the security parameters into a preset JS divergence-based anomaly detection model includes: Extracting a distribution difference segment from the security parameter, wherein the distribution difference segment is a data segment in which a JS divergence calculation result exceeds a preset difference threshold within a continuous time window; Generate an anomaly assessment index based on the duration and divergence amplitude of the distribution difference segment; The corresponding detection algorithm is dynamically selected according to the anomaly assessment index, wherein the local density detection algorithm is used for short-term high-amplitude differences, and the sliding window detection algorithm is used for long-term low-amplitude differences.
5. The network security situation awareness method based on JS divergence according to claim 4 is characterized in that: The method further comprises: After identifying the abnormal data distribution, performing a data consistency check on the pre-processed security data; If the verification finds that the data conflict rate exceeds the preset second threshold, the network behavior completion model is triggered to perform priority completion on the conflict data, wherein the high-priority conflict data is a data segment with a continuous conflict duration exceeding the preset duration.
6. The network security situation awareness method based on JS divergence according to claim 1 is characterized in that: The network behavior completion model includes the following completion steps: Constructing a logical topology model based on the distribution of key nodes of the target network, wherein each topology node corresponds to the logical location of a monitoring unit; Calculate the logical completion weight based on the security parameter correlation of adjacent topological nodes; Combined with the temporal behavior trend of the security parameters, the missing or conflicting topological nodes are completed in a spatiotemporal joint manner.
7. The network security situation awareness method based on JS divergence according to claim 6 is characterized in that: The method further comprises: After the completion is completed, the completion result is verified for logical consistency, where the verification method includes comparing the deviation between the completed data and the actual data of the adjacent nodes; If the deviation exceeds a preset third threshold, the logic completion weight is readjusted and completion is iterated until the deviation is less than the third threshold.
8. The network security situation awareness method based on JS divergence according to claim 1 is characterized in that: Using a preset multi-dimensional threat classification algorithm to classify the security situation data set into threat levels includes: Classify the first-level threat labels according to the security parameter type, wherein the first-level threat labels include traffic anomaly, protocol violation, and attack behavior; Under each level of threat label, the secondary threat sub-labels are further divided based on the severity of the threat; The classified security data is stored in different partitions of the distributed storage system according to the label level.
9. The network security situation awareness method based on JS divergence according to claim 8 is characterized in that: The method further comprises: Configure encryption keys for threat tags based on pre-set data security levels; Upon receiving a data access request, verify whether the key provided by the requester matches the encryption key of the target threat tag; If a match is found, the data access interface for the corresponding threat tag is opened.
10. The network security situation awareness method based on JS divergence according to claim 1 is characterized in that: The method further includes optimizing the network behavior completion model in the following manner: Counting the distribution of completion errors of the security situation dataset at different time periods; Determining parameter adjustments for the completion model based on the error distribution, wherein a logic completion weight is increased during a high error period and a time completion weight is increased during a low error period; The network behavior completion model is iteratively optimized based on the parameter adjustment amount until a completion error rate is less than a preset fourth threshold.