A method for generating network security situation based on multi - perspective monitoring
Through the network security situation generation method based on multi-perspective monitoring, the problems of insufficient multi-level security monitoring coordination, weak data integration and correlation mechanism, and insufficient unknown threat detection capabilities in the existing technology are solved, and comprehensive perception of network security situations and efficient detection of complex attacks are achieved, and the intelligence and decision-making support capabilities of network security management are improved.
Patent Information
- Application Number
- CN202510127594.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-05
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-02-05
AI Technical Summary
When existing network security technologies face the diversified needs of complex network security environments, there are problems such as insufficient multi-level security monitoring coordination, poor data integration and correlation mechanism, insufficient unknown threat detection capabilities, improvement of real-time and dynamic response capabilities, and large-scale data processing and intelligent analysis.
A network security situation generation method based on multi-perspective monitoring is proposed. Data from five perspectives: network traffic, malicious code, network performance, network routing and network log are collected, data preprocessing and standardization are performed, logically layered and integrated through association identifiers and time windows, a unified association relationship is constructed, time series modeling is performed, basic information, advanced features and dynamic features are extracted, and abnormal behavior is used to identify abnormal behaviors, detect unknown threats and predict potential risks, and finally generate network security situations.
It realizes a comprehensive perception of the network security situation, captures potential attack links and hidden security threats, improves the detection capabilities of complex attacks, and enhances the intelligence and decision-making support capabilities of network security management.
Smart Images

Figure CN119583219B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of network security situation generation, and particularly relates to a network security situation generation method based on multi-perspective monitoring. Background Art
[0002] With the rapid development of the Internet and information technology, the network has become an important infrastructure for the operation of modern society. However, the complexity and diversity of network security are increasing continuously, and the attack means are also showing an increasingly complex and concealed trend. In order to cope with the ever-changing security challenges, traditional network security protection technologies, such as firewalls, intrusion detection systems, and antivirus software, have played an important role in specific scenarios. However, when dealing with the diverse requirements of the current complex network security environment, there are still some challenges and room for improvement:
[0003] (1) Insufficient coordination of multi-level security monitoring: The monitoring perspective of current network security technologies usually focuses on a certain level or specific dimension, such as traffic analysis at the network layer or log monitoring at the application layer. This single-level monitoring method may have certain limitations in capturing the full picture of security threats, especially in the correlation analysis of cross-level and multi-dimensional data, and there is further potential for optimization.
[0004] (2) The data integration and correlation mechanism needs to be strengthened: Since network security monitoring technologies are distributed at different levels, such as the network layer, device layer, and user behavior layer, the data lacks a unified identifier and an efficient correlation mechanism between systems, which may lead to a certain degree of dispersion in the data fusion process. This situation increases the complexity of multi-level data integration and feature analysis.
[0005] (3) The detection ability for unknown threats needs to be improved: Currently, many technologies rely on known threat feature libraries or rule matching. The response to advanced persistent threats (APTs) and zero-day attacks is still the focus of research in the field of network security. In complex attack scenarios, there is still a large room for improvement in the detection ability for multi-step operations and concealed behaviors.
[0006] (4) Real-time and dynamic response ability: With the rapid change of the network attack situation, network security technologies need to continuously improve the dynamic response ability to achieve faster and more effective threat defense and security policy adjustment.
[0007] (5) Large-scale data processing and intelligent analysis: With the continuous growth of the data scale and diversity in the network environment, how to efficiently process massive and multi-dimensional data and perform advanced feature extraction has become an important direction for realizing intelligent network security situation generation.
[0008] In summary, in order to meet the requirements of the current network security environment, there is an urgent need for an innovative network security situation generation technology that can comprehensively collect multi-level data, perform cross-level fusion and extract advanced features, and combine advanced technologies such as joint learning to enhance the detection ability of complex attacks. At the same time, through real-time situation generation and dynamic response mechanisms, the intelligence and decision-making support capabilities of network security management are further enhanced. Summary of the Invention
[0009] The object of the present invention is to propose a network security situation generation method based on multi-perspective monitoring for the problems existing in the background technology.
[0010] The technical solution of the present invention: A network security situation generation method based on multi-perspective monitoring includes the following specific implementation steps:
[0011] S1. Monitor and collect data from five perspectives: network traffic, malicious code, network performance, network routing, and network logs;
[0012] S2. Perform data preprocessing and standardization on the collected data; perform data preprocessing and standardization on the collected data, and clean the collected raw data based on a multi-factor cleaning method with dynamic threshold adaptation.
[0013] S3. Logically layer the processed raw data, including network layer, application layer, device layer, user behavior layer, and environment layer;
[0014] S4. Integrate the data of the network layer, application layer, device layer, user behavior layer, and environment layer through association identifiers and time windows to construct a unified association relationship;
[0015] S5. Perform time series modeling on the fused multi-perspective data to capture the dynamic characteristics of the data changing over time;
[0016] S6. Extract basic information, advanced features, and dynamic features to provide input for model learning;
[0017] S7. Use comprehensive features to construct a joint learning model. In the model learning stage, adopt a joint learning strategy of a shared network and multi-task branches, extract general features through the shared layer and combine task-specific branches for refined analysis to identify abnormal behaviors, detect unknown threats, and predict potential risks;
[0018] S8. Generate a network security situation based on statistical analysis and model prediction results. Combine the dynamic prediction results of model learning with the information generated by static analysis to form a network security situation. The generated situation information covers the identification and warning of abnormal events, the distribution of asset quantity and level, the distribution of vulnerability quantity and level, risk assessment and level division, and threat trend prediction.
[0019] Preferably, the specific implementation steps of the logical layering are as follows:
[0020] S21. After data preprocessing is completed, the data from network traffic monitoring, malicious code detection, network performance monitoring, network routing monitoring, and network log monitoring are divided into a network layer, an application layer, a device layer, a user behavior layer, and an environment layer according to attributes and analysis objectives:
[0021] S22. The network layer manages data related to communication activities, including network traffic patterns, protocol behaviors, and routing table updates. The data in the network layer comes from network traffic monitoring and network routing monitoring, reflecting the overall traffic characteristics and link operating status;
[0022] S23. The application layer centrally stores data related to the application running status, including data from network performance monitoring and network log monitoring;
[0023] S24. The device layer manages data related to the device running status and security threats, including the running logs, communication records, and internal alarm information of Internet of Things devices. Classify the data from malicious code detection into the device layer to monitor the running security and potential threats of the device;
[0024] S25. The user behavior layer stores user identity and operation behavior data;
[0025] S26. The environment layer includes external context information and environmental data;
[0026] S27. Attribute the data from different monitoring perspectives to their respective corresponding layers.
[0027] Preferably, the implementation process of integrating the data of the network layer, application layer, device layer, user behavior layer, and environment layer through association identifiers and time windows is as follows:
[0028] S31. Align each data layer with a unified timestamp and identifier field after the data preprocessing stage. Based on the common identifiers in the data, establish a preliminary logical association for records from different sources;
[0029] The IP address IP of the network layer net is aligned with the logged-in IP address IP of the user behavior layer user while the device ID of the device layer, i.e., DeviceID device , is matched with the device ID recorded in the application layer, i.e., DeviceID app , to construct a preliminary identifier mapping relationship through the consistency of field values;
[0030] S32. After the identifier alignment, solve the time difference problem between data from different sources through time window matching. To ensure the consistency of the dynamic event chain, set a reasonable time window ΔT, and associate the records with time differences within the tolerance range;
[0031] When the login event time T at the user behavior layer user and the traffic event time T at the network layer net , if it satisfies |T user -T net |ΔT, the two records are considered to be time-related, and thus the matching in the time dimension enhances the dynamic consistency of the associated data;
[0032] S33. Based on the identifier alignment and time window matching, deeply associate the inter-layer data.
[0033] Preferably, the association process of deeply associating the inter-layer data is as follows:
[0034] A1. Establish an association between the application layer and the user behavior layer. The log data of the application layer is matched with the login records and operation logs of the user behavior layer through the user ID, i.e., UserID, and the session ID, SessionID;
[0035] A2. Between the device layer and the network layer, establish an association through the device ID and the IP address: The running status of the device layer is matched with the device identification of the traffic records in the network layer through the device ID, i.e., DeviceID, and satisfies the following association conditions:
[0036] DeviceID device =DeviceID net ;
[0037] IP device =IP net ;
[0038] In the formula, IP device represents the IP address of the device layer; DeviceID net represents the device ID of the network layer; DeviceID device represents the device ID of the device layer;
[0039] After that, ensure the time consistency of the association through time window matching;
[0040] A3. The association between the network layer and the user behavior layer is established through the login IP address and the network traffic records. The source IP in the login event of the user behavior layer, i.e., IP uesr , is matched with the source IP of the network layer traffic records, i.e., IP net , and satisfies the following conditions:
[0041] IP net = IP uesr ;
[0042] A4. Ensure the association between user operations and network behaviors in the time dimension through time window matching;
[0043] A5. The environment perception layer provides context information for all layers. Through the node identifier in the topology structure, i.e., NodeID, and the IP address in the external threat intelligence, i.e., IP env , establish an association with the data of other layers;
[0044] Among them, the malicious IP, i.e., IP env directly matches the source IP in the network layer traffic, i.e., IP net , and the association conditions are as follows:
[0045] IP env = IP net ;
[0046] A6. Associate with the device layer through the malicious identifier DeviceID env , and then match the login records in the user behavior layer:
[0047] NodeID env ∈ {DeviceID, IP, UserID};
[0048] In the formula, NodeID env represents the malicious node identifier;
[0049] Based on this, quickly identify and locate the threat source, and determine the threat impact scope through the context information.
[0050] Preferably, the construction process of time series modeling is as follows:
[0051] S51. Extract event records from the results of cross-layer data association and generate a time series in chronological order;
[0052] The output of cross-layer data association contains a set of event record sets R assoc , each record r i includes the timestamp T of the event occurrence i , the comprehensive feature vector F i , and the associated identifier ID i :
[0053] R assoc = {r 1 , r 2 , …, r n}, r i = (T i , Fi , ID i );
[0054] Among them, the comprehensive feature vector F i = [f net , f app , f device , f user , f env is the multi-level feature generated during the cross-layer association process. Sort the record set according to the time stamp T i to generate the global time series:
[0055] S = {(T 1 , F 1 ), (T 2 , F 2 ), …, (T n , F n )}, T 1 ≤ T 2 ≤ … ≤ T n ;
[0056] S52. Based on the generated time series, construct a time series model, extract dynamic characteristics. The time series is first segmented into several subsequences according to a fixed time window ΔT. Each time window contains all events within its range. For each time window [T t , T t + ΔT], its subsequence is defined as:
[0057] S t = {(T i , F i )}|T t ≤ T i < T t + ΔT;
[0058] Among them, T t represents the start time of the time window, and T t + ΔT is the end time. The subsequence S t contains all events that meet the conditions;
[0059] S53. After the time window segmentation is completed, generate the comprehensive feature sequence X t within each time window for subsequent dynamic analysis and feature extraction. The representation form of the feature sequence is as follows:
[0060] X t = [F i , F j , …, F k , i, j, k ∈ S t ;
[0061] Among them, F i is the comprehensive feature vector of events within the time window, arranged in the chronological order of event occurrence.
[0062] Preferably, the extraction processes of basic information, advanced features, and dynamic features are as follows:
[0063] S61. Extract multi-dimensional and multi-level features from the results of cross-layer data association and dynamic analysis to provide comprehensive and accurate data support for subsequent model training and situation generation;
[0064] Feature extraction is divided into three parts: directly extract basic information from the processed original data; extract time series features and behavior pattern features from the established time series module; extract statistical features, association relationship features, graph features, context features, and resource features from cross-layer data association to construct the comprehensive feature vector F;
[0065] S62. After feature extraction is completed, prepare training data and labels for the model learning stage;
[0066] S63. For the abnormal event recognition and warning task, the historical attack logs and the warning records of the intrusion detection system label whether there is abnormal behavior for the samples. The label of the normal behavior sample is defined as label = 0, and the abnormal behavior sample is defined as label = 1;
[0067] The label of the risk assessment task is defined as the overall risk score and the corresponding risk level y;
[0068] The scoring label is calculated based on the influence scope and severity of historical events, while the risk level y is discretized through predefined thresholds, including:
[0069] Low risk, that is, y = 1;
[0070] Medium risk, that is, y = 2;
[0071] High risk, that is, y = 3;
[0072] S64. For the threat analysis and prediction task, historical threat events and their time series changes provide prediction targets for the samples, and the label is defined as the threat type y threat and the time series target value T trend ;
[0073] The threat type label y threat is classified and labeled based on historical threat types: y threat ∈{1, 2, …, C};
[0074] Among them, C represents the number of threat categories;
[0075] Time series target value T trend Based on the time series data of historical threat frequencies, annotation is performed through a trend analysis algorithm;
[0076] S65. After the label generation is completed, the data is divided into a training set, a validation set, and a test set according to the task requirements.
[0077] Preferably, the construction process of the comprehensive feature vector F is as follows:
[0078] S71. Basic information F basic , basic characteristics directly extracted from the original data, specifically including:
[0079] Basic elements of network communication, recorded as f ip_port ;
[0080] Network protocol type, recorded as f protocol ;
[0081] Distribution of packet sizes f packet_size ;
[0082] Traffic direction f traffic_dir ;
[0083] Based on this, the features are combined to obtain F basic =[f ip_port , f protocol , f packet_size , f traffic_dir ;
[0084] S72. Time series features F time , dynamic characteristics extracted from time series modeling, and time series features include:
[0085] Occurrence time of traffic peak f time_peak ;
[0086] Distribution of time intervals between events f time_interval ;
[0087] Standard deviation of operation frequency within a time window f time_frequency ;
[0088] Based on this, the features are combined to obtain F time =[f time_peak , f time_interval , f time_frequency ;
[0089] S73. Behavioral pattern features F behavior , extracting the causal relationship between dynamic events, and its specific features include:
[0090] Path length in the attack link f behavior_path ;
[0091] Node dependency f behavior_dependency ;
[0092] Critical behavior trigger probability f behavior_trigger ;
[0093] Feature F is obtained by combining them accordingly behavior =[f behavior_path ,f behavior_dependency ,f behavior_trigger ;
[0094] S74. Statistical feature F stat Extract the distribution and aggregation characteristics of the data. The specific features include:
[0095] Mean value of packet size f mean_packet ;
[0096] Variance of total traffic f var_traffic ;
[0097] Count of abnormal events within a specific time window f count_anomaly ;
[0098] Feature F is obtained by combining them accordingly stat =[f mean_packet ,f var_traffic ,f count_anomaly ;
[0099] S75. Association relationship feature F assoc That is, the association pattern between different data layers, specifically including:
[0100] Frequent pattern of user behavior and device status change f user_device ;
[0101] Conditional probability of API call and traffic anomaly f API_traffic ;
[0102] Interaction frequency between specific resources f resource_freq ;
[0103] Feature F is obtained by combining them accordingly assoc =[f user_device ,f API_traffic ,f resource_freq ;
[0104] S76. Graph feature F graph Based on the graph structure generated by cross-layer association, the specific features include:
[0105] Degree of nodes f graph_degree ;
[0106] Average clustering coefficient of the graph f graph-cluster ;
[0107] Distribution f of the shortest path between nodes path_dist ;
[0108] Based on this, the feature F is obtained by merging graph =[f graph_degree , f graph-cluster , f path_dist ;
[0109] S77. Context feature F context , providing a global perspective on the network background and external threats. The specific features include:
[0110] Whether the device hits the external threat intelligence library f threat_match ;
[0111] Dynamic change frequency f of the network topology topology ;
[0112] Physical parameters f of the device operating environment environment ;
[0113] Based on this, the feature F is obtained by merging context =[f threat_match , f topology , f environment ;
[0114] S78. Resource feature F resource , extracting the resource distribution and status in the network, including:
[0115] Number of devices f resource_count ;
[0116] Priority level f of the core device resource_priority ;
[0117] Current operating status f of the resource resource_status ;
[0118] Based on this, F is obtained by merging resource =[f resource_count , f resource_priority , f resource_status ;
[0119] S79. Through feature normalization and encoding, the dimensionality differences between different features are eliminated. The numerical features are standardized to a distribution with a mean of 0 and a standard deviation of 1, or normalized to the interval [0, 1]; one-hot encoding is used for categorical features, and TF-IDF is used for text features for vectorization processing;
[0120] The comprehensive feature vector F = [F basic , F time , F behavior , F stat, F assoc , F graph , F context , F resource .
[0121] Preferably, the construction process of the federated learning model is as follows:
[0122] S81. Input the comprehensive feature vector F = [F basic , F time , F behavior , F stat , F assoc , F graph , F context , F resource , perform standardization and encoding processing on the features. Numerical features are mapped to the interval [0, 1] through normalization operations, and categorical features are represented by one-hot encoding:
[0123] ;
[0124] f i = [1, 2, …, 0] T ;
[0125] Among them, represents the feature after standardization processing; f i represents the original vector of the i-th feature;
[0126] After processing, all features are merged into the feature matrix X ∈ R N×p ;
[0127] Among them, N is the number of samples, and p is the feature dimension;
[0128] S82. The shared representation layer is responsible for extracting task-irrelevant general information from the input features and generating the shared representation H shared , and performing high-dimensional feature mapping on the input feature matrix X through a fully connected layer:
[0129] H (1) = σ(W (1) X + b (1) );
[0130] Among them, W (1) ∈ R N×p represents the weight matrix; N is the number of samples, and p is the feature dimension;
[0131] S83. Introduce the multi-head attention mechanism, strengthen the correlation between features by calculating the attention weights, and the output feature after the multi-attention mechanism is H shared ;
[0132] The attention score a of each head ijThe calculation is as follows:
[0133] ;
[0134] Among them, score(h i , h j ) represents the similarity measure between features h i and h j ;
[0135] The output of multi-head attention is:
[0136] H (2) = Concat(head 1 , head 2 , …, head M )W 0 ;
[0137] Among them, M represents the number of attention heads; W 0 represents the projection matrix;
[0138] S84. On the basis of the shared representation layer, design an independent branch network for each task for specific task modeling and optimization:
[0139] S8401. Abnormal event recognition and alarm: The input is the shared representation H shared and the time series feature F time , combined with the behavior pattern feature F behavior , that is, the input feature: H event = [H shared , F time , F behavior ;
[0140] Then use the hybrid model to capture time dependence and causal relationship, and the abnormal detection output is:
[0141] y event = σ(W event H event + b event );
[0142] In the formula, y event represents whether the event is abnormal, y event ∈ [0, 1]; W event and b event represent the model parameters for the model to recognize and alarm abnormal events;
[0143] S8402. Risk assessment and grading: Use the shared representation H shared , statistical feature F stat , context feature F context and resource feature F resourcePerform risk scoring and level classification to obtain the risk assessment input feature H risk =[H shared ,F stat ,F context ,F resource , and the shared representation input is:
[0144] R = W risk H risk + b risk ;
[0145] In the formula, R represents the risk score; W risk and b risk represent the model parameters for the model to perform risk assessment;
[0146] Complete risk level classification through the Softmax activation function:
[0147] ;
[0148] In the formula, k represents the risk level; W cls,k represents the weight of the kth risk level;
[0149] S8403. Threat trend prediction: Combine the shared representation H shared and the time series feature F time and the correlation relationship feature F assoc , that is, the threat trend prediction input feature H trend =[H shared ,F time ,F assoc , and predict and analyze the evolution trend of the threat in time and space;
[0150] Capture the trend change through the time series modeling layer GRU: z trend = GRU(H trend ), and predict the occurrence probability of the threat event and generate a trend chart;
[0151] S8404. The total loss function of joint learning combines the independent losses of each task and is defined as follows:
[0152] L = λ event L event + λ risk L risk + λ trend L trend ;
[0153] Among them, λ event , λ risk and λ trend respectively represent the task weights of the anomaly event recognition, risk assessment, and trend analysis tasks; L event , Lrisk and L trend respectively represent the losses of the abnormal event recognition, risk assessment, and trend analysis tasks.
[0154] Preferably, the monitoring process for collecting data from five perspectives of monitoring the network traffic, malicious code, network performance, network routing, and network logs is as follows:
[0155] S91. In the initial stage of network security situation generation, multi-perspective monitoring and data collection capture comprehensive information about the network environment from multiple dimensions. Taking the five perspectives of network traffic monitoring, malicious code detection, network performance monitoring, network routing monitoring, and network log monitoring as the starting points, raw data covering network communication, application behavior, device status, user activities, and environmental dynamics is comprehensively obtained;
[0156] S92. In network traffic monitoring, deploy traffic collection tools to record the traffic patterns and communication details of data packets in real time, including source IP, destination IP, protocol type, and packet size;
[0157] S93. Malicious code detection relies on host-level and network-level monitoring tools to detect file system changes, abnormal process activities, and malicious behavior samples, and assists in further analysis of malicious activities by aggregating feature data;
[0158] S94. Network performance monitoring uses a distributed monitoring platform to collect link status, latency time, packet loss rate, and throughput, revealing the dynamic changes and bottlenecks in network performance;
[0159] S95. Network routing monitoring collects routing table updates, topology changes, and link anomalies through protocol monitoring tools to form a real-time view of network dynamic changes;
[0160] S96. Network log monitoring focuses on the operation logs of systems and devices, and collects event records through a centralized log management tool, including user login information, permission changes, and security events.
[0161] Preferably, the cleaning process for cleaning the collected raw data is as follows:
[0162] S101. Define a multi-factor weight calculation model, comprehensively consider the data time distribution, frequency characteristics, and similarity characteristics, and dynamically generate a cleaning threshold. The formula is as follows:
[0163] ;
[0164] ;
[0165] ;
[0166] ;
[0167] ;
[0168] In the formula, T represents the dynamic cleaning threshold; w t , w f and w s respectively represent the weights of the time distribution factor, the frequency feature factor, and the similarity feature factor, w t + w f + w s = 1; f t (x) represents the time distribution factor, measuring the deviation degree of the data point from the normal time interval; f f (x) represents the frequency feature factor, evaluating the frequency deviation degree of the data point in the statistical distribution; f s (x) represents the similarity feature factor, reflecting the content similarity of the data point with the neighboring points; t x represents the timestamp of the data point; represents the mean value of the timestamps of the same type of data; represents the standard deviation of the timestamps of the same type of data; f x represents the occurrence frequency of the data point; f max represents the maximum frequency in the entire dataset; cos(θ x ) represents the vector similarity between the data point and its neighboring data; represents the feature vector of the target data point x; represents the feature vector of the neighboring data; and respectively represent the vector modulus of the target point and the modulus of the feature vector of the neighboring data point;
[0169] S102. Define the cleaning rule:
[0170] ;
[0171] In the formula, x represents the original value of the data point; μ represents the mean value of the same type of data; σ represents the standard deviation of the same type of data;
[0172] S103. Use an adaptive strategy to adjust the weights during the cleaning process:
[0173] ;
[0174] In the formula, represents the weight of the time distribution factor in the (n + 1)-th round; Δ t represents the contribution gain of the time distribution factor to the cleaning result; α represents the adjustment rate, i.e., the learning rate.
[0175] Compared with the prior art, the above technical solution of the present invention has the following beneficial technical effects:
[0176] The present invention provides a method for generating network security situation based on multi - perspective monitoring, aiming to achieve a comprehensive perception of network security situation through multi - perspective data monitoring and collection, correlation analysis, and model learning:
[0177] (1) The present invention comprehensively collects data from five perspectives: network traffic monitoring, malicious code detection, network performance monitoring, network routing monitoring, and network log monitoring, breaking through the limitations of traditional single - perspective monitoring. By adopting a unified correlation identifier and time synchronization mechanism, combined with a logical layering method, the collected data is divided into network layer, application layer, device layer, user behavior layer, and environment layer. The method effectively eliminates the data island problem, can comprehensively perceive the network security situation, capture potential attack links and hidden security threats, and provides strong support for accurate security event positioning and attack path identification;
[0178] (2) Based on the multi - perspective data after cross - layer correlation, the present invention extracts high - level feature vectors including basic information, high - level static features, and dynamic features, and adopts a joint learning method to complete multiple network security tasks simultaneously. By constructing a shared network to extract common features and combining with specific task - specific branches, the present invention significantly improves the overall performance and collaborative ability for tasks such as abnormal event detection, risk assessment and grading, and threat trend prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0179] Figure 1 is a method flow chart of a method for generating network security situation based on multi - perspective monitoring proposed by the present invention;
[0180] Figure 2 is a schematic diagram of the joint learning process of the joint learning model proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0181] Embodiment 1, as Figures 1 to 2 shown, a method for generating network security situation based on multi - perspective monitoring proposed by the present invention, as Figure 1 shown, the specific implementation steps are as follows:
[0182] S1. Monitor and collect data from five perspectives: network traffic, malicious code, network performance, network routing, and network logs;
[0183] S2. Perform data pre - processing and standardization on the collected data, and clean the collected original data based on a multi - factor cleaning method with dynamic threshold adaptation;
[0184] S3. Perform logical layering on the processed original data, including network layer, application layer, device layer, user behavior layer, and environment layer;
[0185] S4. Integrate the data from the network layer, application layer, device layer, user behavior layer, and environment layer through association identifiers and time windows to construct a unified association relationship;
[0186] S5. Perform time series modeling on the fused multi-perspective data to capture the dynamic characteristics of the data changing over time;
[0187] S6. Extract basic information, advanced features, and dynamic features, including but not limited to time series features, graph features, and behavior patterns, to provide input for model learning;
[0188] S7. Utilize the comprehensive features to construct a joint learning model. In the model learning stage, adopt a joint learning strategy of a shared network and multi-task branches. Extract general features through the shared layer and combine with task-specific branches for refined analysis to identify abnormal behaviors, detect unknown threats, and predict potential risks;
[0189] S8. Generate a network security situation based on statistical analysis and model prediction results. Combine the dynamic prediction results of model learning with the information generated by static analysis to form a network security situation. The generated situation information covers abnormal event identification and alerting, asset quantity and level distribution, vulnerability quantity and level distribution, risk assessment and level division, and threat trend prediction.
[0190] In an optional embodiment, the implementation process of monitoring and collecting data from five perspectives of network traffic, malicious code, network performance, network routing, and network logs is as follows:
[0191] S11. In the initial stage of network security situation generation, multi-perspective monitoring and data collection capture comprehensive information about the network environment from multiple dimensions to provide basic support for subsequent data analysis and security modeling. Taking the five perspectives of network traffic monitoring, malicious code detection, network performance monitoring, network routing monitoring, and network log monitoring as the entry points, comprehensively obtain the original data covering network communication, application behavior, device status, user activities, and environmental dynamics;
[0192] S12. In network traffic monitoring, deploy traffic collection tools (including but not limited to NetFlow or sFlow) to record the traffic patterns and communication details of data packets in real time, including source IP, destination IP, protocol type, and packet size, to provide accurate basic data for identifying traffic anomalies and communication patterns;
[0193] S13. Malicious code detection relies on host-level and network-level monitoring tools to detect file system changes, abnormal process activities, and malicious behavior samples, and assist in further analysis of malicious activities by aggregating feature data (including but not limited to file hash values, behavior logs, and abnormal samples);
[0194] S14. The network performance monitoring uses a distributed monitoring platform to collect link status, latency time, packet loss rate, and throughput, revealing the dynamic changes and bottlenecks in network performance;
[0195] S15. The network routing monitoring uses protocol monitoring tools (including but not limited to SNMP) to collect routing table updates, topology changes, and link anomalies, forming a real-time view of the dynamic changes in the network;
[0196] S16. The network log monitoring focuses on the operation logs of systems and devices, collecting event records (including user login information, permission changes, and security events) through a centralized log management tool (including but not limited to the ELK Stack), providing detailed basis for behavior auditing and anomaly detection.
[0197] In an optional embodiment, the preprocessing process of preprocessing and standardizing the collected data is as follows:
[0198] S21. Clean the collected raw data, including removing invalid fields, redundant records, and repairing missing values;
[0199] For network traffic data, the cleaning process removes incomplete session records in the communication and corrects abnormal timestamps;
[0200] For malicious code detection data, non-critical features with low correlation are removed, including but not limited to invalid file operations or log records of non-abnormal processes;
[0201] S22. Perform data format conversion. For the data formats generated by different collection tools (including but not limited to JSON, CSV, or binary log files), they are uniformly converted into a structured data table format, and field extraction and parsing are performed on text-based logs;
[0202] For example, extract timestamps, event types, and user information from network logs for subsequent logical layer attribution and cross-layer correlation analysis;
[0203] S23. Use normalization and standardization techniques to process numerical data (including but not limited to traffic volume, latency time) to make it have a unified dimension among different monitoring sources. For time series data (including but not limited to link latency in network performance monitoring), data synchronization among different monitoring sources is achieved through time alignment and interpolation methods;
[0204] S24. To improve analysis efficiency and expression ability, feature derivation is carried out for different types of data;
[0205] For example, extract communication frequency and session duration from traffic data, and generate behavior feature vectors from malicious code data;
[0206] Among them, the derived features not only enhance the usability of the data, but also provide higher-dimensional support for the extraction of basic information and advanced features in the logical hierarchy.
[0207] In an optional embodiment, the specific implementation steps of the logical hierarchy are as follows:
[0208] S31. After the data preprocessing is completed, the data from network traffic monitoring, malicious code detection, network performance monitoring, network routing monitoring, and network log monitoring are divided into a network layer, an application layer, a device layer, a user behavior layer, and an environment layer according to attributes and analysis objectives:
[0209] S32. The network layer manages data related to communication activities, including network traffic patterns, protocol behaviors, and routing table updates. The network layer data comes from network traffic monitoring and network routing monitoring, covers the basic situation of network communication, and can reflect the overall traffic characteristics and link operation status;
[0210] S33. The application layer centrally stores data related to the application running status, including but not limited to server logs, API call records, and performance metrics. The data from network performance monitoring and network log monitoring are classified into this layer to provide support for analyzing the stability and performance of application services;
[0211] S34. The device layer manages data related to the device running status and security threats, including the running logs, communication records, and internal alarm information of Internet of Things devices. The data from malicious code detection is attributed to this layer, including but not limited to device status changes and communication logs, for monitoring the running security and potential threats of devices;
[0212] S35. The user behavior layer stores user identity and operation behavior data, including but not limited to login information, operation logs, and permission change records;
[0213] Among them, the source of these data is network log monitoring, especially the system logs and activity logs that record user operation behaviors. By analyzing the user's identity, operation path, and behavior patterns, it supports the tracing of user activities and the detection of abnormal behaviors;
[0214] S36. The environment layer covers external context information and environmental data;
[0215] S37. Accordingly, the data from different monitoring perspectives are attributed to their respective corresponding layers, ensuring the clarity of the data in terms of semantics and functions, and laying a good foundation for subsequent cross-layer associations.
[0216] In an optional embodiment, the implementation process of integrating data from the network layer, application layer, device layer, user behavior layer, and environment layer through associated identifiers and time windows is as follows:
[0217] S41. After the data preprocessing stage, each data layer has a unified timestamp and identifier field. Through identifier alignment, based on the common identifiers in the data (including but not limited to IP addresses, user IDs, device IDs), a preliminary logical association is established for records from different sources;
[0218] For example, the IP address IP in the network layer net is aligned with the login IP address IP in the user behavior layer user while the device ID (DeviceID device ) in the device layer is matched with the device ID (DeviceID app ) recorded in the application layer. A preliminary identifier mapping relationship is constructed through the consistency of field values, laying a foundation for further association;
[0219] S42. After identifier alignment, the time difference problem between data from different sources is solved through time window matching. To ensure the consistency of the dynamic event chain, a reasonable time window ΔT is set, and records with a time difference within the tolerance range are associated;
[0220] For example, when the login event time T user in the user behavior layer and the traffic event time T net in the network layer user satisfy |T net -T | ≤ ΔT, the two records are considered to be time - related, and the matching in the time dimension enhances the dynamic consistency of the associated data;
[0221] S43. Based on identifier alignment and time window matching, the in - depth association of inter - layer data is further carried out. The specific implementation process is as follows:
[0222] S4301. Establish an association between the application layer and the user behavior layer: The log data in the application layer (including but not limited to API call records, application server logs) is matched with the login records and operation logs in the user behavior layer through the user ID (UserID) and session ID (SessionID);
[0223] For example, when a certain user initiates an API call and performs a file operation after logging in, the scattered events are associated through user identification and session information to generate a complete user operation path;
[0224] S4302. Establish an association between the device layer and the network layer through the device ID and the IP address: The operating status of the device layer (including but not limited to device alarms and status changes) is matched with the device identifiers in the traffic records of the network layer (including but not limited to communication logs and abnormal traffic) through the device ID (DeviceID), satisfying the following association conditions:
[0225] DeviceID device =DeviceID net ;
[0226] IP device =IP net ;
[0227] Wherein, IP device represents the IP address of the device layer; DeviceID net represents the device ID of the network layer; DeviceID device represents the device ID of the device layer;
[0228] After that, ensure the time consistency of the association through time window matching;
[0229] For example, when a device runs abnormally, if a sudden increase in traffic on the relevant network interface is detected at the same time, it is determined that the abnormality may be caused by a network attack;
[0230] S4303. The association between the network layer and the user behavior layer is established through the login IP address and the network traffic records: The source IP (IP uesr ) in the login event of the user behavior layer is matched with the source IP (IP net ) in the network layer traffic records, satisfying the following conditions:
[0231] IP net =IP uesr ;
[0232] S4304. Ensure the association between user operations and network behaviors in the time dimension through time window matching;
[0233] For example, when a user logs in and a sudden increase in abnormal traffic on the corresponding IP address is detected, it is marked as a potential threat;
[0234] S4305. The environmental perception layer provides context information for all layers and establishes an association with the data of other layers through the node identifier (NodeID) in the topological structure and the IP address (IP env ) in the external threat intelligence;
[0235] For example, a malicious IP (IP env ) directly matches the source IP (IP net) The associated conditions are as follows:
[0236] IP env = IP net ;
[0237] S4306. Associated with the device layer through a malicious identifier (DeviceID env ) to match the login record in the user behavior layer:
[0238] NodeID env ∈ {DeviceID, IP, UserID};
[0239] In the formula, NodeID env represents the malicious node identifier;
[0240] Based on this, quickly identify and locate the threat source, and determine the threat impact scope through context information;
[0241] For example, when the external threat (IP env ) at a certain moment in the environment layer matches the network layer traffic and user operations, it is confirmed that it attacks the network interface of device A.
[0242] In an alternative embodiment, time series modeling takes the result of cross-layer data association as input, aiming to generate a globally event sequence with time dependence, providing basic data support for subsequent dynamic analysis and feature extraction, that is, extracting event records, organizing events in chronological order, and preprocessing the event sequence within a time window. The specific implementation process is as follows:
[0243] S51. Extract event records from the result of cross-layer data association and generate a time series in chronological order;
[0244] The output of cross-layer data association contains a set of event record sets R assoc , and each record r i includes the timestamp T i of the event occurrence, the comprehensive feature vector F i , and the associated identifier ID i :
[0245] R assoc = {r 1 , r 2 , …, r n}, r i = (T i , F i , ID i );
[0246] Among them, the comprehensive feature vector F i = [f net , fapp , f device , f user , f env are multi-level features generated during the cross-layer association process. Sort the record set according to the time stamp T i to generate a global time series:
[0247] S = {(T 1 , F 1 ), (T 2 , F 2 ), …, (T n , F n )}, T 1 ≤ T 2 ≤ … ≤ T n ;
[0248] S52. Based on the generated time series, perform time series modeling to extract dynamic characteristics. The time series is first segmented into several subsequences according to a fixed time window ΔT. Each time window contains all events within its range. For each time window [T t , T t + ΔT], its subsequence is defined as:
[0249] S t = {(T i , F i )}|T t ≤ T i < T t + ΔT;
[0250] where T t represents the start time of the time window, and T t + ΔT is the end time. The subsequence S t contains all events that meet the conditions;
[0251] S53. After the time window segmentation is completed, generate the comprehensive feature sequence X t for subsequent dynamic analysis and feature extraction. The representation form of the feature sequence is as follows:
[0252] X t = [F i , F j , …, F k , i, j, k ∈ S t ;
[0253] where F i is the comprehensive feature vector of events within the time window, arranged in the order of the occurrence time of events.
[0254] In an alternative embodiment, the extraction process of the basic information, advanced features, and dynamic features is as follows:
[0255] S61. Extract multi-dimensional and multi-level features from the results of cross-layer data association and dynamic analysis to provide comprehensive and accurate data support for subsequent model training and situation generation;
[0256] Feature extraction is divided into three parts: directly extract basic information from the processed raw data; extract time series features and behavior pattern features from the established time series module; extract statistical features, association relationship features, graph features, context features, and resource features from cross-layer data association;
[0257] The extracted features complement each other and jointly construct a complete feature space. The feature information extraction process is as follows:
[0258] S6101. Basic information F basic is the basic feature directly extracted from the raw data, describing the core information of network activities and entity states, specifically including:
[0259] The basic elements of network communication (including but not limited to source IP, destination IP, source port, and destination port), recorded as f ip_port ;
[0260] The network protocol type (including but not limited to TCP, UDP, HTTP), recorded as f protocol ;
[0261] The distribution of packet sizes f packet_size and the traffic direction (upstream or downstream) f traffic_dir ;
[0262] Based on this, the combined feature F basic =[f ip_port , f protocol , f packet_size , f traffic_dir ;
[0263] S6102. Time series feature F time is the dynamic feature extracted from time series modeling, describing the law of data change over time. The time series features include:
[0264] The occurrence time of the traffic peak f time_peak ;
[0265] The distribution of the time intervals between events f time_interval ;
[0266] The standard deviation of the operation frequency within the time window f time_frequency ;
[0267] Feature F is obtained by merging accordingly time =[f time_peak , f time_interval , f time_frequency ;
[0268] S6103. Behavioral pattern feature F behavior Focuses on the causal relationships between dynamic events, and its specific features include:
[0269] The path length f in the attack link behavior_path ;
[0270] Node dependency f behavior_dependency ;
[0271] The critical behavior trigger probability f behavior_trigger ;
[0272] Feature F is obtained by merging accordingly behavior =[f behavior_path , f behavior_dependency , f behavior_trigger ;
[0273] For example, the path length feature extracted from the attack link graph G attack =(V, E) reflects the complexity from the initial attack point to the target node, while the critical behavior trigger probability can reveal which operations are most likely to lead to the occurrence of security events;
[0274] S6104. Statistical feature F stat Describes the distribution and aggregation characteristics of data, providing insights into the global distribution of network activities. Specific features include:
[0275] The mean value f of the packet size mean_packet ;
[0276] The variance f of the total traffic volume var_traffic ;
[0277] The count f of abnormal events within a specific time window count_anomaly ;
[0278] Feature F is obtained by merging accordingly stat =[f mean_packet , f var_traffic , f count_anomaly ;
[0279] For example, a large variance in the total traffic volume may indicate a sudden increase in peak traffic in the network, while the count of abnormal events directly reflects the frequency of potential threats in the network;
[0280] S6105. Association relationship feature F assocUsed to reveal the association patterns between different data layers, helping the model understand the complex interaction relationships between different entities, specifically including:
[0281] The frequent patterns of user behavior and device status changes f user_device ;
[0282] The conditional probability of API calls and traffic anomalies f API_traffic ;
[0283] The interaction frequency between specific resources f resource_freq ;
[0284] Based on this, the feature F is obtained by merging assoc =[f user_device , f API_traffic , f resource_freq ;
[0285] For example, the frequent association between permission changes and abnormal API calls triggered after a user logs in may indicate potential permission abuse;
[0286] S6106. Graph feature F graph The graph structure generated based on cross-layer associations describes the structured information of nodes and edges in the network. The specific features include: the degree of nodes f graph_degree , the average clustering coefficient of the graph f graph-cluster , the distribution of the shortest paths between nodes f path_dist ;
[0287] Among them, the above graph feature F graph reveals the structural properties of the network topology. Nodes with a higher degree may represent core resources, while subgraphs with a higher clustering coefficient may represent the aggregation areas of attack targets;
[0288] Based on this, the feature F is obtained by merging graph =[f graph_degree , f graph-cluster , f path_dist ;
[0289] S6107. Context feature F context Coming from the environmental layer, it provides a global perspective on the network background and external threats. These features can assist the model in quickly locating the threat source. The specific features include:
[0290] Whether the device hits the external threat intelligence library f threat_match ;
[0291] The dynamic change frequency of the network topology f topology ;
[0292] The physical parameters of the device operating environment (including but not limited to temperature) f environment ;
[0293] Combine them accordingly to obtain feature F context =[f threat_match , f topology , f environment ;
[0294] For example, by matching malicious IP addresses or detecting abnormal fluctuations in the device operating environment;
[0295] S6108. Resource feature F resource It describes the resource distribution and status in the network, including the number of devices f resource_count , the priority level f of core devices resource_priority , and the current operating status of resources (including but not limited to online or offline) f resource_status ;
[0296] For example, the priority level of key resources helps the model to focus on high-value targets in threat detection, and abnormal changes in the resource operating status may be a direct signal of an attack behavior;
[0297] Combine them accordingly to obtain F resource =[f resource_count , f resource_priority , f resource_status ;
[0298] S62. Through feature normalization and encoding, eliminate the dimensionality differences between different features, standardize numerical features to a distribution with a mean of 0 and a standard deviation of 1, or normalize them to the interval [0, 1];
[0299] Use one-hot encoding for categorical features and TF-IDF for text features for vectorization processing. Finally, obtain the comprehensive feature vector F = [F basic , F time , F behavior , F stat , F assoc , F graph , F context , F resource ;
[0300] Among them, the comprehensive feature vector F contains both dynamic features (including but not limited to time series features, behavior pattern features) and static features (including but not limited to statistical features, graph features), providing high-quality input support for model training and network security situation generation;
[0301] S63. After feature extraction is completed, prepare the training data and labels for the model learning stage. For tasks that require supervised learning, including but not limited to anomaly event recognition, risk assessment and grading, and threat analysis and prediction, the label y is a necessary input for model training, which defines the mapping relationship between the input features and the output results. The generation of the label is based on historical data, rule matching, or expert annotation to ensure that the training data has high-quality target values;
[0302] 64. For the anomaly event recognition and alert task, the historical attack logs and the alert records of the intrusion detection system (IDS) annotate whether there is abnormal behavior for the samples. The label of the normal behavior samples is defined as label = 0, and the abnormal behavior samples are defined as label = 1;
[0303] The label for the risk assessment task is defined as the overall risk score and the corresponding risk level y;
[0304] The scoring label is calculated based on the impact scope and severity of historical events, while the risk level y is discretized through predefined thresholds, including but not limited to low risk (y = 1), medium risk (y = 2), and high risk (y = 3);
[0305] S65. For the threat analysis and prediction task, the historical threat events and their time series changes provide the prediction targets for the samples, and the label is defined as the threat type y threat and the time series target value T trend ;
[0306] The threat type label y threat is classified and annotated based on historical threat types (including but not limited to DDoS attacks, data breaches), for example, y threat ∈{1,2,…,C};
[0307] where C represents the number of threat categories;
[0308] The time series target value T trend is annotated through a trend analysis algorithm based on the time series data of historical threat frequencies;
[0309] In addition, for tasks that do not require supervised learning, including but not limited to the distribution of vulnerability numbers and grades, and the distribution of asset numbers and grades, the results are directly generated through statistical analysis and rule matching without explicit labels;
[0310] For example, the asset scanning tool automatically classifies assets according to device type, sensitivity, and usage, and the vulnerability scanning tool grades vulnerabilities according to the CVSS score;
[0311] After the label generation is completed, the data is divided into a training set, a validation set, and a test set according to the task requirements to ensure the generalization ability of model learning and the reliability of the results.
[0312] In an optional embodiment, as Figure 2 shown, the construction process of the federated learning model is as follows:
[0313] S71. Input the comprehensive feature vector F = [F basic , F time , F behavior , F stat , F assoc , F graph , F context , F resource . To ensure the consistency of the data, first perform standardization and encoding processing on the features. The numerical features are mapped to the [0, 1] interval through normalization operations, and the categorical features are represented by one-hot encoding:
[0314] ;
[0315] f i = [1, 2, …, 0] T ;
[0316] Among them, represents the feature after standardization processing; f i represents the original vector of the i-th feature;
[0317] After processing, all features are combined into the feature matrix X ∈ R N×p ;
[0318] Among them, N is the number of samples, and p is the feature dimension;
[0319] S72. The shared representation layer is responsible for extracting task-irrelevant general information from the input features and generating the shared representation H shared , and performing high-dimensional feature mapping on the input feature matrix X through a fully connected layer:
[0320] H (1) = σ(W (1) X + b (1) );
[0321] Among them, W (1) ∈ R N×p represents the weight matrix, N is the number of samples, and p is the feature dimension;
[0322] S73. To capture the interaction relationships between features, the multi-head attention mechanism (Multi-Head Attention) is introduced. By calculating the attention weights, the correlation between features is strengthened. The output feature after the multi-attention mechanism is H shared ;
[0323] The attention score a of each head ij is calculated as follows:
[0324] ;
[0325] where score(h i , h j ) represents the similarity measure between features h i and h j ;
[0326] The output of the multi-head attention is:
[0327] H (2) = Concat(head 1 , head 2 , …, head M )W 0 ;
[0328] where M represents the number of attention heads; W 0 represents the projection matrix;
[0329] S74. On the basis of the shared representation layer, an independent branch network is designed for each task, which is used for the modeling and optimization of specific tasks:
[0330] S7401. Abnormal event recognition and alarm: The input is the shared representation H shared and the time series feature F time , combined with the behavior pattern feature F behavior , that is, the input feature: H event = [H shared , F time , F behavior ;
[0331] Then, a hybrid model is used to capture the time dependence and causal relationship. The abnormal detection output is:
[0332] y event = σ(W event H event + b event );
[0333] In the formula, y event ∈ [0, 1] indicates whether the event is abnormal; W event and b eventModel parameters for the model to identify and alarm abnormal events;
[0334] S7402, Risk assessment and level classification: Using the shared representation H shared and static features (statistical feature F stat , context feature F context , resource feature F resource ) to perform risk scoring and level classification, that is, the risk assessment input feature H risk =[H shared ,F stat ,F context ,F resource , and the input of the shared representation is:
[0335] R = W risk H risk + b risk ;
[0336] In the formula, R represents the risk score; W risk and b risk represent the model parameters for the model to perform risk assessment;
[0337] Complete risk level classification through the Softmax activation function: ;
[0338] In the formula, k represents the risk level (high, medium, low); W cls,k represents the weight of the kth risk level;
[0339] S7403, Threat trend prediction: Combining the shared representation H shared and time series features F time and correlation relationship features F assoc , that is, the threat trend prediction input feature H trend =[H shared ,F time ,F assoc , and predict and analyze the evolution trend of threats in time and space;
[0340] Capture trend changes through the time series modeling layer GRU (Gated Recurrent Unit): z trend = GRU(H trend );
[0341] Predict the occurrence probability of threat events and generate a trend chart;
[0342] S7404, The total loss function of joint learning combines the independent losses of each task and is defined as follows:
[0343] L = λ event L event+λ risk L risk +λ trend L trend ;
[0344] Among them, λ event , λ risk and λ trend respectively represent the task weights of anomaly event recognition, risk assessment, and trend analysis tasks; L event , L risk and L trend respectively represent the losses of anomaly event recognition, risk assessment, and trend analysis tasks;
[0345] Accordingly: By extracting task - common features through a shared network and combining with the refined learning of task - specific branches, the joint learning framework realizes the collaborative optimization among multiple tasks, effectively improving the comprehensive perception ability of complex network security situations.
[0346] In an optional embodiment, the process of generating and displaying network security situations is as follows:
[0347] S81. The generation of dynamic information relies on the results of model learning: Through the output of the anomaly detection model, a security event warning list is generated, the content of which includes but is not limited to event type, occurrence time, source IP, target IP, and affected devices;
[0348] To highlight the urgency of high - risk events, the warning list is marked with priorities, facilitating rapid response by security management personnel. Combining with the results of the risk assessment model, the overall risk score and risk - level distribution map of the network are generated. The risk score is mapped to the risk levels of specific assets, highlighting the high - risk assets that need to be key - protected in the network and highlighting these assets for more precise security policy formulation;
[0349] S82. The generation of static information uses non - model methods for statistical analysis, covering asset quantity and level distribution, vulnerability quantity and level distribution, and threat trend analysis;
[0350] The asset quantity and level distribution are based on the results of asset scanning tools. Assets are classified into multiple categories (including but not limited to core devices, ordinary devices, and low - priority devices) according to device type, usage, and sensitivity, providing the distribution of various types of assets in the network;
[0351] The vulnerability quantity and level distribution combine the results of vulnerability scanning tools and are classified and displayed according to vulnerability severity (including but not limited to CVSS scores), providing the vulnerability quantity and level distribution of each device; These results intuitively show the possible security weak points in the network;
[0352] Threat trend analysis is based on time series tools, combines the frequency and type of security incidents, generates an attack trend chart, intuitively shows the change trend of threats in different time periods, and helps predict the possible future threat directions;
[0353] S83. The integration of multi-dimensional information realizes the display of the global situation by organically combining dynamic and static information;
[0354] The dynamically generated alarm events are combined with the asset distribution information to form an alarm asset list, indicating the specific assets involved in each alarm event and the potential impact scope;
[0355] The risk level and vulnerability distribution data are superimposed to generate a chart of the threatened degree of high-risk assets, clarifying the devices and vulnerabilities that need to be repaired first;
[0356] The threat trend analysis results are combined with the static information of the asset status to generate a global network security situation map, providing an overall security perspective for managers;
[0357] S84. The visual display of the situation information adopts various forms, which is convenient for decision-makers to quickly understand and analyze;
[0358] The security incident map intuitively shows the sources and targets of events from a geographical or topological perspective, helping to track the threat path;
[0359] The risk distribution map shows the risk level distribution of assets and vulnerabilities in the form of bar charts and pie charts, clearly marking the priorities;
[0360] The trend analysis chart shows the change of threats over time in the form of line charts or heat maps, which is convenient for identifying threat growth points and peak periods;
[0361] In addition, visualization tools are used to support interactive query and detailed analysis functions, providing deeper support for managers;
[0362] S85. The finally generated situation report includes dynamic alarm events, static asset and vulnerability information, risk level assessment and trend analysis charts, constituting a comprehensive and easy-to-understand network security situation report, providing a reliable basis for security operation and decision-making.
[0363] In an optional embodiment, the cleaning process of the collected raw data is as follows:
[0364] S91. Define a multi-factor weight calculation model, comprehensively consider the data time distribution, frequency characteristics and similarity characteristics, and dynamically generate a cleaning threshold. The formula is as follows:
[0365] ;
[0366] ;
[0367] ;
[0368] ;
[0369] ;
[0370] wherein, T represents the dynamic cleaning threshold; w t , w f and w s respectively represent the weights of the time distribution factor, the frequency feature factor, and the similarity feature factor, w t + w f + w s = 1; f t (x) represents the time distribution factor, measuring the deviation degree of the data point from the normal time interval; f f (x) represents the frequency feature factor, evaluating the frequency deviation degree of the data point in the statistical distribution; f s (x) represents the similarity feature factor, reflecting the content similarity of the data point with its neighboring points; t x represents the timestamp of the data point; represents the mean value of the timestamps of the same type of data; represents the standard deviation of the timestamps of the same type of data; f x represents the occurrence frequency of the data point; f max represents the maximum frequency in the entire dataset; cos(θ x ) represents the vector similarity between the data point and its neighboring data; represents the feature vector of the target data point x; represents the feature vector of the neighboring data; and respectively represent the vector modulus of the target point and the modulus of the feature vector of the neighboring data point;
[0371] S92. Define the cleaning rule:
[0372] ;
[0373] wherein, x represents the original value of the data point; μ represents the mean value of the same type of data; σ represents the standard deviation of the same type of data;
[0374] S93. Use an adaptive strategy to adjust the weights during the cleaning process:
[0375] ;
[0376] wherein, represents the weight of the time distribution factor in the (n + 1)-th round; Δ tIndicates the contribution gain of the time distribution factor to the cleaning result; α represents the adjustment rate, that is, the learning rate.
[0377] The embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited thereto, and various changes can be made without departing from the spirit of the present invention within the scope of knowledge possessed by those skilled in the art.
Claims
1. A network security situation generation method based on multi-view monitoring, characterized in that: The specific implementation steps include the following: S1, monitor and collect data from five perspectives: network traffic, malicious code, network performance, network routing and network logs; S2. Preprocess and standardize the collected data, and clean the collected raw data based on a dynamic threshold adaptive multi-factor cleaning method; S3, logically stratify the processed raw data, including network layer, application layer, device layer, user behavior layer and environment layer; S4: Integrate the data of the network layer, application layer, device layer, user behavior layer, and environment layer through association identifiers and time windows to build a unified association relationship. The integration process is as follows: S41, aligning the data layers that have unified timestamp and identifier fields after the data preprocessing stage, and establishing preliminary logical associations between records from different sources based on the common identifiers in the data; Network layer IP address net Login IP address of the user behavior layer user Aligned, and the device ID of the device layer, that is, DeviceID device , and the device ID recorded in the application layer, i.e. DeviceID app , to match, and build a preliminary identifier mapping relationship through the consistency of field values; S42. After the identifiers are aligned, the time difference problem between data from different sources is solved through time window matching. To ensure the consistency of the dynamic event chain, a reasonable time window ΔT is set to associate the records whose time difference is within the tolerance range; If the login event time T of the user behavior layer user and the traffic event time T of the network layer net Satisfy |T user -T net |≤ΔT, then the two records are considered to be temporally associated; S43, deeply associating the inter-layer data based on identifier alignment and time window matching; S5. Perform time series modeling on the fused multi-view data to capture the dynamic characteristics of data changing over time; S6, extract basic information, advanced features and dynamic features to provide input for model learning; S7. Use comprehensive features to build a joint learning model. In the model learning stage, a joint learning strategy of shared network and multi-task branches is adopted. Common features are extracted through the shared layer and combined with task-specific branches for detailed analysis to identify abnormal behaviors, detect unknown threats, and predict potential risks. S8. Generate network security situation based on statistical analysis and model prediction results. Combine the dynamic prediction results of model learning with the information generated by static analysis to form a network security situation. The generated situation information covers abnormal event identification and alarm, asset quantity and level distribution, vulnerability quantity and level distribution, risk assessment and level classification, and threat trend prediction.
2. A method for generating network security situation based on multi-view monitoring according to claim 1, characterized in that: The specific implementation steps of logical stratification are as follows: S21. After data preprocessing is completed, the data from network traffic monitoring, malicious code detection, network performance monitoring, network routing monitoring and network log monitoring are divided into network layer, application layer, device layer, user behavior layer and environment layer according to attributes and analysis objectives: S22. Network layer management and communication activity-related data, including network traffic patterns, protocol behaviors, and routing table updates. Network layer data is derived from network traffic monitoring and network routing monitoring, reflecting overall traffic characteristics and link operation status. S23, the application layer centrally stores data related to the application running status, including data from network performance monitoring and network log monitoring; S24. The device layer manages data related to the device operation status and security threats, including the operation logs, communication records and internal alarm information of IoT devices, and classifies the data from malicious code detection into the device layer to monitor the operation security and potential threats of the device; S25, the user behavior layer stores user identity and operation behavior data; S26, the environment layer includes external context information and environmental data; S27. Assign the data from different monitoring perspectives to their respective corresponding levels.
3. A method for generating network security situation based on multi-view monitoring according to claim 1, characterized in that: The association process of deeply associating data between layers is as follows: A1. Establish an association between the application layer and the user behavior layer. The log data of the application layer is matched with the login records and operation logs of the user behavior layer through the user ID, i.e., UserID, and the session ID, SessionID; A2. Establish an association between the device layer and the network layer through the device ID and IP address: The running status of the device layer is matched with the device ID of the traffic record of the network layer through the device ID, i.e., DeviceID, and meets the following association conditions: DeviceID device =DeviceID net ; IP device =IP net ; Where, IP device Indicates the IP address of the device layer; DeviceID net Indicates the device ID of the network layer; DeviceID device Indicates the device ID of the device layer; Afterwards, the temporal consistency of the association is established through time window matching; A3. The association between the network layer and the user behavior layer is established through the login IP address and network traffic records. The source IP in the login event of the user behavior layer, i.e., IP uesr , and the source IP of the network layer traffic record, i.e. IP net , to match, the following conditions are met: IP net =IP uesr ; A4. Through time window matching, establish associations between user operations and network behaviors in the time dimension; A5. The environment perception layer provides context information for all layers through the node identifier in the topology structure, namely NodeID, and the IP address in the external threat intelligence, namely IP env , establish associations with other layer data; Among them, malicious IP, that is, IP env Directly match the source IP in the network layer traffic, that is, IP net , the associated conditions are as follows: IP env =IP net ; A6. Through malicious identification DeviceID env Associated with the device layer, and then matched to the login record of the user behavior layer: NodeID env ∈{DeviceID,IP,UserID}; Where, NodeID env Indicates the malicious node identifier; Based on this, the threat source can be quickly identified and located, and the scope of threat impact can be determined through contextual information.
4. The method for generating network security situation based on multi-view monitoring according to claim 1, characterized in that: The construction process of time series modeling is as follows: S51, extracting event records from the result of cross-layer data association, and generating a time series in chronological order; The output of cross-layer data association contains a set of event records R assoc , each record r i Includes the timestamp T of the event i , comprehensive feature vector F i , and the associated identifier ID i : R assoc ={r1,r2,…,r n },r i =(T i ,F i ,ID i ); Among them, the comprehensive feature vector F i =[f net , f app , f device , f user , f env ] is a multi-level feature generated in the cross-layer association process, according to the timestamp T i Sort a collection of records to generate a global time series: S={(T1,F1),(T2,F2),…,(T n ,F n )},T1≤T2≤…≤T n ; S52. Based on the generated time series, a time series model is constructed to extract dynamic characteristics. The time series is first divided into several subsequences according to a fixed time window ΔT. Each time window contains all events within its range. For each time window [T t , T t +ΔT], whose subsequence is defined as: S t ={(T i ,F i )}|T t ≤T i <T t +ΔT; Among them, T t Indicates the start time of the time window, T t +ΔT is the end time, subsequence S t Contains all events that meet the conditions; S53, after the time window segmentation is completed, generate a comprehensive feature sequence X in each time window t , the representation of the feature sequence is as follows: X t =[F i ,F j ,…,F k ],i,j,k∈S t ; Among them, F i It is the comprehensive feature vector of events in the time window, arranged in the chronological order of the events.
5. The method for generating network security situation based on multi-view monitoring according to claim 1 is characterized in that: The extraction process of basic information, advanced features and dynamic features is as follows: S61, extracting features from the results of cross-layer data association and dynamic analysis; Feature extraction is divided into three parts: extracting basic information directly from the processed raw data; extracting time series features and behavior pattern features from the built time series module; extracting statistical features, association relationship features, graph features, context features and resource features from cross-layer data associations to construct a comprehensive feature vector F; S62, after feature extraction is completed, prepare training data and labels; S63. For the abnormal event identification and alarm task, read the historical attack log and the alarm record of the intrusion detection system to identify whether there is abnormal behavior. The label of the normal behavior sample is defined as label=0, and the label of the abnormal behavior sample is defined as label=1; Define the label of the risk assessment task as the overall risk score and the corresponding risk level y; Rating Tags The risk level y is discretized based on the impact and severity of historical events, including: Low risk, i.e. y=1; Medium risk, i.e. y=2; High risk, i.e. y=3; S64, define a label, the label is defined as threat type y threat and the time series target value T trend ; Threat Type Label threat Classification and annotation based on historical threat types: y threat ∈{1,2,…,C}; Where C represents the number of threat categories; Time series target value T trend Based on the time series data of historical threat frequency, it is annotated through trend analysis algorithm; S65. After the label generation is completed, the data is divided into training set, validation set and test set according to task requirements.
6. A method for generating network security situation based on multi-view monitoring according to claim 5, characterized in that: The construction process of the comprehensive feature vector F is as follows: S71. Basic Information F basic , the basic features extracted directly from the raw data, including: The basic elements of network communication are recorded as f ip_port ; Network protocol type, recorded as f protocol ; The distribution of packet sizes f packet_size ; Flow directionf traffic_dir ; Based on this, we can get feature F basic =[f ip_port , f protocol , f packet_size , f traffic_dir ]; S72, time series feature F time ,The dynamic characteristics extracted from time series modeling,time series features include: The peak flow time f time_peak ; The distribution of time intervals between events f time_interval ; The standard deviation f of the operating frequency within the time window time_frequency ; Based on this, we can get feature F time =[f time_peak , f time_interval , f time_frequency ]; S73, Behavior Pattern Characteristics F behavior , extracting the causal relationship between dynamic events, its specific features include: The path length f in the attack link behavior_path ; Node Dependencyf behavior_dependency ; Key behavior trigger probability f behavior_trigger ; Based on this, we can get feature F behavior =[f behavior_path , f behavior_dependency , f behavior_trigger ]; S74, Statistical Features F stat , extract the distribution and aggregation characteristics of the data, including: The mean packet size f mean_packet ; The variance of the total flow rate f var_traffic ; The count of abnormal events within a specific time window f count_anomaly ; Based on this, we can get feature F stat =[f mean_packet , f var_traffic , f count_anomaly ]; S75, association relationship feature F assoc , that is, the association patterns between different data layers, including: Frequent patterns of user behavior and device state changes user_device ; Conditional probability f of API calls and traffic anomalies API_traffic ; The frequency of interaction between specific resources f resource_freq ; Based on this, we can get the feature F assoc =[f user_device , f API_traffic , f resource_freq ]; S76, graph feature F graph , a graph structure generated based on cross-layer associations, with specific features including: The degree of the node f graph_degree ; The average clustering coefficient f of the graph graph-cluster ; Distribution of the shortest paths between nodes f path_dist ; Based on this, we can get feature F graph =[f graph_degree , f graph-cluster , f path_dist ]; S77, context feature F context , providing a global view of network context and external threats, with specific features including: Whether the device hits the external threat intelligence database threat_match ; The frequency of dynamic changes in network topology f topology ; Physical parameters of the equipment operating environment environment ; Based on this, we can get the feature F context =[f threat_match , f topology , f environment ]; S78, Resource Characteristics F resource , extract the resource distribution and status in the network, including: Number of devices resource_count ; The priority level of the core device is f resource_priority ; The current running status of the resource resource_status ; Based on this combination, we get F resource =[f resource_count , f resource_priority , f resource_status ]; S79, standardize the numerical features, use one-hot encoding for categorical features, and use TF-IDF for vectorization of text features; Get the comprehensive feature vector F=[F basic , F time , F behavior , F stat , F assoc , F graph , F context , F resource ].
7. A method for generating network security situation based on multi-view monitoring according to claim 6, characterized in that: The construction process of the joint learning model is as follows: S81, input comprehensive feature vector F = [F basic , F time , F behavior , F stat , F assoc , F graph , F context , F resource ], the features are standardized and encoded, the numerical features are mapped to the [0,1] interval through normalization operations, and the categorical features are represented by one-hot encoding: ; in, represents the features after normalization; f i Represents the original vector of the i-th feature; After processing, all features are merged into the feature matrix X∈R N×p ; Among them, N is the number of samples, p is the feature dimension; S82, the shared representation layer is responsible for extracting task-independent general information from the input features and generating a shared representation H shared , high-dimensional feature mapping is performed on the input feature matrix X through the fully connected layer: H (1) =σ(W (1) X+b (1) ); Among them, W (1) ∈R N×p represents the weight matrix, N is the number of samples, and p is the feature dimension; S83, introduce the multi-head attention mechanism, and strengthen the correlation between features by calculating the attention weight. The output feature after the multi-attention mechanism is H shared ; The attention score a of each head ij The calculation is as follows: ; Among them, score(h i ,h j ) represents the feature h i and h j Similarity measure between ; The output of multi-head attention is: H (2) =Concat(head1,head2,…,head M )W0; Where M represents the number of attention heads; W0 represents the projection matrix; S84. Based on the shared representation layer, an independent branch network is designed for each task to model and optimize specific tasks: S8401, abnormal event identification and alarm: input is shared to indicate H shared and time series features F time , combined with the behavioral pattern characteristics F behavior , that is, input feature: H event =[H shared , F time , F behavior ]; Then, a hybrid model is used to capture time dependency and causality, and the anomaly detection output is: y event =σ(W event H event +b event ); In the formula, y event Indicates whether the event is abnormal, y event ∈[0,1];W event and b event Model parameters that represent the model's abnormal event recognition and alarm; S8402, Risk Assessment and Classification: Using Shared Representation H shared , Statistical features F stat , contextual features F context and resource characteristics F resource Perform risk scoring and classification to obtain the risk assessment input feature H risk =[H shared ,F stat ,F context ,F resource ], the shared representation input is: R=W risk H risk +b risk ; Where R represents the risk score; W risk and b risk The model parameters representing the risk assessment of the model; Risk level classification is completed through the Softmax activation function: ; Where k represents the risk level; W cls,k represents the weight of the k-th risk level; S8403, Threat Trend Forecast: Combined with Shared Representation H shared and time series features F time and the association feature F assoc , that is, the threat trend prediction input feature H trend =[H shared , F time , F assoc ], predict and analyze the evolution trend of threats in time and space; Capturing trend changes through the time series modeling layer GRU: trend =GRU(H trend ), and predict the probability of threat events and generate trend graphs; S8404. The total loss function of joint learning combines the independent losses of each task and is defined as follows: L=λ event L event +λ risk L risk +λ trend L trend ; Among them, λ event , risk and λ trend They represent the task weights of abnormal event identification, risk assessment and trend analysis tasks respectively; L event , L risk and L trend They represent the losses of abnormal event identification, risk assessment, and trend analysis tasks respectively.
8. The method for generating network security situation based on multi-view monitoring according to claim 1, characterized in that: The monitoring process of collecting data from five perspectives, namely network traffic, malicious code, network performance, network routing and network logs, is as follows: S91. In the initial stage of network security situation generation, multi-perspective monitoring and data collection captures all-round information about the network environment from multiple dimensions, taking network traffic monitoring, malicious code detection, network performance monitoring, network routing monitoring and network log monitoring as the entry points, and comprehensively acquires raw data covering network communications, application behaviors, device status, user activities and environmental dynamics; S92. In network traffic monitoring, deploy traffic collection tools to record the traffic patterns and communication details of data packets in real time, including source IP, destination IP, protocol type and packet size; S93. Malicious code detection relies on host-level and network-level monitoring tools to detect file system changes, abnormal process activities, and malicious behavior samples, and assists in analyzing malicious activities by collecting feature data; S94. Network performance monitoring uses a distributed monitoring platform to collect link status, delay time, packet loss rate and throughput, revealing the dynamic changes and bottlenecks of network performance; S95, Network routing monitoring collects routing table updates, topology changes and link anomalies through protocol monitoring tools to form a real-time view of network dynamic changes; S96. Network log monitoring focuses on the operation logs of systems and devices, and collects event records through centralized log management tools, including user login information, permission changes and security events.
9. The method for generating network security situation based on multi-view monitoring according to claim 1, characterized in that: The cleaning process of the collected raw data is as follows: S101. Define a multi-factor weight calculation model, comprehensively consider the data time distribution, frequency characteristics and similarity characteristics, and dynamically generate a cleaning threshold. The formula is as follows: ; ; ; ; ; Where T represents the dynamic cleaning threshold; w t 、w f and w s They represent the weights of time distribution factor, frequency characteristic factor and similarity characteristic factor respectively, w t +w f +w s =1; f t (x) represents the time distribution factor, which measures the degree of deviation of the data point from the normal time interval; f f (x) represents the frequency characteristic factor, which evaluates the frequency deviation of the data point in the statistical distribution; f s (x) represents the similarity feature factor, which reflects the content similarity between the data point and its neighboring points; t x The timestamp representing the data point; Indicates the timestamp mean of the same type of data; Indicates the timestamp standard deviation of the same type of data; f x Indicates the frequency of occurrence of data points; f max Represents the maximum frequency in the entire data set; cos(θ x ) represents the vector similarity between a data point and its neighboring data; The feature vector representing the target data point x; Feature vector representing neighboring data; and Represent the vector modulus of the target point and the modulus of the feature vector of the neighboring data points respectively; S102. Define cleaning rules: ; In the formula, x represents the original value of the data point; μ represents the mean of the same type of data; σ represents the standard deviation of the same type of data; S103, use adaptive strategy to adjust weights during cleaning: ; In the formula, represents the time distribution factor weight of the n+1th round; Δ t It represents the contribution gain of the time distribution factor to the cleaning result; α represents the adjustment rate, that is, the learning rate.
Citation Information
Patent Citations
Network security detection method and system
CN118101250A
Network security comprehensive protection system based on deep learning
CN118890187A