Online information security comprehensive analysis and monitoring method and system
By performing structured processing and feature extraction of network traffic data, combined with protocol state transition entropy values and multi-source risk propagation graphs, the problems of poor protocol compatibility and inaccurate risk assessment in existing technologies are solved, achieving high-precision anomaly detection and risk management.
Patent Information
- Application Number
- CN202511543946.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-28
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2045-10-28
AI Technical Summary
Existing technologies cannot accurately match the security behavior baselines of different protocols in online information security analysis and monitoring, resulting in high false alarm or false negative rates. Furthermore, risk assessment models fail to accurately reflect the overall network security situation, affecting the relevance and effectiveness of response commands.
By structuring network traffic data, extracting time-frequency-protocol domain features, combining protocol state transition entropy values for dynamic anomaly detection, constructing a multi-source risk propagation graph and conducting a comprehensive risk assessment, generating risk response instructions and updating anomaly monitoring thresholds, and forming a comprehensive analysis report.
It significantly improves the accuracy and environmental adaptability of anomaly detection, reduces false alarm and false negative rates, enables precise control of cybersecurity risks and targeted response instructions, and ensures refined management of cybersecurity risks.
Smart Images

Figure CN121012701A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of information monitoring technology, and in particular to a method and system for comprehensive analysis and monitoring of online information security. Background Technology
[0002] In current online information security analysis and monitoring methods, anomaly detection thresholds mostly adopt fixed values or general dynamic adjustment methods, which lack adaptability to protocol characteristics. Existing technologies do not optimize thresholds in conjunction with protocol state transition rules, but only adjust detection standards based on the general statistical characteristics of traffic data. This leads to high false alarm rates or missed alarms when protocol behavior is complex and variable. For example, when faced with frequent state transitions of the TCP protocol or the connectionless nature of the UDP protocol, fixed thresholds cannot accurately match the security behavior baselines of different protocols, making it difficult to effectively distinguish between normal protocol fluctuations and abnormal attack behavior. This reduces the accuracy and anti-interference capability of anomaly detection and fails to meet the security monitoring needs of diverse protocol scenarios.
[0003] Meanwhile, existing risk assessment models often simply overlay abnormal traffic data or ignore the correlation between network topology and protocol type when calculating comprehensive risk. They fail to construct an effective risk propagation analysis mechanism. Most methods assess risk based only on the anomaly level of a single node, without considering the propagation path of risk between different network devices and the attenuation effect of protocol type on propagation intensity. This results in risk scores failing to accurately reflect the overall network security situation. For example, they do not set differentiated risk propagation attenuation coefficients for different protocols such as TCP and UDP, making it impossible to quantify the diffusion effect of risk in different protocol links. This leads to a large discrepancy between risk level classification and actual threat level, which in turn affects the targeting and effectiveness of subsequent response instructions, making it difficult to achieve precise control over network security risks. Therefore, meeting the security monitoring needs under diverse protocol scenarios and achieving precise control over network security risks has become an urgent problem to be solved. Summary of the Invention
[0004] This invention provides a method and system for comprehensive analysis and monitoring of online information security, in order to solve the problems mentioned in the background art.
[0005] To achieve the above objectives, the present invention provides a comprehensive online information security analysis and monitoring method, comprising: S1, perform structured processing on network traffic data to obtain structured data; S2, perform time-frequency-protocol domain feature extraction on the structured data to obtain a structured feature vector containing protocol-specific indicators; S3, dynamically detect anomalies in the structured feature vector using a preset anomaly detection threshold to obtain anomaly identification data for the network traffic data; S4. Based on the anomaly identification data and combined with the network topology data, calculate the comprehensive risk score, classify the risk level of the comprehensive risk score, and obtain the risk level of the network traffic data. S5, perform response instruction mapping on the risk level to obtain the risk response instruction for the network traffic data; S6. Based on the risk response instruction, the dynamic anomaly detection process is updated through the performance index feedback mechanism to obtain the updated anomaly monitoring threshold. S7. Construct a comprehensive network information security analysis report based on the risk level, the risk response instruction, and the updated anomaly monitoring threshold.
[0006] In a preferred embodiment, the step of structuring network traffic data to obtain structured data includes: The network traffic data is denoised to obtain initial network traffic data; The initial network traffic data is parsed using multiple protocols to obtain the protocol parsing data of the network traffic data; The protocol parsing data is serialized to obtain structured data of the network traffic data.
[0007] In a preferred embodiment, the step of extracting time-frequency-protocol domain features from the structured data to obtain a structured feature vector containing protocol-specific indicators includes: Temporal features are extracted from the structured data to obtain temporal feature vectors; Frequency domain features are extracted from the structured data to obtain frequency domain feature vectors; Protocol domain features are extracted from the structured data by calculating the protocol state transition entropy value to obtain a protocol domain feature vector. The mathematical expression of the Shannon entropy formula used to calculate the protocol state transition entropy is as follows: ; In the formula, This is the protocol state transition entropy value. Protocol status The probability of occurrence in the state sequence For the first Each protocol status, For protocol index, This represents the total number of protocol states. The time-domain feature vector, the frequency-domain feature vector, and the protocol-domain feature vector are fused to obtain the structured feature vector.
[0008] In a preferred embodiment, the step of dynamically detecting anomalies in the structured feature vector using a preset anomaly detection threshold to obtain anomaly identification data for the network traffic data includes: The preset anomaly monitoring threshold is dynamically adjusted based on the protocol state transition entropy value to obtain the dynamic anomaly monitoring threshold. The structured feature vector is used to detect anomalies using the dynamic anomaly monitoring threshold to obtain anomaly identification data for the network traffic data.
[0009] In a preferred embodiment, the step of calculating a comprehensive risk score based on anomaly identification data and network topology data includes: Obtain network topology data; Based on the anomaly identification data and the network topology data, a directed risk propagation graph of the network traffic data is constructed; Based on the directed risk propagation graph, a breadth-first search logic is used to calculate the node risk value, wherein the mathematical expression of the breadth-first search logic is as follows: ; In the formula, For nodes The risk value, This is the node identifier of the current node. Adjacent nodes The risk value, This serves as the node identifier for adjacent nodes. Protocol type The corresponding attenuation factor, Protocol type, For nodes The set of adjacent nodes; The overall risk score of the node risk value is calculated by summing all the node risk values.
[0010] In a preferred embodiment, the step of classifying the risk level of the comprehensive risk score to obtain the risk level of the network traffic data includes: The comprehensive risk score is compared at multiple levels based on a preset risk level threshold to obtain the threshold comparison result. Based on this, a corresponding risk level identifier is assigned to the threshold comparison result through a risk level mapping rule to obtain the risk level of the network traffic data.
[0011] In a preferred embodiment, the step of mapping the risk level to a response instruction to obtain the risk response instruction for the network traffic data includes: The risk level is mapped to a response instruction to obtain the operation instructions for the network traffic data; The operation instructions are processed for protocol adaptation to obtain risk response instructions for the network traffic data.
[0012] In a preferred embodiment, the step of updating the dynamic anomaly detection process based on the risk response instruction through a performance indicator feedback mechanism to obtain the updated anomaly monitoring threshold includes: Based on the risk response instructions, false alarm events and alarm events are extracted; Based on the false alarm events and the alarm events, calculate the total number of false alarms and the total number of alarms; Calculate the false alarm rate based on the total number of false alarms and the total number of alarms; Based on the false alarm rate, the dynamic anomaly monitoring threshold is updated to obtain the updated dynamic anomaly monitoring threshold.
[0013] In a preferred embodiment, the step of constructing a comprehensive network information security analysis report based on the risk level, the risk response instruction, and the updated anomaly monitoring threshold includes: Integrate the risk level, the risk response instructions, and the updated anomaly monitoring thresholds to generate comprehensive report data; The comprehensive report data is input into a preset report format template for report formatting to generate a comprehensive network information security analysis report.
[0014] To address the aforementioned problems, this invention also provides an online information security comprehensive analysis and monitoring system. The system includes: a data structuring processing module, a structured feature extraction module, a dynamic anomaly monitoring module, a comprehensive risk assessment module, an intelligent response decision-making module, a data monitoring feedback module, and a comprehensive analysis report generation module, wherein: The data structuring module is used to perform structuring processing on network traffic data to obtain structured data; The structured feature extraction module is used to extract time-frequency-protocol domain features from the structured data to obtain a structured feature vector containing protocol-specific indicators. The dynamic anomaly monitoring module is used to perform dynamic anomaly detection on the structured feature vector through a preset anomaly monitoring threshold to obtain anomaly identification data of the network traffic data. The comprehensive risk assessment module is used to calculate a comprehensive risk score based on anomaly identification data and network topology data, classify the risk level of the comprehensive risk score, and obtain the risk level of the network traffic data. The intelligent response decision module is used to perform response instruction mapping on the risk level to obtain the risk response instruction for the network traffic data; The data monitoring and feedback module is used to update the dynamic anomaly detection process based on the risk response instruction through a performance indicator feedback mechanism, so as to obtain the updated anomaly monitoring threshold. The comprehensive analysis report generation module is used to construct a comprehensive network information security analysis report based on the risk level, the risk response instruction, and the updated anomaly monitoring threshold.
[0015] Compared with the prior art, the present invention has the following beneficial effects: 1. In this invention, the dynamic anomaly monitoring threshold adjustment mechanism based on protocol state transition entropy significantly improves the accuracy and environmental adaptability of anomaly detection. This mechanism calculates the protocol state transition entropy value using the Shannon entropy formula, deeply binding the threshold to the randomness of protocol behavior. When the protocol state exhibits high entropy, the threshold is automatically widened to avoid normal fluctuations being misjudged as anomalies. When the protocol state is in a low entropy state, the threshold is tightened to enhance the capture of subtle abnormal behaviors. Compared with traditional fixed thresholds or general dynamic thresholds, this mechanism can accurately match the security behavior baseline of different protocols, effectively reducing false alarm and false negative rates. At the same time, it enhances the adaptability to diverse protocol scenarios in complex network environments, ensuring that anomaly detection neither misses potential threats nor interferes with normal network communication.
[0016] 2. In this invention, a protocol-aware multi-source risk propagation graph and a multi-layer attenuation model are constructed. This model consists of three key elements: model structure (directed risk propagation graph), core mechanism (protocol-related attenuation factors), and calculation process (breadth-first search traversal). Together, these elements enable the attenuation-based propagation calculation of risks across multiple nodes in the network topology, significantly improving the comprehensiveness and accuracy of risk assessment and providing a reliable basis for subsequent response decisions. The model first constructs a directed risk propagation graph by combining anomaly identification data and network topology data, clearly presenting the propagation path of risks among network devices. Then, by introducing differentiated attenuation factors based on protocol types, it uses breadth-first search logic to calculate the risk value of each node, accurately quantifying the diffusion and attenuation effect of risks in different protocol links. Compared to traditional risk assessment methods that simply superimpose abnormal data or ignore the impact of protocols, this model can fully integrate multi-source information such as network topology, protocol characteristics, and abnormal data, making the comprehensive risk score more consistent with the overall network security situation, and the risk level classification highly matched with the actual threat level. This ensures that subsequent risk response instructions are more targeted, achieving refined management of network security risks. Attached Figure Description
[0017] Figure 1This is a flowchart illustrating a comprehensive online information security analysis and monitoring method according to an embodiment of the present invention. Figure 2 A functional module diagram of an online information security comprehensive analysis and monitoring system provided in an embodiment of the present invention; The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0018] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0019] This application provides a method for comprehensive analysis and monitoring of online information security. The executing entity of this method includes, but is not limited to, at least one of the following electronic devices that can be configured to execute the method provided in this application: a server, a terminal, etc. In other words, the method for comprehensive analysis and monitoring of online information security can be executed by software or hardware installed on a terminal device or a server device. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.
[0020] Reference Figure 1 The diagram shown is a flowchart illustrating a comprehensive online information security analysis and monitoring method according to an embodiment of the present invention. In this embodiment, the comprehensive online information security analysis and monitoring method includes: S1, perform structured processing on network traffic data to obtain structured data; In this embodiment of the invention, the step of performing structured processing on network traffic data to obtain structured data includes: The network traffic data is denoised to obtain initial network traffic data; The initial network traffic data is parsed using multiple protocols to obtain the protocol parsing data of the network traffic data; The protocol parsing data is serialized to obtain structured data of the network traffic data.
[0021] It should be noted that the essence of noise reduction is to balance processing latency and noise detection accuracy by setting the window size to 1024 bytes, and to identify the noise processing method by comparing the entropy value of the window with the entropy threshold. If the entropy value is lower than the threshold, it is marked as noise and filtered.
[0022] Furthermore, the denoising process is based on a sliding window noise filtering method. This method divides the data into fixed-size windows using a sliding window and calculates the entropy value of each window. Low-entropy windows are then filtered out as noise data based on an entropy threshold. The mathematical expression of the Shannon entropy formula used for entropy calculation is as follows: ; In the formula, The entropy value of the window. Byte value The probability of appearance in the window. For a range of byte values, For byte index, where the byte value ranges from 0 to 255.
[0023] Furthermore, the entropy value of the window is calculated using the Shannon entropy formula, representing the degree of randomness of the data byte values within the window. Essentially, it is used to quantify the uncertainty and information content of the data. In denoising, the entropy value is used to identify noise windows. Low entropy values indicate that the data may contain repetitive or regular patterns, and thus it is filtered to improve data quality.
[0024] Furthermore, the window size is set to 1024 bytes, a value determined by balancing processing latency with noise detection accuracy to optimize computational efficiency in real-time network traffic analysis.
[0025] Furthermore, the entropy threshold is set based on the statistical characteristics of historical network traffic data. It is calibrated experimentally to distinguish between normal and noisy data. Specifically, the entropy threshold is set according to the entropy distribution of typical protocol traffic. Low-entropy windows are marked as noise. The entropy distribution of typical protocol traffic is set by collecting 30 days of historical network traffic data, analyzing the window entropy distribution of different protocol sessions, and calculating the 25th percentile statistical characteristics to distinguish between normal and noisy data. Low-entropy windows are marked as noise and filtered, thereby optimizing the accuracy and efficiency of noise reduction processing.
[0026] Furthermore, the probability of a byte value appearing in the window is obtained by calculating the ratio of the number of times each byte value appears to the total number of bytes in the window.
[0027] It should be noted that multi-protocol parsing separates different protocol sessions by parsing the protocol fields in network traffic data. This is used to decompose mixed protocol traffic and provide a protocol-specific data foundation for subsequent feature extraction.
[0028] It should be noted that protocol parsing data is structured protocol information obtained from multi-protocol parsing. It is a binary representation of the protocol header or payload, used to quantify protocol behavior and provide input for time-frequency-protocol domain feature extraction.
[0029] It should be noted that serialization is the process of converting protocol-parsed data into a standard byte sequence format. Through encoding operations, the data is made easier to store and transmit, and the data format is standardized to ensure the consistency of structured data.
[0030] It should be noted that structured data includes protocol parsing data consisting of protocol type and session identifier, serialization fields consisting of timestamp and data length, and noise-reduced traffic bytes. Essentially, it is a high-dimensional vector data set that is easy for computers to process.
[0031] Furthermore, network traffic data refers to the original transmission data units and their derivative data that carry information such as source, destination, content, timing and behavior in digital communication networks. It includes three levels. The first level is the data packet, which is the most primitive communication unit and includes a header and a payload. The second level is the flow, which is a logical collection of multiple data packets with the same key attributes. It is the most commonly used data abstraction in traffic analysis. The third level consists of metadata and derived data, which is information further calculated and extracted from the original data packets and streaming data.
[0032] S2, perform time-frequency-protocol domain feature extraction on the structured data to obtain a structured feature vector containing protocol-specific indicators; In this embodiment of the invention, the step of extracting time-frequency-protocol domain features from the structured data to obtain a structured feature vector containing protocol-specific indicators includes: Temporal features are extracted from the structured data to obtain temporal feature vectors; Frequency domain features are extracted from the structured data to obtain frequency domain feature vectors; Protocol domain features are extracted from the structured data by calculating the protocol state transition entropy value to obtain a protocol domain feature vector. The mathematical expression of the Shannon entropy formula used to calculate the protocol state transition entropy is as follows: ; In the formula, This is the protocol state transition entropy value. Protocol status The probability of occurrence in the state sequence For the first Each protocol status, For protocol index, This represents the total number of protocol states. The time-domain feature vector, the frequency-domain feature vector, and the protocol-domain feature vector are fused to obtain the structured feature vector.
[0033] It should be noted that temporal feature extraction is a method of obtaining features by analyzing the statistical characteristics of network traffic data in the time dimension. It is determined by calculating the average time interval of data packet arrival. It is used to capture the changing patterns of traffic over time, provide time-related behavioral features for anomaly detection, and enhance the ability to perceive dynamic changes in network traffic.
[0034] It should be noted that frequency domain feature extraction is a method of obtaining features by converting time-domain signals into frequency-domain representations. By using fast Fourier transform to calculate the frequency components of traffic data, the periodicity of traffic is represented, which provides a basis for detecting periodic anomalies and improves the accuracy of identifying hidden frequency features.
[0035] It should be noted that protocol domain feature extraction refers to the process of calculating the protocol state transition entropy value based on the protocol type field in the structured data to obtain the protocol domain feature vector, which is essentially a quantitative representation of the randomness of protocol behavior.
[0036] It should be noted that the protocol state transition entropy value represents the degree of randomness of protocol state transitions. High entropy values indicate complex protocol behaviors such as frequent state transitions, while low entropy values indicate stable protocol behaviors such as fixed state sequences. It is used to provide a protocol context-aware indicator for dynamic threshold adjustment and optimize the sensitivity of anomaly detection.
[0037] Furthermore, the protocol state is obtained by parsing the protocol fields in the structured data, specifically including the states in the protocol state machine, such as the TCP protocol states SYN_SENT, ESTABLISHED, FIN_WAIT, etc. These states are obtained by parsing the protocol header or state transition sequence and are used to construct a state sequence to calculate the entropy value.
[0038] Furthermore, the protocol domain feature vector is a vector composed of protocol state transition entropy values. It is a feature representation of protocol behavior and is essentially a numerical representation of the randomness of protocol states. It is used to contribute protocol-specific information in feature fusion and to provide protocol-dimensional features for structured feature vectors.
[0039] Furthermore, the total number of protocol states is predefined based on the protocol type. For example, for the TCP protocol, the total number of states is 11, which is the standard TCP state number. Different protocols have different total numbers of states, which are determined by the protocol specification. This ensures that the entropy calculation is tailored to the behavior pattern of a specific protocol, thereby improving the accuracy of feature extraction.
[0040] It should be noted that feature fusion is based on the time-domain feature vector, the frequency-domain feature vector, and the protocol-domain feature vector. The feature vectors are concatenated into a single high-dimensional vector through vector concatenation, and principal component analysis is used to reduce the dimensionality of the concatenated high-dimensional vector to obtain the structured feature vector. This process is used to reduce redundancy and improve feature representation efficiency, ensuring that the structured feature vector balances computational performance and feature integrity.
[0041] Furthermore, principal component analysis (PCA) reduces dimensionality by calculating the eigenvalues of the covariance matrix of each eigenvector and retaining some principal components. This ensures that the structured eigenvectors balance computational efficiency and feature integrity. The number of principal components retained is determined by calculating the eigenvalues of the covariance matrix and selecting the number of principal components with a cumulative contribution rate of 95%. This ensures that most of the information is retained after dimensionality reduction, balancing computational efficiency and feature integrity, and avoiding information loss.
[0042] It should be noted that structured feature vectors are low-dimensional vectors obtained after feature fusion and dimensionality reduction. They contain compressed features in the time domain, frequency domain, and protocol domain. Essentially, they are a comprehensive feature representation of network traffic data, used to provide input for anomaly detection and achieve efficient and accurate security analysis.
[0043] S3, dynamically detect anomalies in the structured feature vector using a preset anomaly detection threshold to obtain anomaly identification data for the network traffic data; In this embodiment of the invention, the step of dynamically detecting anomalies in the structured feature vector using a preset anomaly monitoring threshold to obtain anomaly identification data for the network traffic data includes: The preset anomaly monitoring threshold is dynamically adjusted based on the protocol state transition entropy value to obtain the dynamic anomaly monitoring threshold. The structured feature vector is used to detect anomalies using the dynamic anomaly monitoring threshold to obtain anomaly identification data for the network traffic data.
[0044] It should be noted that the preset anomaly monitoring threshold is an initial threshold pre-set based on the statistical characteristics of historical network traffic data. It is used to distinguish between normal and abnormal traffic, and represents the upper limit of normal fluctuations in network traffic behavior. When the feature vector exceeds this threshold, it may indicate an anomaly. It is used to provide a benchmark reference for anomaly detection in the initial stage, ensuring that the detection process is based on evidence and improving the stability and reliability of the detection system.
[0045] Furthermore, the preset anomaly monitoring threshold is based on 30 days of historical network traffic data. Structured feature vectors are extracted, and the 95th percentile of these vectors is calculated to obtain a statistical distribution. This statistical distribution is used as the preset anomaly monitoring threshold to improve the stability and reliability of anomaly detection.
[0046] It should be noted that the dynamic threshold adjustment is based on the protocol state transition entropy value and the preset anomaly monitoring threshold. The dynamic anomaly monitoring threshold is generated by multiplication. Specifically, it is obtained by multiplying the preset anomaly monitoring threshold by 1 and the sum of the protocol state transition entropy values. This is used to adaptively adjust the detection sensitivity according to the degree of randomness of the protocol state transition.
[0047] Furthermore, when the protocol state transition entropy value is high, it indicates that the protocol behavior is complex and variable, and the threshold is automatically relaxed to reduce false alarms; When the entropy value is low, it indicates that the protocol behavior is stable, and the threshold is tightened to improve detection accuracy; By adjusting thresholds in real time, anomaly detection can better reflect the actual behavior of current network traffic, improving the accuracy and adaptability of detection while reducing the risk of misjudgment due to changes in the network environment.
[0048] It should be noted that the dynamic anomaly monitoring threshold is a real-time threshold that has been dynamically adjusted. It represents the critical value used for anomaly detection under the current network state and reflects the dynamic detection boundary based on the uncertainty of protocol state. It fluctuates with changes in protocol behavior and is used to provide a dynamic benchmark for anomaly detection, ensuring that the detection process can respond to real-time changes in network traffic and optimize the timeliness and effectiveness of anomaly identification.
[0049] It should be noted that anomaly detection compares structured feature vectors with dynamic anomaly monitoring thresholds to identify feature points that exceed the thresholds as anomalies. This is used to mark potential security threats and abnormal behaviors in network traffic, providing input data for subsequent risk assessments and helping to promptly detect potential network attacks or faults.
[0050] It should be noted that the anomaly identification data includes information such as the location of the anomaly point, timestamp, protocol type, feature vector value, and degree of anomaly. It represents detailed data of the detected abnormal event and is essentially a quantitative identifier of abnormal behavior in network traffic. It is used to provide basic data for the comprehensive risk assessment module, participate in the calculation of risk scores and the generation of response instructions, and support subsequent source tracing analysis and report generation.
[0051] S4. Based on the anomaly identification data and combined with the network topology data, calculate the comprehensive risk score, classify the risk level of the comprehensive risk score, and obtain the risk level of the network traffic data. In this embodiment of the invention, the calculation of a comprehensive risk score based on anomaly identification data and network topology data includes: Obtain network topology data; Based on the anomaly identification data and the network topology data, a directed risk propagation graph of the network traffic data is constructed; Based on the directed risk propagation graph, a breadth-first search logic is used to calculate the node risk value, wherein the mathematical expression of the breadth-first search logic is as follows: ; In the formula, For nodes The risk value, This is the node identifier of the current node. Adjacent nodes The risk value, This serves as the node identifier for adjacent nodes. Protocol type The corresponding attenuation factor, Protocol type, For nodes The set of adjacent nodes; The overall risk score of the node risk value is calculated by summing all the node risk values.
[0052] It should be noted that obtaining network topology data is a process of actively collecting network device connection relationships through network scanning tools and network management protocols. Specific methods include using the SNMP protocol to poll network devices to obtain interface information, using the NetFlow protocol to collect traffic statistics, and combining network device configuration information to extract the physical and logical connection relationships between devices.
[0053] Furthermore, the acquired data includes network device identifiers consisting of IP addresses and MAC addresses, device type data labeling routers, switches, and firewalls, connection port information, link bandwidth capacity, and adjacency tables between devices. These data together constitute a complete description of the network topology, providing basic structural information for subsequent risk propagation analysis.
[0054] It should be noted that the construction of the directed risk propagation graph is based on the anomaly identification data and network topology data, mapping network devices and sessions as nodes, and mapping risk propagation paths as directed edges to generate the directed risk propagation graph, where nodes represent network devices or session entities, and edges represent the direction of risk propagation.
[0055] Furthermore, the essence of a directed risk propagation graph is to utilize the connectivity relationships in network topology data and the anomaly points in anomaly identification data, and store node relationships through an adjacency matrix, thus providing a graph structure foundation for risk propagation calculation.
[0056] It should be noted that the directed risk propagation graph is a graph structure model built based on network topology data and anomaly identification data. It is used to map network devices as nodes and the risk propagation path between devices as directed edges. It vividly represents the potential spread path and impact range of network security risks in the network. The node weight represents the risk level of the device, and the edge weight represents the intensity of risk propagation. This graph structure makes the originally abstract risk propagation process computable and visualized, providing a mathematical model basis for quantitative analysis of the chain reaction of risks in the network.
[0057] Furthermore, the edge weight, or the attenuation factor corresponding to the protocol type, is a predefined coefficient based on the protocol type, representing the attenuation strength of the risk propagating along the connection path.
[0058] Node weight is the node risk value. The initial value of the node risk value comes from the anomaly identification data. Specifically, when constructing a directed risk propagation graph, if a node is marked as having an anomaly by the anomaly monitoring module, then the node will be assigned an initial risk value. Nodes in the graph that are not directly marked as an anomaly but are connected to other anomaly nodes have their initial risk value set to 0.
[0059] It should be noted that the risk value of each node represents the potential risk level of network devices or sessions. It is generated by accumulating the risk values of adjacent nodes and the attenuation factor, and is used to quantify the diffusion effect of risk in the topology.
[0060] It should be noted that the node risk value is a quantitative indicator calculated on the directed risk propagation graph using a graph traversal algorithm. It represents the degree of security risk of a specific network device or session in the current network environment, reflecting the likelihood of the node being threatened and its importance in the risk propagation process. It provides basic input for comprehensive risk scoring and helps identify key risk points in the network, enabling security protection resources to be prioritized for deployment on key nodes with higher risk values, thereby achieving more precise and efficient risk management.
[0061] Furthermore, since the risk values of adjacent nodes have not yet been calculated, the first calculated risk value adopts an initialization method based on anomaly identification data. Specifically, the anomaly degree value of the current node is used as its initial risk value. The anomaly degree value comes from the anomaly score corresponding to the node in the anomaly identification data and is set to 0. This initialization mechanism ensures that the risk propagation calculation has a reasonable starting point, avoids logical contradictions caused by the node calculation order, and ensures that the risk value of the first calculated node can accurately reflect its actual security status.
[0062] It should be noted that the attenuation factor corresponding to the protocol type is a predefined weighting coefficient based on the protocol type, used to adjust the risk propagation intensity. It is a predefined weighting coefficient based on protocol characteristics and historical security event analysis. The predefined method is to assign values by statistically analyzing the risk propagation characteristics of different protocols in historical attack events and combining expert experience.
[0063] Furthermore, the attenuation factors for specific protocol types are set as follows: TCP 0.8, UDP 0.6, ICMP 0.4, HTTP 0.7, and HTTPS 0.5. The attenuation factors for each protocol type adjust the intensity of risk propagation between different protocol sessions according to the characteristics of the protocol. For example, the risk attenuation of connection-oriented TCP is slower, while the risk attenuation of connectionless UDP is faster. This makes the risk propagation model more consistent with the security characteristics of different protocols in the actual network environment.
[0064] It should be noted that calculating the comprehensive risk score is essentially summing the risk values of all nodes to generate a quantitative index of overall network risk. The higher the comprehensive risk score, the greater the overall network risk, which is used for subsequent risk level classification.
[0065] Furthermore, the neighbor node set is all nodes extracted from the directed risk propagation graph and collected through a graph traversal algorithm to ensure that risk calculation covers the entire network topology.
[0066] It should be noted that the comprehensive risk score is an aggregated value of the overall network risk. It is a scalar index obtained by summing the risk values of all nodes. It is a quantitative representation of the overall risk level of the network and is used to provide input for risk level classification. The value ranges from 0 to 100. The higher the value, the greater the network risk. It is used to guide security response decisions.
[0067] In this embodiment of the invention, the step of classifying the risk level of the comprehensive risk score to obtain the risk level of the network traffic data includes: The comprehensive risk score is compared at multiple levels based on a preset risk level threshold to obtain the threshold comparison result. Based on this, a corresponding risk level identifier is assigned to the threshold comparison result through a risk level mapping rule to obtain the risk level of the network traffic data.
[0068] It should be noted that the essence of multi-level threshold comparison is the process of comparing the comprehensive risk score with the preset risk level threshold in real time, which is used to define the boundary of the risk level.
[0069] Furthermore, the multi-level threshold comparison has two thresholds: a low-to-medium risk threshold and a medium-to-high risk threshold. The low-to-medium risk threshold is 50, and the medium-to-high risk threshold is 80. When the comprehensive risk score is less than or equal to the low-to-medium risk threshold, the threshold comparison result is 1. When the comprehensive risk score is less than or equal to the medium-high risk threshold, the threshold comparison result is 2; When the overall risk score is greater than the medium-high risk threshold, the threshold comparison result is 3.
[0070] It should be noted that the threshold comparison result refers to the classification result obtained by comparing the comprehensive risk score calculated in real time with the preset multi-level thresholds. The values are 1, 2 and 3, which are used to map the continuous risk score to discrete risk levels and reflect the degree of deviation of the current security status of the network from the historical benchmark.
[0071] It should be noted that the risk level identifier is assigned by mapping the value of the threshold comparison result to the risk level. When the threshold comparison result is 1, the assigned risk level identifier is low risk. When the threshold comparison result is 2, the assigned risk level is identified as medium risk; When the threshold comparison result is 3, the assigned risk level is identified as high risk.
[0072] It should be noted that the risk level is a network security status classification obtained through threshold comparison and label allocation. It represents a qualitative assessment of the overall risk level of the network and is a discretized mapping result of the comprehensive risk score in a preset threshold system. It reflects the gradual change of the network from safe to dangerous and is used to provide a classification decision basis for security response.
[0073] Furthermore, low-risk levels require only routine monitoring, medium-risk levels require enhanced monitoring, and high-risk levels trigger an emergency response immediately, thereby achieving refined management and optimized resource allocation for network security.
[0074] S5, perform response instruction mapping on the risk level to obtain the risk response instruction for the network traffic data; In this embodiment of the invention, the step of mapping the risk level to a response instruction to obtain the risk response instruction for the network traffic data includes: The risk level is mapped to a response instruction to obtain the operation instructions for the network traffic data; The operation instructions are processed for protocol adaptation to obtain risk response instructions for the network traffic data.
[0075] It should be noted that response instruction mapping is a process of matching risk level identifiers with corresponding basic response actions. The specific mapping method is as follows: establish a correspondence between risk levels and response levels, and perform one-to-one matching and mapping based on the correspondence. This is used to transform abstract risk levels into specific executable response categories, ensuring that different risk levels can trigger security responses of appropriate strength.
[0076] Furthermore, low-risk levels are mapped to "notification" instructions, medium-risk levels to "enhanced monitoring" instructions, and high-risk levels to "emergency response" instructions.
[0077] It should be noted that the operation command is a standardized response command obtained after the response command is mapped. It is a general security response description that has not yet been adapted to the protocol and represents the direction of the response strategy initially determined by the system based on the risk level.
[0078] It should be noted that the protocol adaptation process is a process of adjusting the command parameters based on the operation command and network protocol type to obtain the risk response command. The protocol adaptation process includes parsing BGP routing policies and IP blocking rules.
[0079] Furthermore, the essence of protocol adaptation processing is to use a protocol-aware engine to convert operation instructions into executable network commands; To adjust BGP routing commands, add AS path attributes; To block malicious IP commands, set up IP blacklist entries and ensure that response commands are compatible with network devices.
[0080] It should be noted that the risk response instructions are a set of executable commands that are finally generated after protocol adaptation, including specific operation instructions, exception identification data and execution parameters; The operational instructions include: generating security alerts and sending them to the management platform, automatically adjusting firewall policies to block TCP connections of specified IPs, issuing flow table rules to switches to limit UDP traffic, and initiating deep packet inspection for specific protocol sessions. These instructions directly drive network security devices and software to perform specific protective actions, enabling rapid containment and handling of identified risks, thereby forming a closed-loop management system from risk detection to response and handling, effectively enhancing the proactive defense capabilities of network security.
[0081] Furthermore, the execution parameters are dynamically parsed and injected from the router configuration table in the network configuration library through protocol adaptation processing, including protocol-specific fields, operation rules, and time-limit settings.
[0082] S6. Based on the risk response instruction, the dynamic anomaly detection process is updated through the performance index feedback mechanism to obtain the updated anomaly monitoring threshold. In this embodiment of the invention, the step of updating the dynamic anomaly detection process based on the risk response instruction through a performance indicator feedback mechanism to obtain the updated anomaly monitoring threshold includes: Based on the risk response instructions, false alarm events and alarm events are extracted; Based on the false alarm events and the alarm events, calculate the total number of false alarms and the total number of alarms; Calculate the false alarm rate based on the total number of false alarms and the total number of alarms; Based on the false alarm rate, the dynamic anomaly monitoring threshold is updated to obtain the updated dynamic anomaly monitoring threshold.
[0083] It should be noted that false alarm events and alarm events are erroneous alarm records that have been verified and confirmed by rules, extracted from the alarm records after the execution of risk response instructions, as well as all alarm records within the same time period.
[0084] Furthermore, the total number of false alarms refers to the number of erroneous alarms that are verified and confirmed by the rules after the risk response command is executed, while the total number of alarms refers to the total number of alarm events within the same time period.
[0085] It should be noted that the false alarm rate is calculated by dividing the total number of false alarms by the total number of alarms. The false alarm rate represents the proportion of false alarms in the system alarms. It is used to quantify the accuracy of anomaly detection, evaluate detection performance, and provide feedback for adjusting the dynamic anomaly monitoring threshold. This optimizes detection accuracy in the network communication environment to reduce false alarms and improve security monitoring efficiency.
[0086] It should be noted that the process of updating the dynamic anomaly detection is achieved by updating the dynamic anomaly monitoring threshold. This is based on the false alarm rate and adjusts the old threshold through multiplication. Specifically, the update method is to multiply the dynamic anomaly monitoring threshold before the update by 1 and subtract the difference in the false alarm rate to obtain the updated dynamic anomaly monitoring threshold. This update method dynamically adjusts the detection sensitivity according to the false alarm rate. When the false alarm rate is high, the threshold is automatically lowered to tighten the detection and reduce false alarms. When the false alarm rate is low, the threshold is maintained or appropriately relaxed to balance detection efficiency, thereby enabling adaptive anomaly detection in digital communication networks and improving the system's adaptability and robustness to dynamic traffic changes.
[0087] It should be noted that the updated dynamic anomaly monitoring threshold is a real-time detection threshold adjusted based on performance feedback. It represents the dynamic boundary value that distinguishes normal and abnormal traffic under network conditions. It fluctuates with changes in network protocol behavior and traffic patterns, and is used to provide an adaptive benchmark for subsequent anomaly detection. This ensures that the detection process can respond to traffic changes in real time, improve detection accuracy and efficiency, reduce false alarms and false negatives, and enhance the reliability and real-time performance of overall security monitoring.
[0088] S7. Construct a comprehensive network information security analysis report based on the risk level, the risk response instruction, and the updated anomaly monitoring threshold.
[0089] In this embodiment of the invention, the step of constructing a comprehensive network information security analysis report based on the risk level, the risk response instruction, and the updated anomaly monitoring threshold includes: Integrate the risk level, the risk response instructions, and the updated anomaly monitoring thresholds to generate comprehensive report data; The comprehensive report data is input into a preset report format template for report formatting to generate a comprehensive network information security analysis report.
[0090] It should be noted that the comprehensive report data is generated by integrating risk levels, risk response instructions, and updated anomaly monitoring thresholds into a structured dataset. The specific method includes extracting risk levels, parsing the operation commands and parameters in the risk response instructions, and recording the updated dynamic anomaly monitoring threshold values. These data are then aggregated to obtain the comprehensive report data.
[0091] It should be noted that the report formatting process involves inputting the comprehensive report data into a preset report format template. The template engine then populates the data and applies style rules to generate standardized report content, including adding a report title, timestamp, risk level summary, response instruction details, threshold update records, and visualization charts. This process aims to transform the raw data into an easy-to-read and distribute document format, improving the report's operability and readability, and supporting cybersecurity management decisions.
[0092] like Figure 2 The diagram shown is a functional block diagram of an online information security comprehensive analysis and monitoring system provided in an embodiment of the present invention.
[0093] The online information security comprehensive analysis and monitoring system 100 described in this invention can be installed in an electronic device. Depending on the functions implemented, the online information security comprehensive analysis and monitoring system 100 may include a data structuring processing module 101, a structured feature extraction module 102, a dynamic anomaly monitoring module 103, a comprehensive risk assessment module 104, an intelligent response decision-making module 105, a data monitoring feedback module 106, and a comprehensive analysis report generation module 107. The modules described in this invention can also be referred to as units, which are a series of computer program segments that can be executed by the processor of an electronic device and can perform a fixed function, stored in the memory of the electronic device.
[0094] In this embodiment, the functions of each module / unit are as follows: The data structuring module is used to perform structuring processing on network traffic data to obtain structured data; The structured feature extraction module is used to extract time-frequency-protocol domain features from the structured data to obtain a structured feature vector containing protocol-specific indicators. The dynamic anomaly monitoring module is used to perform dynamic anomaly detection on the structured feature vector through a preset anomaly monitoring threshold to obtain anomaly identification data of the network traffic data. The comprehensive risk assessment module is used to calculate a comprehensive risk score based on anomaly identification data and network topology data, classify the risk level of the comprehensive risk score, and obtain the risk level of the network traffic data. The intelligent response decision module is used to perform response instruction mapping on the risk level to obtain the risk response instruction for the network traffic data; The data monitoring and feedback module is used to update the dynamic anomaly detection process based on the risk response instruction through a performance indicator feedback mechanism, so as to obtain the updated anomaly monitoring threshold. The comprehensive analysis report generation module is used to construct a comprehensive network information security analysis report based on the risk level, the risk response instruction, and the updated anomaly monitoring threshold.
[0095] In the several embodiments provided by this invention, it should be understood that the disclosed methods and systems can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.
[0096] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0097] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.
[0098] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0099] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0100] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for comprehensive analysis and monitoring of online information security, characterized in that, The method includes: S1, perform structured processing on network traffic data to obtain structured data; S2, perform time-frequency-protocol domain feature extraction on the structured data to obtain a structured feature vector containing protocol-specific indicators. The specific method for time-frequency-protocol domain feature extraction is as follows: Temporal features are extracted from the structured data to obtain temporal feature vectors; Frequency domain features are extracted from the structured data to obtain frequency domain feature vectors; Protocol domain features are extracted from the structured data by calculating the protocol state transition entropy value to obtain a protocol domain feature vector. The mathematical expression of the Shannon entropy formula used to calculate the protocol state transition entropy is as follows: ; In the formula, This is the protocol state transition entropy value. Protocol status The probability of occurrence in the state sequence For the first Each protocol status, For protocol index, This represents the total number of protocol states. The time-domain feature vector, the frequency-domain feature vector, and the protocol-domain feature vector are fused to obtain the structured feature vector; S3, dynamically detect anomalies in the structured feature vector using a preset anomaly detection threshold to obtain anomaly identification data for the network traffic data; S4. Based on the anomaly identification data and combined with the network topology data, calculate the comprehensive risk score, classify the risk level of the comprehensive risk score, and obtain the risk level of the network traffic data. S5, perform response instruction mapping on the risk level to obtain the risk response instruction for the network traffic data; S6. Based on the risk response instruction, the dynamic anomaly detection process is updated through the performance index feedback mechanism to obtain the updated anomaly monitoring threshold. S7. Construct a comprehensive network information security analysis report based on the risk level, the risk response instruction, and the updated anomaly monitoring threshold.
2. The online information security comprehensive analysis and monitoring method as described in claim 1, characterized in that, The process of structuring network traffic data to obtain structured data includes: The network traffic data is denoised to obtain initial network traffic data; The initial network traffic data is parsed using multiple protocols to obtain the protocol parsing data of the network traffic data; The protocol parsing data is serialized to obtain structured data of the network traffic data.
3. The online information security comprehensive analysis and monitoring method as described in claim 1, characterized in that, The step of dynamically detecting anomalies in the structured feature vector using a preset anomaly detection threshold to obtain anomaly identification data for the network traffic data includes: The preset anomaly monitoring threshold is dynamically adjusted based on the protocol state transition entropy value to obtain the dynamic anomaly monitoring threshold. The structured feature vector is used to detect anomalies using the dynamic anomaly monitoring threshold to obtain anomaly identification data for the network traffic data.
4. The online information security comprehensive analysis and monitoring method as described in claim 1, characterized in that, The calculation of the comprehensive risk score based on anomaly identification data and network topology data includes: Obtain network topology data; Based on the anomaly identification data and the network topology data, a directed risk propagation graph of the network traffic data is constructed; Based on the directed risk propagation graph, a breadth-first search logic is used to calculate the node risk value, wherein the mathematical expression of the breadth-first search logic is as follows: ; In the formula, For nodes The risk value, This is the node identifier of the current node. Adjacent nodes The risk value, This serves as the node identifier for adjacent nodes. Protocol type The corresponding attenuation factor, Protocol type, For nodes The set of adjacent nodes; The overall risk score of the node risk value is calculated by summing all the node risk values.
5. The online information security comprehensive analysis and monitoring method as described in claim 1, characterized in that, The process of classifying the risk level of the comprehensive risk score to obtain the risk level of the network traffic data includes: The comprehensive risk score is compared at multiple levels based on a preset risk level threshold to obtain the threshold comparison result. Based on this, a corresponding risk level identifier is assigned to the threshold comparison result through a risk level mapping rule to obtain the risk level of the network traffic data.
6. The online information security comprehensive analysis and monitoring method as described in claim 1, characterized in that, The process of mapping response instructions to the risk level to obtain risk response instructions for the network traffic data includes: The risk level is mapped to a response instruction to obtain the operation instructions for the network traffic data; The operation instructions are processed for protocol adaptation to obtain risk response instructions for the network traffic data.
7. The online information security comprehensive analysis and monitoring method as described in claim 3, characterized in that, The process of updating the dynamic anomaly detection procedure based on the risk response instruction through a performance indicator feedback mechanism to obtain the updated anomaly monitoring threshold includes: Based on the risk response instructions, false alarm events and alarm events are extracted; Based on the false alarm events and the alarm events, calculate the total number of false alarms and the total number of alarms; Calculate the false alarm rate based on the total number of false alarms and the total number of alarms; Based on the false alarm rate, the dynamic anomaly monitoring threshold is updated to obtain the updated dynamic anomaly monitoring threshold.
8. The online information security comprehensive analysis and monitoring method as described in claim 1, characterized in that, The comprehensive network information security analysis report constructed based on the risk level, the risk response instructions, and the updated anomaly monitoring threshold includes: Integrate the risk level, the risk response instructions, and the updated anomaly monitoring thresholds to generate comprehensive report data; The comprehensive report data is input into a preset report format template for report formatting to generate a comprehensive network information security analysis report.
9. A comprehensive online information security analysis and monitoring system, characterized in that, The system includes: The data structuring module is used to perform structuring processing on network traffic data to obtain structured data; The structured feature extraction module is used to extract time-frequency-protocol domain features from the structured data to obtain a structured feature vector containing protocol-specific indicators. The specific method for time-frequency-protocol domain feature extraction is as follows: Temporal features are extracted from the structured data to obtain temporal feature vectors; Frequency domain features are extracted from the structured data to obtain frequency domain feature vectors; Protocol domain features are extracted from the structured data by calculating the protocol state transition entropy value to obtain a protocol domain feature vector. The mathematical expression of the Shannon entropy formula used to calculate the protocol state transition entropy is as follows: ; In the formula, This is the protocol state transition entropy value. Protocol status The probability of occurrence in the state sequence For the first Each protocol status, For protocol index, This represents the total number of protocol states. The time-domain feature vector, the frequency-domain feature vector, and the protocol-domain feature vector are fused to obtain the structured feature vector; The dynamic anomaly monitoring module is used to perform dynamic anomaly detection on the structured feature vector through a preset anomaly monitoring threshold to obtain anomaly identification data of the network traffic data. The comprehensive risk assessment module is used to calculate a comprehensive risk score based on anomaly identification data and network topology data, classify the risk level of the comprehensive risk score, and obtain the risk level of the network traffic data. The intelligent response decision module is used to perform response instruction mapping on the risk level to obtain the risk response instruction for the network traffic data. The data monitoring and feedback module is used to update the dynamic anomaly detection process based on the risk response instruction through a performance indicator feedback mechanism, so as to obtain the updated anomaly monitoring threshold. The comprehensive analysis report generation module is used to construct a comprehensive network information security analysis report based on the risk level, the risk response instructions, and the updated anomaly monitoring threshold.
Citation Information
Patent Citations
Information security management method based on data processing
CN120128361A
Network security analysis method and system based on big data
CN120415850A
Automobile part enterprise supply chain risk early warning method based on artificial intelligence
CN120494628A
Data security monitoring method based on risk early warning
CN120546955A
Enhanced encrypted traffic analysis via integrated entropy estimation and neural network-based feature hybridization
US20250286903A1