Risk monitoring method based on intelligent association and global situation of multi-source data

By integrating multi-source heterogeneous data and utilizing graph neural networks and causal reasoning technology to generate a network security situation map, the shortcomings of existing network security systems in multi-source data processing and global situation awareness are addressed, achieving efficient threat detection and rapid response.

CN120675823AActive Publication Date: 2025-09-19INFORMATION & COMM CO OF STATE GRID SHAANXI ELECTRIC POWER CO LTD

Patent Information

Application Number
CN202511180622.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-22
Publication Date
2025-09-19
Estimated Expiration
2045-08-22

AI Technical Summary

Technical Problem

Existing network security systems lack comprehensive analysis capabilities when faced with multi-source heterogeneous data, making it difficult to cope with complex attack chains. They also have deficiencies in real-time and global situational awareness, resulting in low threat detection efficiency and difficult operation and maintenance.

Method used

By introducing causal reasoning and improved graph neural networks (GNN), multi-source heterogeneous data is integrated, attack behavior correlation maps are generated, anomaly detection is performed, explicit and potential attack links are obtained, and risk assessment reports are generated based on the network security situation dynamic map to trigger security policies.

Benefits of technology

It significantly improves the restoration capability and prediction accuracy of complex attack scenarios, reduces false alarm and missed alarm rates, realizes real-time threat detection and situation updates, reduces the operation and maintenance burden, and improves the global network security situation awareness capability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120675823A_ABST
    Figure CN120675823A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of network security, and particularly relates to an intelligent association and global situation risk monitoring method based on multi-source data, which comprises the following steps: acquiring a multi-source heterogeneous data set; the multi-source heterogeneous data set is preprocessed, and preprocessed multi-source data is obtained; obtaining an attack behavior association graph according to the multi-source data, and performing anomaly detection on the association graph by using a graph neural network to obtain an explicit attack link and a potential attack link; obtaining a network security situation dynamic graph based on the explicit attack link and the potential attack link; and generating a risk assessment report according to the network security situation dynamic graph, triggering a corresponding security policy, and performing risk monitoring according to the security policy. According to the method, the accuracy and response speed of threat detection are remarkably improved, the global network security situation awareness capability is enhanced, and an intelligent solution is provided for security protection in a complex network environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of network security technology, and in particular relates to a risk monitoring method based on intelligent association and global situation of multi-source data. Background Art

[0002] With the rapid development of internet technology, cyberattack methods are becoming increasingly diverse and complex, making traditional network security measures inadequate to cope with the modern threat landscape. Current security systems often rely on a single data source for threat detection and lack the ability to comprehensively analyze heterogeneous data from multiple sources. This limitation of a single data source renders systems vulnerable to complex attack vectors, particularly in scenarios such as advanced persistent threats (APTs), zero-day vulnerability exploits, and large-scale distributed attacks.

[0003] While existing network security systems (such as intrusion detection systems (IDS) and security information and event management (SIEM) systems) have improved security protection capabilities to a certain extent, they still have many shortcomings. For example, traditional security systems generate a large number of alerts, but due to the lack of contextual analysis, operations and maintenance personnel struggle to quickly locate key threats. Furthermore, existing systems rely primarily on rule matching or simple statistical analysis, which makes them inadequate when faced with complex attack chains. Regarding real-time performance, traditional systems struggle to process massive amounts of real-time data, making them unable to meet the demands of real-time threat detection in modern network environments. Furthermore, traditional systems typically focus only on localized threats and lack dynamic monitoring and visualization of the global network security status.

[0004] In recent years, the application of artificial intelligence and big data technologies in cybersecurity has gradually gained momentum, offering new insights into addressing these challenges. For example, machine learning and deep learning technologies can improve the accuracy of threat detection by training models to identify anomalous behavior. However, these technologies still face challenges in practical applications, such as poor data quality and insufficient model generalization. Graph computing and knowledge graph technologies can uncover hidden attack links by constructing attack behavior association graphs, but existing methods suffer from high computational complexity when processing large amounts of heterogeneous data and have limited capabilities for handling unstructured data. Stream computing and distributed frameworks can improve system timeliness by processing massive amounts of data in real time, but existing frameworks still need improvement in data fusion and dynamic analysis.

[0005] In modern network environments, threat detection alone is no longer sufficient to ensure network security. Decision-makers need to understand the security status of the entire network so they can take timely action. Global situational awareness technology integrates real-time and historical data to generate a dynamic network security situation map, helping decision-makers understand the overall security status and predict future trends. However, existing technologies still have significant room for improvement in the real-time, accuracy, and visualization capabilities of situational awareness.

[0006] As the frequency and complexity of cyberattacks increase, manual intervention is no longer sufficient for rapid response. Automated response mechanisms, combining Bayesian networks with SOAR (Security Orchestration, Automation, and Response) technology, can automatically generate risk assessment reports and trigger security policies, significantly improving emergency response efficiency. However, existing automated systems still need to improve the flexibility and effectiveness of policy orchestration.

[0007] In summary, existing network security technologies have significant shortcomings in multi-source data integration, intelligent analysis, global situational awareness, and automated response. To address these challenges, an intelligent analysis method that can integrate multi-source heterogeneous data is urgently needed. Summary of the Invention

[0008] To solve the above technical problems, the present invention proposes a risk monitoring method based on intelligent correlation and global situation of multi-source data. By introducing causal reasoning and an improved graph neural network (GNN), this method significantly improves the restoration capability and prediction accuracy of complex attack scenarios, and supports high-concurrency data processing and dynamic strategy orchestration.

[0009] To achieve the above objectives, the present invention provides a risk monitoring method based on intelligent correlation and global situation of multi-source data, comprising:

[0010] Acquire multi-source heterogeneous datasets;

[0011] Preprocessing the multi-source heterogeneous data set to obtain preprocessed multi-source data;

[0012] Obtain an attack behavior association graph based on the multi-source data, perform anomaly detection on the association graph using a graph neural network, and obtain explicit attack links and potential attack links;

[0013] Based on explicit attack links and potential attack links, a dynamic picture of network security situation is obtained;

[0014] A risk assessment report is generated based on the network security situation dynamic graph, and corresponding security policies are triggered, and risk monitoring is performed according to the security policies.

[0015] Optionally, obtaining the multi-source heterogeneous dataset includes:

[0016] Collect log information, status information, local flash memory data, and threat intelligence data from network devices, security devices, and servers through multiple protocols to form an initial multi-source heterogeneous data set.

[0017] The initial multi-source heterogeneous data set is partitioned using a partitioning strategy to obtain the multi-source heterogeneous data set.

[0018] Optionally, preprocessing the multi-source heterogeneous data set to obtain the preprocessed multi-source data includes:

[0019] Cleaning, deduplication, and format standardization are performed on the multi-source heterogeneous data set to obtain initially processed multi-source data, where the initially processed multi-source data includes structured data and unstructured data;

[0020] Parsing the structured data using a regularization method to obtain target data, and unifying the format of the target data to obtain processed structured data;

[0021] Converting the unstructured data into vectors using a word embedding method to obtain processed unstructured data;

[0022] The isolation forest method is used to remove abnormal data points in the processed structured data and the processed unstructured data to obtain multi-source data after preprocessing.

[0023] Optionally, the structured data is parsed using a regularization method to obtain target data, and the target data is formatted uniformly. The obtained processed structured data includes:

[0024] ;

[0025] ;

[0026] in, (*) is the regular expression parsing function, For the original log, For the parsed structured data, For predefined standardized templates, For the processed structured data, (*) is the data format unification function.

[0027] Optionally, converting the unstructured data into a vector using a word embedding method includes:

[0028] ;

[0029] in, For unstructured log text, is the vector representation of unstructured log text, (*) is the word embedding function.

[0030] Optionally, obtaining the attack behavior association graph according to the multi-source data includes:

[0031] Based on a graph computing model, entities in the multi-source data are modeled as nodes, event relationships are modeled as edges, and edge weights are calculated to obtain the attack behavior association graph.

[0032] Optionally, use a graph neural network to detect anomalies in the association graph and obtain potential attack links, including:

[0033] Use the graph neural network to perform anomaly detection on the association graph to obtain abnormal nodes;

[0034] Calculate the probability of the attack path in the path formed by the abnormal node using the Bayesian model to obtain an explicit attack link;

[0035] Frequent subgraph mining is performed on the association graph to obtain potential attack links.

[0036] Optionally, based on the explicit attack links and potential attack links, a method for obtaining a dynamic graph of network security situation is:

[0037] ;

[0038] in, is the fused data, is the real-time data at the current moment, For historical data, Optionally, generating a risk assessment report based on the network security situation dynamic graph and triggering corresponding security policies include:

[0039] Calculating risk probability using a Bayesian model based on the network security situation dynamic graph;

[0040] Generate an assessment report based on the risk probability;

[0041] Based on the assessment report, corresponding security policies are triggered.

[0042] The present invention also provides a risk monitoring system based on intelligent association and global situation of multi-source data, including: a data acquisition module, a data pre-processing module, an intelligent association analysis module, a global situation awareness module and a linkage disposal module;

[0043] The data acquisition module is used to collect log information, status information and external threat intelligence data of network devices, security devices and servers through multiple protocols to form a multi-source heterogeneous data set;

[0044] The data preprocessing module is used to process the multi-source heterogeneous data set to obtain preprocessed multi-source data;

[0045] The intelligent association analysis module is used to obtain an attack behavior association graph based on the multi-source data, perform anomaly detection on the association graph using a graph neural network, and obtain explicit attack links and potential attack links;

[0046] The global situation awareness module is used to obtain a dynamic diagram of network security situation based on explicit attack links and potential attack links;

[0047] The linkage handling module is used to generate a risk assessment report based on the network security situation dynamic diagram, trigger corresponding security policies, and perform risk monitoring according to the security policies.

[0048] Compared with the prior art, the present invention has the following advantages and technical effects:

[0049] The present invention integrates multi-source heterogeneous data and uses machine learning algorithms (such as deep learning and graph neural networks) for comprehensive analysis, significantly reducing false alarm and missed alarm rates, accurately identifying hidden patterns in complex attack chains, and adopting distributed computing frameworks (such as Spark and Flink) and streaming computing technologies to achieve real-time threat detection and situation updates, processing massive amounts of data in milliseconds, greatly improving system timeliness, and combining automated orchestration technology to automatically generate risk assessment reports, trigger emergency response strategies, and shorten threat handling time; in addition, it integrates real-time and historical data to generate dynamic network security situation maps, providing a global perspective to fully grasp the security status, quickly locate high-risk areas and support long-term strategy formulation, and combining natural language processing (NLP), knowledge graphs and causal reasoning technologies to deeply mine unstructured data (such as social media discussions), and provide It can provide early warning of potential threats and reveal attack logic, thus improving the level of intelligence. It can significantly reduce invalid alarms and lower the burden of operation and maintenance through efficient integration and intelligent analysis of multi-source data. The automated response mechanism can further reduce manual intervention and improve operation and maintenance efficiency. It is widely applicable to enterprise network security, cloud security, industrial Internet, smart cities and other fields. It can monitor abnormal behavior in real time, predict failures and ensure the security of critical infrastructure. It not only solves the key technical problems in the current network security field, but also provides guidance for future security reinforcement through technologies such as causal reasoning. It is highly scalable to promote continuous innovation, improve threat detection and response capabilities in terms of economic benefits and reduce corporate economic losses, and improve the security of critical infrastructure in terms of social benefits to help build a safer cyberspace. It is an important milestone in responding to the complex security challenges of the modern network environment and promoting the development of network security technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of this application. The exemplary embodiments and descriptions of this application are intended to explain this application and do not constitute an improper limitation on this application. In the accompanying drawings:

[0051] Figure 1 This is a flow chart of a risk monitoring method based on intelligent association of multi-source data and global situation according to an embodiment of the present invention;

[0052] Figure 2 is a flow chart of multi-source data collection and preprocessing according to an embodiment of the present invention;

[0053] Figure 3 Schematic diagram of the working principle of the intelligent association analysis module according to an embodiment of the present invention;

[0054] Figure 4 This is a flow chart of policy arrangement of the linkage processing module according to an embodiment of the present invention. DETAILED DESCRIPTION

[0055] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0056] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0057] Integrating multi-source data not only improves data comprehensiveness and accuracy but also provides richer context for subsequent intelligent analysis. Internal data includes network device logs, security device alarms, and server operating status. External data includes malicious IP address databases and malicious domain name lists provided by threat intelligence platforms. Unstructured data includes user behavior descriptions and security-related discussions on social media. Comprehensive analysis of this data allows for a more comprehensive reconstruction of attack scenarios and identification of potential threats.

[0058] This embodiment proposes a risk monitoring method based on intelligent association of multi-source data and global situation, such as Figure 1 As shown, the specific steps include:

[0059] Preprocessing the multi-source heterogeneous data set to obtain preprocessed multi-source data;

[0060] Obtain an attack behavior association graph based on the multi-source data, perform anomaly detection on the association graph using a graph neural network, and obtain explicit attack links and potential attack links;

[0061] Based on explicit attack links and potential attack links, a dynamic picture of network security situation is obtained;

[0062] A risk assessment report is generated based on the network security situation dynamic graph, and corresponding security policies are triggered, and risk monitoring is performed according to the security policies.

[0063] Specifically, this embodiment integrates multi-source heterogeneous data (such as network device logs, security device alarms, and external threat intelligence) and combines natural language processing and knowledge graph technology to perform semantic extraction on unstructured data to build a unified knowledge representation model. On this basis, advanced technologies such as graph neural networks, Bayesian reasoning, and frequent subgraph mining are used to mine hidden attack links and abnormal patterns, and based on big data streaming computing technology, real-time and historical data are integrated to generate a dynamic network security situation map. At the same time, the distribution, trend, and impact range of attack events are displayed through visualization technology, and security policies are triggered in combination with automated orchestration technology to achieve efficient emergency response. The present invention significantly improves the accuracy and response speed of threat detection, enhances the global network security situation awareness capability, and provides an intelligent solution for security protection in complex network environments.

[0064] Furthermore, obtaining the multi-source heterogeneous data set includes:

[0065] Collect log information, status information, local flash memory data, and threat intelligence data from network devices, security devices, and servers through multiple protocols to form an initial multi-source heterogeneous data set.

[0066] The initial multi-source heterogeneous data set is partitioned using a partitioning strategy to obtain the multi-source heterogeneous data set.

[0067] Specifically, the data collection module actively collects log information, status information and local flash memory data of network devices (such as routers, switches), security devices (such as firewalls, IDS / IPS) and servers through multiple protocols (such as Telnet, SNMP, Syslog, NetFlow, etc.), and obtains the latest threat intelligence data (such as malicious IP address database, malicious domain name list, etc.) from external threat intelligence platforms to form a multi-source heterogeneous data set. The data collection frequency adopts a periodic polling mechanism, and the collection frequency is Can be adjusted dynamically according to needs:

[0068] ;

[0069] Where t is the collection interval in seconds. To optimize the storage and query performance of massive data, distributed storage technologies (such as HDFS and Cassandra) are used to store data to ensure high data reliability and high throughput. The data partitioning strategy is as follows:

[0070] ;

[0071] in, A unique identifier used to determine which partition the data should be stored in. is the identifier of the data source node, is the timestamp, (*) is a hash function. In addition, to reduce bandwidth consumption, the data acquisition module supports data compression and incremental transmission.

[0072] Furthermore, preprocessing the multi-source heterogeneous data set to obtain the preprocessed multi-source data includes:

[0073] Cleaning, deduplication, and format standardization are performed on the multi-source heterogeneous data set to obtain initially processed multi-source data, where the initially processed multi-source data includes structured data and unstructured data;

[0074] Parsing the structured data using a regularization method to obtain target data, and unifying the format of the target data to obtain processed structured data;

[0075] Converting the unstructured data into vectors using a word embedding method to obtain processed unstructured data;

[0076] The isolation forest method is used to remove abnormal data points in the processed structured data and the processed unstructured data to obtain multi-source data after preprocessing.

[0077] Specifically, data preprocessing cleans, removes duplicates, and standardizes the format of collected multi-source data to ensure data consistency and availability. Regular expressions are used to parse the original logs and extract key fields (such as timestamp, source IP, destination IP, protocol type, etc.). The formula is as follows:

[0078] ;

[0079] in, (*) is the regular expression parsing function, For the original log, For parsed structured data, convert data in different formats into a unified JSON format:

[0080] ;

[0081] in, For the processed structured data, For predefined standardized templates, (*) is a normalization function.

[0082] For unstructured log text, use word embedding technology (such as Word2Vec, BERT) to convert it into a vector representation:

[0083] ;

[0084] in, is the log text, V is the vector representation of the unstructured log text, is the word embedding function. To detect abnormal data points, the Isolation Forest algorithm is used, and the formula is as follows:

[0085] ;

[0086] in, is the abnormality score, (*) is the isolation forest algorithm, is the input dataset, is the abnormal proportion parameter. In addition, the data preprocessing module also supports dimensionality reduction based on feature engineering to improve the efficiency of subsequent analysis. Specifically, It can be processed structured data , or it can be processed unstructured data.

[0087] Furthermore, obtaining the attack behavior association graph based on the multi-source data includes:

[0088] Based on a graph computing model, entities in the multi-source data are modeled as nodes, event relationships are modeled as edges, and edge weights are calculated to obtain the attack behavior association graph.

[0089] Furthermore, we use graph neural networks to detect anomalies in the association graph and obtain explicit and potential attack links, including:

[0090] Use the graph neural network to perform anomaly detection on the association graph to obtain abnormal nodes;

[0091] Calculate the probability of the attack path in the path formed by the abnormal node using the Bayesian model to obtain an explicit attack link;

[0092] Frequent subgraph mining is performed on the association graph to obtain potential attack links.

[0093] Specifically, the intelligent association analysis module is based on a graph computing framework (such as Neo4j, GraphX), which models entities (such as IP addresses, user accounts, device IDs, etc.) in multi-source data as nodes and event relationships as edges to construct an attack behavior association graph. The calculation formula is as follows:

[0094] ;

[0095] in, For nodes and nodes The similarity of To adjust the parameters. Use the graph neural network (GNN) to detect anomalies on the graph and calculate the node anomaly score :

[0096] ;

[0097] in, For nodes The neighbor set of For nodes The eigenvector of .

[0098] In order to infer the attack path, based on the Bayesian reasoning model:

[0099] ;

[0100] in, is the attack path probability under given evidence, is the prior path probability, is the probability of evidence, For possible attack paths, For observed evidence, use a frequent subgraph mining algorithm (such as gSpan) to mine the attack links hidden in the attack behavior association graph:

[0101] ;

[0102] in, is a set of frequent subgraphs, (*) is the frequent subgraph mining algorithm, is the input graph (attack behavior association graph), is the minimum support threshold. In order to restore the attack scenario, causal inference technology is used:

[0103] ;

[0104] in, is a cause-effect diagram, (*) is the causal inference function, Causal reasoning is based on prior knowledge constraints, such as known attack patterns and security rules. By analyzing the causal relationships in the graph, it infers the time sequence and logical dependencies of attack behaviors. Combined with the constraints, it can more accurately restore the attack scenario and help security analysts understand the attacker's intentions and methods.

[0105] Furthermore, based on explicit attack links and potential attack links, the method for obtaining a dynamic diagram of network security situation is as follows:

[0106] Global situational awareness uses big data streaming computing technologies (such as Apache Flink and Kafka Streams) to fuse and analyze real-time and historical data to generate a dynamic network security situation map. Sliding window technology is used to fuse real-time and historical data. Historical data refers to data collected in the past, while real-time data refers to data collected at the current moment. The collected data is multi-source heterogeneous data:

[0107] ;

[0108] in, is the fused data, is the real-time data at the current moment, For historical data, is the weight coefficient.

[0109] Calculating the Cybersecurity Posture Score:

[0110] ;

[0111] in, is the indicator weight, is the indicator value, is the situation score, and n is the number of indicators.

[0112] Use time series analysis (such as ARIMA model, LSTM) to predict the evolution trend of the situation:

[0113] ;

[0114] in, is the current state value, is the regression coefficient, is the noise term, is the state value at the previous moment, is the state value of the first two moments.

[0115] Use heat maps, topology maps, and other forms to display the distribution and propagation paths of attack events:

[0116] ;

[0117] in, For heat map, (*) is the rendering function, is the spatial distribution of attack events, is the color mapping rule.

[0118] Specifically, the collected data includes situational awareness data, such as the number of malicious IP accesses, the frequency of attack behaviors, changes in network traffic, the proportion of abnormal traffic, etc. .

[0119] exist middle is the real-time data at the current moment, is the historical data (aggregation result within the sliding window), is the weight coefficient, which is used to control the contribution ratio of real-time data and historical data.

[0120] From , for the fused situation awareness data at time t, weighted calculation of the situation score at time t is performed.

[0121] Can be modified to ,in, Score the current situation. Score the situation at the previous moment, Score the situation at the next moment and predict the situation score at the next moment based on the situation scores at the current and past moments.

[0122] Using heat maps, topology maps, and other forms to display the distribution and propagation paths of attack events has nothing to do with the previous content. It is an abstract representation used to describe how to generate heat maps through data rendering.

[0123] Heatmap: The final generated heat map is a visualization result used to intuitively show the spatial distribution of attack events.

[0124] Render: Rendering function, which represents the process of converting the input data (AttackDistribution) and color mapping rules (ColorMapping) into a heat map.

[0125] AttackDistribution: Spatial distribution data of attack events, usually a two-dimensional or three-dimensional matrix, recording the number or intensity of attack events at different locations (such as network nodes, IP addresses, geographic locations, etc.).

[0126] ColorMapping: Color mapping rules define how to assign colors based on the intensity of attack events (such as quantity, frequency, etc.).

[0127] Furthermore, generating a risk assessment report based on the network security situation dynamic graph and triggering corresponding security policies include:

[0128] Calculating risk probability using a Bayesian model based on the network security situation dynamic graph;

[0129] Generate an assessment report based on the risk probability;

[0130] Based on the assessment report, corresponding security policies are triggered.

[0131] Specifically, the linkage disposal module automatically generates a risk assessment report based on the global situation awareness results and triggers corresponding security policies (such as blocking malicious IPs, isolating infected devices, etc.). Use Bayesian networks to calculate risk probabilities :

[0132] ;

[0133] in, is the conditional variable, As its parent node.

[0134] Trigger predefined security policies based on risk assessment results:

[0135] ;

[0136] in, is the risk threshold, Score the current risk, For specific response actions (such as blocking malicious IP, isolating devices, etc.), (*) is the trigger function, which decides whether to take action based on RiskScore and Threshold.

[0137] Implementing policy orchestration using SOAR (Security Orchestration, Automation, and Response) technology:

[0138] ;

[0139] in, is a set of executable actions, is context information, The final automated workflow is used to execute a series of security policies. (*) is an orchestration function used to combine ActionSet and Context to generate an automated workflow.

[0140] To verify the effectiveness of the strategy, conduct emergency response drills using simulated attack scenarios (Red Team vs Blue Team):

[0141] ;

[0142] in, For the simulation results, For the real situation, Policy effectiveness, which indicates the actual effect of the security policy in the simulated attack scenario. (*) is the evaluation function, which is used to compare the simulation results with the actual situation and calculate the effectiveness of the strategy.

[0143] This embodiment also provides a risk monitoring system based on intelligent correlation and global situation of multi-source data, including: a data acquisition module, a data pre-processing module, an intelligent correlation analysis module, a global situation awareness module and a linkage disposal module;

[0144] The data acquisition module is used to collect log information, status information and external threat intelligence data of network devices, security devices and servers through multiple protocols to form a multi-source heterogeneous data set;

[0145] The data preprocessing module is used to process the multi-source heterogeneous data set to obtain preprocessed multi-source data;

[0146] The intelligent association analysis module is used to obtain an attack behavior association graph based on the multi-source data, perform anomaly detection on the association graph using a graph neural network, and obtain explicit attack links and potential attack links;

[0147] The global situation awareness module is used to obtain a dynamic diagram of network security situation based on explicit attack links and potential attack links;

[0148] The linkage handling module is used to generate a risk assessment report based on the network security situation dynamic diagram, trigger corresponding security policies, and perform risk monitoring according to the security policies.

[0149] The present embodiment will be described in detail below with reference to the accompanying drawings:

[0150] like Figure 2As shown, data collection is the foundation of this embodiment and is responsible for acquiring multi-source heterogeneous data from network devices, security devices, servers, and external threat intelligence platforms. In actual deployment, it is first necessary to install data collection agents on key nodes of the network (such as routers, switches, firewalls, servers, etc.). These agents actively collect device status information and log information through various protocols (such as SNMP, Syslog, NetFlow, and API interfaces). For example, traffic statistics of routers are collected through the SNMP protocol, alarm logs of firewalls are collected through the Syslog protocol, and the latest malicious IP address database and domain name list are obtained from the external threat intelligence platform through the API interface. The collection frequency can be dynamically adjusted according to actual conditions. For example, a shorter collection interval (such as once per second) can be set during high-traffic periods, and a longer collection interval (such as once per minute) can be set during low-traffic periods.

[0151] like Figure 2 As shown in Figure 1, the main task of data preprocessing is to clean, deduplicate, and standardize the format of collected multi-source data to ensure data consistency and availability. For structured log data, regular expressions are used to extract key fields such as timestamp, source IP address, destination IP address, and protocol type. Taking Syslog format logs as an example, the following regular expression is designed:

[0152] .

[0153] After parsing, the raw logs are converted into a unified JSON format for subsequent analysis. Natural language processing techniques are used to extract semantic meaning from unstructured log text (such as free text descriptions). For example, the BERT model is used to generate a vector representation of the log text, converting a log entry such as "User admin attempted to log in but failed" into a 1024-dimensional vector. Furthermore, to detect anomalous data points, the Isolation Forest algorithm is applied. Assuming a 1% anomaly rate, 100 isolation trees are constructed for anomaly detection.

[0154] like Figure 3 As shown, intelligent association analysis is based on graph computing frameworks (such as Neo4j and GraphX), modeling entities in multi-source data as nodes and event relationships as edges to construct an attack behavior association graph. For example, a node could represent the IP address "192.168.1.1," another node could represent the user "admin," and the edge between the two could represent login behavior. Edge weights are calculated using the formula:

[0155] ;

[0156] Based on this, we use graph neural networks (GNNs) to detect anomalies in graphs and identify potential attack links and abnormal patterns. For example, we analyze frequently occurring subgraphs to uncover hidden attack paths. We use the gSpan algorithm to set a minimum support threshold of 0.05 to identify frequent subgraphs.

[0157] Global situational awareness uses big data streaming computing technologies (such as Apache Flink and Kafka Streams) to fuse and analyze real-time and historical data to generate a dynamic network security situation map. Sliding window technology is used to fuse real-time and historical data. For example, if the sliding window size is set to 5 minutes and the step size is 1 minute, the fusion formula is:

[0158] ;

[0159] Define multiple security indicators (such as attack frequency, number of affected assets, and threat level), assign different weights, and calculate the overall situation score.

[0160] To further predict the evolution of the situation, an LSTM model is used to forecast changes in the situation over the next 30 minutes, using input features including historical situation scores from the past hour. Ultimately, heat maps and topology diagrams are used to display the distribution, propagation paths, and impact range of attack events, providing decision makers with an intuitive global view.

[0161] like Figure 4 As shown in the figure, the coordinated response automatically generates a risk assessment report based on the global situational awareness results and triggers the corresponding security policies. Risk probabilities are calculated based on Bayesian networks. For example, given conditional variables (such as attack type, impact scope, and duration), an overall risk score is inferred. When the risk score exceeds a preset threshold (such as 80 points), security policies are automatically triggered, such as blocking malicious IP addresses and isolating infected devices. SOAR (Security Orchestration, Automation, and Response) technology is used to implement policy orchestration, defining the following action sequence:

[0162] 1. Block malicious IP addresses;

[0163] 2. Send alarm notifications to operation and maintenance personnel;

[0164] 3. Launch a vulnerability scanning tool to check the affected devices.

[0165] The above are merely preferred embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A risk monitoring method based on intelligent association of multi-source data and global situation, characterized by: include: Acquire multi-source heterogeneous datasets; Preprocessing the multi-source heterogeneous data set to obtain preprocessed multi-source data; Obtain an attack behavior association graph based on the multi-source data, perform anomaly detection on the association graph using a graph neural network, and obtain explicit attack links and potential attack links; Based on the explicit attack links and potential attack links, a dynamic diagram of network security situation is obtained; A risk assessment report is generated based on the network security situation dynamic graph, and corresponding security policies are triggered, and risk monitoring is performed according to the security policies.

2. The risk monitoring method based on intelligent association and global situation of multi-source data according to claim 1 is characterized in that: Acquiring the multi-source heterogeneous data set includes: Collect log information, status information, local flash memory data, and threat intelligence data from network devices, security devices, and servers through multiple protocols to form an initial multi-source heterogeneous data set. The initial multi-source heterogeneous data set is partitioned using a partitioning strategy to obtain the multi-source heterogeneous data set.

3. The risk monitoring method based on intelligent association and global situation of multi-source data according to claim 1 is characterized in that: Preprocessing the multi-source heterogeneous data set to obtain the preprocessed multi-source data includes: Cleaning, deduplication, and format standardization are performed on the multi-source heterogeneous data set to obtain initially processed multi-source data, where the initially processed multi-source data includes structured data and unstructured data; Parsing the structured data using a regularization method to obtain target data, and unifying the format of the target data to obtain processed structured data; Converting the unstructured data into vectors using a word embedding method to obtain processed unstructured data; The isolation forest method is used to remove abnormal data points in the processed structured data and the processed unstructured data to obtain multi-source data after preprocessing.

4. The risk monitoring method based on intelligent association and global situation of multi-source data according to claim 3 is characterized in that: The structured data is parsed using a regularization method to obtain target data, and the target data is formatted uniformly. The processed structured data includes: ; ; in, (*) is the regular expression parsing function, For the original log, For the parsed structured data, For predefined standardized templates, For the processed structured data, (*) is the data format unification function.

5. The risk monitoring method based on intelligent association and global situation of multi-source data according to claim 3 is characterized in that: Converting the unstructured data into vectors using word embedding includes: ; in, For unstructured log text, is the vector representation of unstructured log text, (*) is the word embedding function.

6. The risk monitoring method based on intelligent association and global situation of multi-source data according to claim 1 is characterized in that: Acquiring the attack behavior association graph according to the multi-source data includes: Based on a graph computing model, entities in the multi-source data are modeled as nodes, event relationships are modeled as edges, and edge weights are calculated to obtain the attack behavior association graph.

7. The risk monitoring method based on intelligent association and global situation of multi-source data according to claim 6 is characterized in that: Use graph neural networks to detect anomalies in association graphs and obtain potential attack links, including: Use the graph neural network to perform anomaly detection on the association graph to obtain abnormal nodes; Calculate the probability of the attack path in the path formed by the abnormal node using the Bayesian model to obtain an explicit attack link; Frequent subgraph mining is performed on the association graph to obtain potential attack links.

8. The risk monitoring method based on intelligent association and global situation of multi-source data according to claim 1 is characterized in that: Based on the explicit attack links and potential attack links, a method for obtaining a dynamic diagram of network security situation is as follows: ; in, is the fused data, is the real-time data at the current moment, For historical data, is the weight coefficient.

9. The risk monitoring method based on intelligent association and global situation of multi-source data according to claim 1 is characterized in that: Generating a risk assessment report based on the network security situation dynamic graph and triggering corresponding security policies include: Calculating risk probability using a Bayesian model based on the network security situation dynamic graph; Generate an assessment report based on the risk probability; Based on the assessment report, corresponding security policies are triggered.

10. A risk monitoring system based on intelligent association of multi-source data and global situation implemented according to the method according to any one of claims 1 to 9, characterized in that: include: Data acquisition module, data preprocessing module, intelligent correlation analysis module, global situation awareness module and linkage disposal module; The data acquisition module is used to collect log information, status information and external threat intelligence data of network devices, security devices and servers through multiple protocols to form a multi-source heterogeneous data set; The data preprocessing module is used to process the multi-source heterogeneous data set to obtain preprocessed multi-source data; The intelligent association analysis module is used to obtain an attack behavior association graph based on the multi-source data, perform anomaly detection on the association graph using a graph neural network, and obtain explicit attack links and potential attack links; The global situation awareness module is used to obtain a dynamic diagram of network security situation based on explicit attack links and potential attack links; The linkage handling module is used to generate a risk assessment report based on the network security situation dynamic diagram, trigger corresponding security policies, and perform risk monitoring according to the security policies.

Citation Information

Patent Citations

  • Network information security protection system

    CN118353702A

  • Network security threat tracing method and system based on correlation analysis

    CN119324817A

  • Tunnel full-period construction feature information fusion and quality tracing method and system

    CN120087821A

  • Network security situation awareness prediction method and device based on knowledge graph

    CN120128434A

  • Methods and apparatuses for event detection

    WO2024067950A1

Cited By

  • Network security situation awareness method and application

    CN120880781A

  • Network security situation awareness method and system

    CN120915595A

  • A cyber security situation awareness method and system

    CN120915595B

  • Network attack path tracking method and system based on three-domain communication event structure

    CN121037104A

  • Method and device for constructing network attack behavior chain and active defense, and computer equipment

    CN121151134A