Network security operation monitoring management system and method based on big data

By using multi-source data stream processing, dynamic knowledge graphs, and meta-reinforcement learning to generate adaptive defense strategies, the problem of existing systems being unable to self-optimize has been solved, achieving efficient and accurate network security operation and management.

CN122069074APending Publication Date: 2026-05-19GUANGDONG POWER GRID CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGDONG POWER GRID CO LTD
Filing Date
2026-02-04
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing network security operation monitoring and management systems are difficult to automatically adjust to changes in attack techniques and tactics, and the systems are difficult to continuously optimize and adapt to themselves, resulting in passive and lagging security operations.

Method used

It employs a multi-source secure data stream processing pipeline, a network security situation awareness engine, an adaptive policy generator, and a policy execution and optimization closed loop. Through dynamic knowledge graphs and meta-reinforcement learning mechanisms, it generates adaptive defense policies and performs simulation evaluation and optimization in a digital twin sandbox.

Benefits of technology

It significantly enhances the initiative, accuracy, and overall resilience of security operations, improves the ability to detect complex attack chains and the operability of response decisions, ensures the level of automation in threat response and business continuity, and avoids business risks caused by trial and error of defense measures in real networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122069074A_ABST
    Figure CN122069074A_ABST
Patent Text Reader

Abstract

The invention discloses a network security operation monitoring management system and method based on big data, and relates to the technical field of network security, and the system comprises a multi-source security data stream processing pipeline, a network security situation awareness engine, a self-adaptive strategy generator and a strategy execution and optimization closed loop. According to the method, multi-source heterogeneous data is converted into a standardized event stream in real time, deep behavior modeling and threat hunting are carried out based on a dynamic knowledge graph, an attack intention chain with high confidence is generated, then the knowledge graph and a meta reinforcement learning mechanism are combined, a self-adaptive defense strategy giving consideration to safety restraint and service continuity is generated, and the security and the service continuity are improved. Finally, a strategy effect is verified safely through digital twin simulation, an evaluation result is fed back bidirectionally, and a threat detection model at the front end and a decision-making model at the rear end are optimized synchronously, so that the system has a continuous self-evolution capability, and the initiative, accuracy and overall toughness of safe operation are remarkably enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of network security technology, specifically to a network security operation monitoring and management system and method based on big data. Background Technology

[0002] With the rapid development of information technology, cyberattacks are increasingly characterized by large scale, complexity, concealment, and persistence. The traditional security operation model that relies on single protection devices and manual analysis is facing severe challenges. In recent years, technologies such as knowledge graphs, graph neural networks, and reinforcement learning have been attempted to be introduced into the field of cybersecurity to improve threat correlation analysis and automated response.

[0003] The existing network security operation monitoring and management system has the following shortcomings:

[0004] 1. Patent document CN117640257B discloses a data processing method and system for network security operation based on big data. "This invention discloses a data processing method and system for network security operation based on big data, relating to the field of network security technology. When the system is running, it collects real-time network traffic data and IP address related information to form traffic data groups and IP address data groups. It preprocesses these groups to obtain traffic-related information, forming a first dataset, and obtains IP address-related information, forming a second dataset. It matches the traffic-related information and IP address-related information to obtain the number of one-to-many or one-to-one relationship matches for IP addresses, records these matches, and forms a third dataset. It then calculates and obtains an anomaly index Yczs, which is matched with a preset anomaly warning threshold Y to obtain an anomaly level strategy scheme, which is then implemented. When traffic anomalies and IP anomalies are correlated, it can effectively detect them and take real-time protective measures." However, the aforementioned document suffers from technical problems: it is difficult to automatically adjust to changes in attack tactics, and the system lacks the ability to continuously self-optimize and adapt, leading to passive and lagging security operations. Summary of the Invention

[0005] The purpose of this invention is to provide a network security operation monitoring and management system and method based on big data, so as to solve the technical problems mentioned in the background.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a network security operation monitoring and management system based on big data, characterized in that it includes:

[0007] A multi-source secure data stream processing pipeline is used to access and parse heterogeneous secure data in real time and output a standardized time-series event stream.

[0008] The network security situation awareness engine connects to a multi-source security data stream processing pipeline to continuously store event streams and perform entity behavior modeling and threat correlation analysis based on a dynamically evolving network knowledge graph, in order to output a high-confidence attack intent chain.

[0009] An adaptive policy generator, connected to a network security situation awareness engine, is used to receive the attack intent chain and generate a targeted defense policy sequence based on a pre-built response action knowledge graph and a meta-reinforcement learning mechanism.

[0010] The strategy execution and optimization closed loop connects to the adaptive strategy generator, which is used to execute the defense strategy sequence in a controllable environment and evaluate the effect through digital twin simulation. The feedback signal is then used to optimize the behavior model of the network security situation awareness engine and the decision model of the adaptive strategy generator.

[0011] Preferably, the network security situation awareness engine includes a dynamic knowledge graph unit and a threat hunting analysis unit. The dynamic knowledge graph unit is used to construct and maintain a multi-hop cognitive graph that represents network assets, vulnerabilities, services, and historical interaction relationships. The threat hunting analysis unit is used to dynamically encode the behavioral patterns of entities in the multi-hop cognitive graph using a temporal graph neural network, and based on the abnormal perturbation of the encoded vector, to perform an active graph search on the multi-hop cognitive graph that conforms to the attack tactics, so as to identify and reconstruct the attack intent chain.

[0012] Preferably, the process by which the threat hunting analysis unit performs active graph search and reconstructs the attack intent chain includes, starting with an entity whose behavior encoding shows statistically significant deviation, exploring multiple steps and probabilistic paths along the associated edges on the multi-hop cognitive graph, calculating the threat confidence of each exploration path, and combining the deviation of the path nodes, the historical rarity of the associated edges, and the matching degree with the preset attack pattern. Paths with threat confidence exceeding a dynamic threshold are aggregated and semantically described to form an attack intent chain that includes attack steps, tactical intent, and impact assessment.

[0013] Preferably, the adaptive policy generator contains a meta-reinforcement learning policy engine to maintain a meta-policy network pre-trained in multiple scenarios. When the attack intent chain is input, the meta-policy network retrieves candidate actions from the response action knowledge graph based on the current system state, and generates a set of customized defense policy sequences that minimize the impact of the attack while taking into account business continuity through fast gradient updates.

[0014] Preferably, the strategy execution and optimization closed loop includes a strategy simulation and evaluation unit, which is specifically used for:

[0015] In a logically isolated digital twin sandbox environment, the defense strategy sequence is executed without loss of performance, and simulated attacks are injected to quantify the blocking effect, response latency, and potential impact on normal business. The quantitative evaluation results are converted into feedback signals, which are used to adjust the behavior encoding model parameters and anomaly judgment threshold of the threat hunting analysis unit in the network security situation awareness engine, and to update the policy network parameters of the meta-reinforcement learning policy engine in the adaptive policy generator.

[0016] Preferably, the operational steps of this big data-based network security operation monitoring and management method are as follows:

[0017] S1. Through a multi-source security data stream processing pipeline, multi-source security data is accessed and standardized in real time to generate a time-series event stream; S2. Through a network security situational awareness engine, behavioral modeling and threat correlation analysis are performed based on a dynamic knowledge graph to identify and output attack intent chains; S3. Through an adaptive policy generator, an adaptive defense policy sequence is generated based on the attack intent chain using meta-reinforcement learning; S4. Through a policy execution and optimization closed loop, the policy sequence is executed, and through simulation evaluation and feedback, the behavioral model and decision model are optimized in a closed loop.

[0018] Preferably, S2 includes constructing and continuously updating a multi-hop cognitive graph as a dynamic knowledge graph, using a temporal graph neural network to encode and model the behavior of entities in the graph to obtain a dynamic behavior baseline, and when a behavior deviation from the baseline is detected, initiating an active graph search on the cognitive graph to aggregate and reconstruct the discovered abnormal paths into an attack intent chain.

[0019] Preferably, step S3 includes mapping the attack intent chain to the current system state to a response action knowledge graph, invoking a pre-trained meta-reinforcement learning model, and quickly inferring the optimal defense action sequence, the goal of which is to minimize service interruption while suppressing the threat.

[0020] Preferably, S4 includes simulating the execution of a defense strategy sequence in a digital twin sandbox and evaluating security effectiveness and operational risks. Based on the evaluation results, a dual-path optimization signal is generated: the first path is used to correct the baseline or threshold in the behavior modeling, and the second path is used to update the decision strategy of the meta-reinforcement learning model.

[0021] Compared with the prior art, the beneficial effects of the present invention are:

[0022] 1. This invention transforms multi-source heterogeneous data into standardized event streams in real time, and performs deep behavioral modeling and threat hunting based on dynamic knowledge graphs to produce high-confidence attack intent chains. Then, by combining knowledge graphs and meta-reinforcement learning mechanisms, it generates adaptive defense strategies that balance security containment and business continuity. Finally, it verifies the effectiveness of the strategies through digital twin simulation and provides bidirectional feedback of the evaluation results to simultaneously optimize the front-end threat detection model and the back-end decision model, enabling the system to have continuous self-evolution capabilities and significantly enhancing the initiative, accuracy, and overall resilience of security operations.

[0023] 2. This invention constructs and maintains a multi-hop cognitive graph through a dynamic knowledge graph unit, integrating contextual information of assets, vulnerabilities, and historical interactions, providing a structured foundation for in-depth analysis. The threat hunting analysis unit uses a temporal graph neural network to dynamically encode and model entity behavior, thereby establishing an accurate behavioral baseline. When a statistically significant deviation in behavioral encoding is detected, the system can initiate an active graph search on the multi-hop cognitive graph that conforms to the attack tactics, starting from the abnormal entity. Through multi-step, probabilistic path exploration, the deviation of path nodes, the historical rarity of associated edges, and the matching degree with preset attack patterns are comprehensively calculated to obtain a high-confidence threat confidence, thereby significantly improving the ability to discover hidden and complex attack chains, the depth of analysis, and the operability of response decisions.

[0024] 3. This invention utilizes a meta-policy network pre-trained for multiple scenarios. Upon receiving an attack intent chain, it intelligently retrieves compliant candidate response actions from a knowledge graph based on real-time system status. Then, it fine-tunes and optimizes these actions through rapid gradient updates, thereby generating a customized defense strategy sequence within seconds. This not only significantly improves the automation level and decision-making speed of threat response, but also ensures, through optimized target design, that the generated strategy can effectively curb the spread of attacks, minimize security impact, and guarantee the continuity of critical business operations. This solves the prominent problems of rigid rules, poor adaptability, and neglect of business risks in traditional automated responses.

[0025] 4. This invention, through a strategy simulation evaluation unit, performs lossless simulation execution and adversarial testing of defense strategy sequences in a logically isolated digital twin sandbox, thereby achieving prior security verification of strategy effectiveness. It can accurately quantify the blocking effect, response latency, and potential impact on normal business, and transform this objective evaluation into a two-way feedback signal. On the one hand, it dynamically optimizes the behavior encoding model and anomaly judgment threshold of the threat hunting analysis unit in the situational awareness engine, improving detection sensitivity and accuracy. On the other hand, it updates the strategy network parameters of the meta-reinforcement learning engine in the adaptive strategy generator in real time, enhancing its intelligent decision-making capability. This not only avoids business risks caused by trial and error of defense measures in real networks, but also drives the two core capabilities of situational awareness and strategy generation to evolve synergistically in a closed loop, forming a continuously self-improving dynamic defense system. Attached Figure Description

[0026] Figure 1 This is a schematic diagram of the overall system architecture and data flow of the present invention;

[0027] Figure 2 This is a schematic diagram of the internal workflow of the network security situation awareness engine of the present invention;

[0028] Figure 3 This is a schematic diagram of the internal workflow of the adaptive strategy generator of the present invention;

[0029] Figure 4 This is a schematic diagram of the strategy execution and optimization closed loop of the present invention;

[0030] Figure 5 This is a schematic diagram illustrating the overall system workflow and method steps of the present invention. Detailed Implementation

[0031] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0032] Please see Figure 1The present invention provides an embodiment of a network security operation monitoring and management system based on big data, comprising: a multi-source security data stream processing pipeline for real-time access and parsing of heterogeneous security data, outputting a standardized time-series event stream; a network security situation awareness engine connected to the multi-source security data stream processing pipeline for continuously storing the event stream and performing entity behavior modeling and threat correlation analysis based on a dynamically evolving network knowledge graph to output a high-confidence attack intent chain; an adaptive policy generator connected to the network security situation awareness engine for receiving the attack intent chain and generating a targeted defense policy sequence based on a pre-built response action knowledge graph and a meta-reinforcement learning mechanism; and a policy execution and optimization closed loop connected to the adaptive policy generator for executing the defense policy sequence in a controllable environment and evaluating its effectiveness through digital twin simulation, and synchronously using feedback signals to optimize the behavior model of the network security situation awareness engine and the decision model of the adaptive policy generator.

[0033] Furthermore, after the user starts the system, the multi-source security data stream processing pipeline automatically connects to heterogeneous security data streams from firewalls, intrusion detection systems, endpoint logs, etc. The pipeline parses and unifies the data format in real time, outputting a standardized time-series event stream, thereby eliminating data silos and ensuring the consistency and timeliness of data in subsequent analysis. This lays an efficient foundation for the entire process. The standardized event stream is continuously input into the network security situation awareness engine. The engine relies on a dynamically evolving network knowledge graph to model the behavior of entities and associate potential threats in real time. Through the evolution and correlation analysis of the graph, the system automatically identifies abnormal patterns and outputs high-confidence attack intent chains, thereby significantly improving the accuracy and depth of threat detection. This helps users focus on key attack patterns from massive alerts. The attack intent chain is then fed into the adaptive policy generator. The generator combines a pre-built knowledge graph of response actions with a meta-reinforcement learning mechanism to quickly infer a sequence of defense strategies for the current attack chain. The strategy generation not only considers threat suppression but also business continuity, thereby achieving dynamic and intelligent response decisions. The defense strategy sequence is first executed in a lossless simulation within a digital twin sandbox in the strategy execution and optimization closed loop. The system simulates real network environments and attacks, quantitatively evaluates the blocking effect, response latency, and business impact of the strategies, and transforms the evaluation results into feedback signals. On the one hand, it optimizes the behavior model of the situational awareness engine, improves the sensitivity of future threat identification, and on the other hand, it updates the decision model of the strategy generator, enhances its ability to adapt to new threats, ensures the security and reliability of the strategies, and drives the overall system to continuously evolve in continuous use, forming a virtuous cycle of becoming smarter with use.

[0034] Please see Figure 2The present invention provides an embodiment of a network security operation monitoring and management system based on big data. The network security situation awareness engine includes a dynamic knowledge graph unit and a threat hunting analysis unit. The dynamic knowledge graph unit is used to construct and maintain a multi-hop cognitive graph that represents network assets, vulnerabilities, services and historical interaction relationships. The threat hunting analysis unit is used to dynamically encode the behavior patterns of entities in the multi-hop cognitive graph using a temporal graph neural network, and to perform active graph search on the multi-hop cognitive graph that conforms to attack tactics based on the abnormal perturbation of the encoded vector, so as to identify and reconstruct the attack intent chain.

[0035] Furthermore, after system startup, the dynamic knowledge graph unit continuously runs, automatically integrating network asset information, vulnerability databases, service dependencies, and historical interaction logs between entities to construct and update a panoramic multi-hop cognitive graph in real time. This graph not only statically describes who has what, but also dynamically depicts who interacts with whom and how, providing a structured contextual foundation for deep behavioral analysis. The threat hunting analysis unit utilizes a temporal graph neural network to continuously encode the behavior of entities in the cognitive graph. This process is not based on fixed rules, but rather on model learning to establish a dynamic and quantified behavioral baseline for each entity's normal behavioral patterns in the relational network. When an entity's behavioral encoding vector shows a statistically significant deviation, the system automatically marks it as a suspicious starting point, triggering an active graph search. Hunting is no longer passively waiting for alerts, but rather... Starting from the cognitive graph, the system explores multi-step, probabilistic paths along the associated edges to simulate potential paths for attackers to move laterally or escalate privileges. For each potential path explored, the system calculates its threat confidence. This calculation comprehensively considers three dimensions: the deviation of the behavior of each node on the path, the historical rarity of the associated edges used by the path, and the overall matching degree of the path with the preset attack tactical pattern. This effectively distinguishes between ordinary anomalies and chain behaviors with clear attack intentions. Finally, the system intelligently aggregates and semantically describes paths with threat confidence exceeding a dynamic threshold, ultimately forming a structured attack intent chain as output. This clearly illustrates the complete attack process, the underlying tactical objectives, and the potential impact, thereby elevating low-level discrete alerts to high-level, actionable threat intelligence.

[0036] In the threat hunting analysis unit of the network security situation awareness engine, the temporal graph neural network can be implemented using models such as graph attention network or graph convolutional network. Taking the graph attention network model as an example, its input is the adjacency matrix and node feature matrix of the multi-hop cognitive graph. It aggregates the neighbor node information through a multi-layer attention mechanism and outputs the dynamic behavior encoding vector of each entity node. The dataset used to train this model can include network interaction logs and asset status snapshots during the historical normal operation and maintenance cycle. The learning goal is to enable the model to accurately represent the behavior pattern of entities under normal conditions.

[0037] The process by which the threat hunting analysis unit performs active graph search and reconstructs the attack intent chain includes: starting with entities whose behavioral encoding shows statistically significant deviations, exploring multi-step, probabilistic paths along the associated edges on a multi-hop cognitive graph; calculating the threat confidence of each exploration path; the threat confidence calculation formula is: Threat Confidence = α × Average Node Deviation + β × Reciprocal of Edge Historical Rarity + γ × Attack Pattern Matching Degree, where α, β, and γ are adjustable weight coefficients, and α + β + γ = 1; Average Node Deviation refers to the average deviation of the current behavioral encoding of all nodes on the path from their historical baseline; Reciprocal of Edge Historical Rarity is the reciprocal of the frequency of the edges involved in the path in historical interactions; Attack Pattern Matching Degree is obtained by comparing the similarity between the path pattern and the tactical techniques in the knowledge base; and aggregating and semantically describing paths with threat confidence exceeding a dynamic threshold to form an attack intent chain that includes attack steps, tactical intent, and impact assessment.

[0038] Please see Figure 3 The present invention provides an embodiment of a network security operation monitoring and management system based on big data. The adaptive policy generator contains a meta-reinforcement learning policy engine, which maintains a meta-policy network pre-trained in multiple scenarios. The training process of the meta-reinforcement learning policy engine includes: collecting a dataset of historical attack scenarios, system states and corresponding effective response actions; using algorithms such as near-end policy optimization; and performing multi-task pre-training on the policy network in a simulated environment to enable it to learn basic response policies under different threat modes. When an attack intent chain is input, the meta-policy network retrieves candidate actions from the response action knowledge graph based on the current system state, and generates a set of customized defense policy sequences that minimize the impact of attacks and take into account business continuity through fast gradient updates.

[0039] Furthermore, when the attack intent chain is input into the adaptive policy generator, the system first calls its core meta-reinforcement learning policy engine. This engine has a built-in meta-policy network pre-trained in multiple scenarios, which means that the system does not start from scratch but has prior knowledge and a rapid adaptation foundation to deal with new threats. The meta-policy network integrates the current attack intent chain and the real-time system state, and then retrieves a set of logically related candidate defense actions from the pre-built response action knowledge graph. The graph ensures that the actions comply with security specifications and are interpretable. The engine does not directly apply fixed actions, but uses the retrieved candidate actions as a basis to fine-tune the meta-policy network through rapid gradient updates. Within seconds, a customized defense policy sequence is calculated, and this sequence optimizes two core objectives: minimizing the potential impact of the attack and maximizing business continuity. Finally, the system outputs this policy sequence, which is not only a list of actions, but also clarifies the logical order, conditions and timing of execution. For example, firstly, suspicious IPs are temporarily blocked from accessing non-business ports of the database at the edge routing layer. At the same time, a deep forensics process is started on the server. If subsequent data leakage attempts are detected, network isolation is immediately initiated.

[0040] In the meta-reinforcement learning policy engine of the adaptive policy generator, meta-reinforcement learning can adopt algorithmic frameworks such as proximal policy optimization. During the pre-training phase, its meta-policy network is trained in various simulated attack scenarios to learn the mapping from attack features and system states to basic defense actions. After pre-training, when faced with a new attack intent chain, the network can quickly adjust policy parameters based on candidate actions retrieved from the response action knowledge graph through a few gradient updates, generating a customized defense policy sequence that balances security containment and business impact for the current specific threat.

[0041] Please see Figure 4 The present invention provides an embodiment of a network security operation monitoring and management system based on big data. The policy execution and optimization closed loop includes a policy simulation evaluation unit, which is specifically used to: perform lossless simulation execution of defense policy sequences in a logically isolated digital twin sandbox environment, and inject simulated attacks to quantify the blocking effect, response delay and potential impact on normal business of the policy, and convert the quantitative evaluation results into feedback signals. On the one hand, it is used to adjust the behavior encoding model parameters and anomaly judgment threshold of the threat hunting analysis unit in the network security situation awareness engine, and on the other hand, it is used to update the policy network parameters of the meta-reinforcement learning policy engine in the adaptive policy generator.

[0042] Furthermore, upon receiving a defense strategy sequence, the system does not execute it directly in the real network. Instead, it imports it into a logically isolated digital twin sandbox environment, which is a precise mirror of the real network topology and service status. In this environment, the system first performs lossless simulation execution of the strategy sequence while simultaneously injecting simulated attacks to securely trigger the defense strategy and observe its actual interaction. The system automatically collects data during the simulation process and quantitatively evaluates the strategy from three core dimensions: blocking effect, response latency, and potential impact on normal services. The system then transforms the strategy effect into objective and measurable indicators, converts the quantitative evaluation results into structured feedback signals, and simultaneously drives optimization in two key areas:

[0043] First-path optimization: Feedback signals are used to adjust the internal parameters of the threat hunting analysis unit in the network security situation awareness engine. For example, the parameters of the behavior coding model are corrected to make it more sensitive to abnormal behavior patterns that lead to the effectiveness of the policy, or the anomaly judgment threshold is dynamically adjusted to reduce similar false alarms or missed detections in the future.

[0044] The second optimization approach involves using feedback signals to update the policy network parameters of the meta-reinforcement learning policy engine in the adaptive policy generator. Successful policy actions are reinforced, while ineffective or negatively impactful actions are weakened, enabling the policy generation model to make better decisions when facing similar threats in the future.

[0045] Please see Figure 5 The present invention provides an embodiment of a network security operation monitoring and management method based on big data. The working steps of the network security operation monitoring and management method based on big data are as follows:

[0046] S1. Through a multi-source security data stream processing pipeline, multi-source security data is accessed and standardized in real time to generate a time-series event stream; S2. Through a network security situational awareness engine, behavioral modeling and threat correlation analysis are performed based on a dynamic knowledge graph to identify and output attack intent chains; S3. Through an adaptive policy generator, an adaptive defense policy sequence is generated based on the attack intent chain using meta-reinforcement learning; S4. Through a policy execution and optimization closed loop, the policy sequence is executed, and through simulation evaluation and feedback, the behavioral model and decision model are optimized in a closed loop.

[0047] S2 includes constructing and continuously updating a multi-hop cognitive graph as a dynamic knowledge graph, using a temporal graph neural network to encode and model the behavior of entities in the graph to obtain a dynamic behavior baseline, and when a behavior deviation from the baseline is detected, initiating an active graph search on the cognitive graph to aggregate and reconstruct the discovered abnormal paths into an attack intent chain.

[0048] S3 includes mapping the attack intent chain to the current system state to a response action knowledge graph, calling a pre-trained meta-reinforcement learning model, and quickly inferring the optimal defense action sequence. The goal of the sequence is to suppress the threat while minimizing business interruption.

[0049] S4 includes simulating the execution of defense strategy sequences in a digital twin sandbox and evaluating their security effectiveness and operational risks. Based on the evaluation results, a dual-path optimization signal is generated. The first path is used to correct the baseline or threshold in the behavior modeling, and the second path is used to update the decision strategy of the meta-reinforcement learning model.

[0050] Furthermore, upon system startup, the multi-source security data stream processing pipeline begins operation, automatically and in parallel accessing various heterogeneous logs and traffic data from cloud platforms, network devices, terminal hosts, and application systems. Through a built-in parser and standardized rules, the chaotic data is transformed in real-time into a unified formatted time-series event stream with precise timestamps, achieving efficient and accurate analysis throughout the entire process. This resolves the analysis delays and errors caused by inconsistent data formats. Simultaneously, within the network security situation awareness engine, the system first constructs and continuously updates a multi-hop cognitive graph integrating assets, vulnerabilities, and entity relationships. Second, it analyzes the historical behavior of entities in the graph using a time-series graph neural network, automatically learning and establishing dynamic behavioral baselines. When any entity's behavior significantly deviates from the baseline, the engine immediately initiates an active graph search on the cognitive graph, starting from that entity, to explore suspicious related paths. Ultimately, multiple high-threat paths are aggregated, semantically represented, and reconstructed into a clear attack intent chain, generating a comprehensive attack strategy. The attack intent chain and the current system state are input into the adaptive policy generator. The system first maps the attack context to the response action knowledge graph to obtain a set of compliant candidate actions. Then, it calls the pre-trained meta-reinforcement learning model for rapid inference. Within seconds, the model outputs an optimal defense action sequence through internal fine-tuning. This sequence aims to maximize threat blocking while embedding a trade-off logic to minimize business interruption, achieving intelligent and accurate response decisions. The policy sequence is sent into the policy execution and optimization closed loop. The system performs lossless simulation execution of the policy in a digital twin sandbox and injects simulated attacks to accurately assess its security effectiveness and operational risks. After the assessment, the system automatically generates dual-path optimization signals. The first signal is transmitted in reverse to the situational awareness engine to correct the dynamic baseline or anomaly judgment threshold of behavior modeling and improve the accuracy of future detection. The second signal is used to update the decision policy parameters of the meta-reinforcement learning model to enhance its decision-making ability to deal with similar threats.

[0051] The system works by transforming multi-source heterogeneous data into standardized event streams in real time, and then performing deep behavioral modeling and threat hunting based on a dynamic knowledge graph to generate high-confidence attack intent chains. This is followed by combining the knowledge graph with meta-reinforcement learning mechanisms to generate adaptive defense strategies that balance security containment and business continuity. Finally, the effectiveness of these strategies is verified through digital twin simulation, and the evaluation results are fed back bidirectionally to simultaneously optimize both the front-end threat detection model and the back-end decision-making model. This enables the system to continuously evolve, significantly enhancing the initiative, accuracy, and overall resilience of security operations. The multi-hop cognitive graph, constructed and maintained through dynamic knowledge graph units, integrates contextual information about assets, vulnerabilities, and historical interactions. This information provides a structured foundation for in-depth analysis. The threat hunting analysis unit uses a temporal graph neural network to dynamically encode and model entity behavior, thereby establishing a precise behavioral baseline. When a statistically significant deviation in the behavioral encoding is detected, the system can initiate an active graph search on a multi-hop cognitive graph that conforms to the attack tactics, starting from the abnormal entity. Through multi-step, probabilistic path exploration, it comprehensively calculates the deviation of path nodes, the historical rarity of associated edges, and the matching degree with preset attack patterns to obtain a high-confidence threat confidence. This significantly improves the ability to detect, analyze, and respond to covert and complex attack chains, as well as the operability of the decision-making. By utilizing a meta-policy network pre-trained in multiple scenarios, it can receive... Upon reaching the attack intent chain, and combining real-time system status, compliant candidate response actions are intelligently retrieved from the knowledge graph. Based on this, fine-tuning and optimization are performed through rapid gradient updates, generating a customized defense strategy sequence within seconds. This not only significantly improves the automation level and decision-making speed of threat response, but also ensures, through optimized target design, that the generated strategy can effectively curb attack spread, minimize security impact, and strive to ensure the continuity of critical business operations. This solves the prominent problems of rigid rules, poor adaptability, and neglect of business risks in traditional automated responses. Through the strategy simulation and evaluation unit, the defense strategy sequence is simulated and executed without loss in a logically isolated digital twin sandbox. By resisting testing, the system achieves prior security verification of the strategy's effectiveness. It can accurately quantify the blocking effect, response latency, and potential impact on normal business operations, and transform this objective assessment into a two-way feedback signal. On the one hand, it dynamically optimizes the behavior encoding model and anomaly judgment threshold of the threat hunting analysis unit in the situational awareness engine, improving detection sensitivity and accuracy. On the other hand, it updates the policy network parameters of the meta-reinforcement learning engine in the adaptive policy generator in real time, enhancing its intelligent decision-making capabilities. This not only avoids business risks caused by trial and error of defense measures in real networks, but also drives the two core capabilities of situational awareness and policy generation to evolve synergistically in a closed loop, forming a continuously self-improving dynamic defense system.

[0052] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, it is intended that all variations falling within the meaning and scope of equivalents of the claims be included within the present invention.

Claims

1. A network security operation monitoring and management system based on big data, characterized in that: include: A multi-source secure data stream processing pipeline is used to access and parse heterogeneous secure data in real time and output a standardized time-series event stream. The network security situation awareness engine connects to a multi-source security data stream processing pipeline and is used for entity behavior modeling and threat correlation analysis based on a dynamically evolving network knowledge graph in the form of a multi-hop cognitive graph. It outputs an attack intent chain, and the multi-hop cognitive graph integrates network assets, vulnerabilities, services and their historical interaction relationships. An adaptive policy generator, connected to a network security situation awareness engine, is used to receive the attack intent chain and generate a targeted defense policy sequence based on a pre-built response action knowledge graph and a meta-reinforcement learning mechanism. The strategy execution and optimization closed loop connects to the adaptive strategy generator, which is used to execute the defense strategy sequence in a controllable environment and evaluate the effect through digital twin simulation. The feedback signal is then used to optimize the behavior model of the network security situation awareness engine and the decision model of the adaptive strategy generator.

2. The network security operation monitoring and management system based on big data according to claim 1, characterized in that: The network security situation awareness engine includes a dynamic knowledge graph unit and a threat hunting analysis unit. The dynamic knowledge graph unit is used to construct and maintain the multi-hop cognitive graph. The threat hunting analysis unit is used to dynamically encode the behavioral patterns of entities in the multi-hop cognitive graph using a temporal graph neural network based on graph attention network or graph convolutional network, and to perform active graph search on the multi-hop cognitive graph based on abnormal perturbations of the encoded vectors to identify and reconstruct the attack intent chain.

3. The network security operation monitoring and management system based on big data according to claim 2, characterized in that: The process by which the threat hunting analysis unit performs active graph search and reconstructs the attack intent chain includes: starting with entities whose behavioral encoding shows statistically significant deviations, exploring multi-step, probabilistic paths along the associated edges on the multi-hop cognitive graph, calculating the threat confidence of each exploration path, and calculating the threat confidence by comprehensively considering the deviation of path nodes, the historical rarity of associated edges, and the matching degree with a preset attack pattern. The specific calculation formula is: Threat Confidence = α × Average Node Deviation + β × Reciprocal of Historical Rarity of Edges + γ × Attack Pattern Matching Degree, where α, β, and γ are adjustable weight coefficients and satisfy α + β + γ = 1. Paths with threat confidence exceeding a dynamic threshold are aggregated and semantically described to form an attack intent chain.

4. The network security operation monitoring and management system based on big data according to claim 1, characterized in that: The adaptive policy generator contains a unary reinforcement learning policy engine to maintain a meta-policy network pre-trained in multiple scenarios. When the attack intent chain is input, the meta-policy network retrieves candidate actions from the response action knowledge graph based on the current system state, and generates a set of defense policy sequences that minimize the impact of the attack while taking into account business continuity through fast gradient updates.

5. The network security operation monitoring and management system based on big data according to claim 1, characterized in that: The strategy execution and optimization closed loop includes a strategy simulation and evaluation unit, which is specifically used for: In a logically isolated digital twin sandbox environment, the defense strategy sequence is executed without loss of quality, and simulated attacks are injected to quantify the blocking effect, response latency, and potential impact on normal business operations. The quantitative evaluation results are converted into feedback signals, which are used to adjust the behavior encoding model parameters and anomaly detection threshold of the threat hunting analysis unit in the network security situation awareness engine, and to update the policy network parameters of the meta-reinforcement learning policy engine in the adaptive policy generator.

6. A big data-based network security operation monitoring and management method, applicable to the big data-based network security operation monitoring and management system described in any one of claims 1-5, characterized in that: The working steps of this big data-based network security operation monitoring and management method are as follows: S1. Through the multi-source security data stream processing pipeline, multi-source security data is accessed and standardized in real time to generate a time-series event stream; S2. Through the network security situation awareness engine, construct and maintain a multi-hop cognitive graph that integrates network assets, vulnerabilities and historical interaction relationships. Use a temporal graph neural network to encode and model the behavior of entities in the graph to obtain a dynamic behavior baseline. When a behavior deviates from the baseline, initiate an active graph search on the multi-hop cognitive graph to aggregate and reconstruct the discovered abnormal paths into an attack intent chain. S3. Using an adaptive policy generator, the attack intent chain is mapped to the current system state and then to the response action knowledge graph. A pre-trained meta-reinforcement learning model is invoked to quickly infer and generate a defense policy sequence. S4. Through strategy execution and optimization closed loop, the defense strategy sequence is simulated and executed in the digital twin sandbox, and simulated attacks are injected for adversarial testing to evaluate security effectiveness and operational risks. Feedback signals are generated based on the evaluation results to optimize the behavior encoding model and meta-reinforcement learning decision-making strategy in a closed loop.

7. The network security operation monitoring and management method based on big data according to claim 6, characterized in that: In step S2, the threat confidence of the path obtained by the active graph search is calculated. The threat confidence is a combination of the deviation of the path nodes, the historical rarity of the associated edges, and the matching degree with the preset attack pattern.

8. The network security operation monitoring and management method based on big data according to claim 6, characterized in that: In S3, the meta-reinforcement learning model is a meta-policy network pre-trained by a proximal policy optimization algorithm, which generates the defense policy sequence through fast gradient updates.

9. A network security operation monitoring and management method based on big data according to claim 6, characterized in that: In step S4, the feedback signal includes a first optimization signal and a second optimization signal. The first optimization signal is used to correct the dynamic baseline or anomaly detection threshold in behavior modeling, and the second optimization signal is used to update the policy network parameters of the meta-reinforcement learning model.