Network security operation management platform
By using components such as an intelligent decision-making center and a digital twin module, dynamic response strategies are generated, which solves the problem of automated response interruption of the SOAR system under unknown attacks and achieves efficient security operation and management.
Patent Information
- Application Number
- CN202511295157.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-11
- Publication Date
- 2026-01-16
AI Technical Summary
Existing SOAR systems lack adaptability in the face of unknown attacks, leading to interruptions in automated responses, reliance on human experience, and low operational efficiency.
Employing components such as an intelligent decision-making hub, a digital operation twin module, and a feedback and evolutionary learning module, the system generates dynamic response strategies through reinforcement learning models. Combined with a virtual simulation environment and natural language interaction, it achieves automated response and strategy optimization for security incidents.
It enables efficient automatic response in unknown attack scenarios, reduces reliance on human experience, and improves security operation efficiency and the ability to combat new types of attacks.
Smart Images

Figure CN121356809A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of network security technology, and specifically relates to a network security operation and management platform. Background Technology
[0002] With the continuous evolution of network attack techniques, sophisticated new attacks such as Advanced Persistent Threats (APTs) and zero-day attacks are emerging in an endless stream, placing higher demands on network security operations. Currently, SOAR (Security Orchestration, Automation, and Response) technology, as a mainstream security operations tool, automates security incident responses through predefined scripts, which alleviates the workload of security operations personnel to some extent. However, predefined scripts can only cover known attack scenarios. For unprecedented zero-day attacks and APT attacks, the scripts become completely ineffective due to the lack of corresponding handling logic. In such cases, security analysts need to intervene manually and formulate temporary handling plans based on experience, leading to automation interruptions and a significant decrease in operational efficiency. To address this, a network security operations management platform has been designed and improved. Summary of the Invention
[0003] To address the aforementioned shortcomings of existing technologies, this invention solves the problems of rigid scripts, poor adaptability, and reliance on human experience in existing SOAR systems by synergizing core components such as an intelligent decision-making center, a digital operation twin module, and a feedback and evolutionary learning module. This enables dynamic response to security incidents, simulation verification of strategies, and continuous system evolution, reducing reliance on human intervention and improving security operation efficiency and the ability to combat new types of attacks.
[0004] To solve the above-mentioned technical problems, the present invention adopts the following technical solution: include: The intelligent decision-making center is used to receive security events and operational data from the security toolchain, and generate dynamic and adaptive response strategies based on preset operational goals, constraints and reinforcement learning models. The strategy execution engine, which is communicatively connected to the intelligent decision-making center, is used to decompose the response strategy into executable task instructions and schedule and drive the corresponding tools in the underlying security toolchain to execute the task instructions. The digital operations twin module is communicatively connected to the intelligent decision-making center and the strategy execution engine. It is used to build and maintain a virtual simulation environment that is synchronized with the real security operation environment. The intelligent decision-making center uses the digital operations twin module to simulate and evaluate the effects of candidate response strategies. The feedback and evolutionary learning module is used to collect the execution result data of the strategy execution engine in the real environment and the feedback from the operators, and to continuously train and optimize the reinforcement learning model in the intelligent decision-making center. The Natural Language Interaction and Operational Knowledge Generation Module is used to receive natural language instructions or queries from operations personnel and call the corresponding functions of the platform to process or respond. At the same time, it automatically extracts and summarizes knowledge from security events, response actions and final results to generate readable operational reports and strategy summaries.
[0005] Preferably, the intelligent decision-making center includes: The strategy generation unit has a built-in reinforcement learning algorithm to generate multiple candidate strategies based on the current security situation, operational objectives and historical data; The strategy evaluation unit interacts with the digital operation twin module to perform parallel simulation and deduction of candidate strategies in the digital operation twin module, and outputs the optimal response strategy based on the deduction results.
[0006] Preferably, the evaluation indicators used by the strategy evaluation unit include the degree of risk reduction, business impact index, resource consumption cost, and response efficiency.
[0007] Preferably, the digital operation twin module synchronizes with real security toolchains and network topology asset data through API interfaces, and uses simulation technology to simulate attack behaviors, defense actions, and their interactive effects.
[0008] Preferably, the feedback and evolutionary learning module quantifies the strategy execution effect by establishing a reward function, which integrates multiple dimensions such as risk reduction, operational efficiency improvement, resource consumption cost, and human feedback score.
[0009] Preferably, it further includes: The human-machine collaboration interface is used to present the strategies recommended by the intelligent decision-making center, their simulation results, and evaluation criteria to the operations personnel, and to receive the final decision instructions, modification opinions, or feedback scores from the operations personnel.
[0010] Preferably, the human-machine collaboration interface displays the attack path, strategy deduction process, and expected impact in a visual manner.
[0011] Preferably, the natural language interaction and operational knowledge generation module integrates a large language model to generate a comprehensive report based on instructions, including root causes, actions taken, and effect evaluations.
[0012] Preferably, it further includes: The unified data lake is used to store and govern raw event data collected from the underlying security toolchain, policy data generated by the platform, simulation data, execution result data, and feedback data, providing data support for the intelligent decision-making center, digital operation twin module, and feedback and evolutionary learning module.
[0013] It also includes the following steps: Receive security incident and operational data; Based on reinforcement learning models and digital operations twin simulations, dynamic response strategies are generated and evaluated. Execute the preferred response strategy or the strategy confirmed by humans; Collect execution feedback data to optimize the reinforcement learning model; Receive queries or instructions through natural language interaction and generate operational knowledge reports.
[0014] Compared with the prior art, the present invention has the following beneficial effects: 1. This platform is based on an intelligent decision-making center reinforcement learning model. It can perform deep neural network fitting of Q-function and nearest neighbor policy optimization algorithm transformation based on attack characteristics collected in real time from a unified data lake. Then, it quickly constructs the state-action space of the attack scenario through the PPO algorithm to generate disposal strategies. Then, through the 1:1 virtual mirror constructed by the digital operation twin module, it can perform "attack and defense simulation" on candidate strategies without affecting real business. By quantitatively calculating indicators such as risk reduction degree and business impact index, it selects the strategy with the best defense effect and the least business loss, ensuring a balance between core business security and defense effect. Then, through the task decomposition module of the strategy execution engine, it can automatically identify action logic dependencies and generate the optimal task sequence. It can also avoid execution delays caused by resource conflicts through parameter optimization. At the same time, the tool scheduling module automatically matches the optimal tool and configures backup plans based on the three-dimensional filtering logic of "availability-load rate-efficiency", thereby shortening the strategy execution time. During the strategy execution, the automatic anomaly handling mechanism of the monitoring module ensures that the entire task will not fail to execute when a certain task fails. The execution failure can be quickly marked through visualization, which is convenient for operation personnel to quickly locate. This allows operation personnel to troubleshoot more efficiently and significantly shortens the time required for troubleshooting. 2. By setting up a feedback and evolutionary learning module, the platform quantifies the strategy execution results into model training data, continuously updates the reinforcement learning model, gradually reduces the platform's reliance on security analysts, and improves operational capabilities. Furthermore, through a human-machine collaboration interface, the platform supports intervention in strategies for high-risk scenarios such as core databases and trading systems, and human feedback can be directly used to optimize the reinforcement learning model. This ensures business security in high-risk scenarios while preventing the loss of experience due to staff turnover. Attached Figure Description
[0015] Figure 1 This is a framework diagram of a network security operation and management platform according to the present invention. Detailed Implementation
[0016] To enable those skilled in the art to better understand the present invention, the technical solution of the present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0017] The accompanying drawings are for illustrative purposes only and are schematic diagrams, not actual images. They should not be construed as limiting the scope of this patent. To better illustrate the embodiments of the present invention, some parts in the drawings may be omitted, enlarged, or reduced, and do not represent the actual dimensions of the product. It is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.
[0018] In the accompanying drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components. In the description of the present invention, it should be understood that if terms such as "upper," "lower," "left," "right," "inner," and "outer" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, they are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, the terms used to describe positional relationships in the drawings are only for illustrative purposes and should not be construed as limiting the present patent. For those skilled in the art, the specific meaning of the above terms can be understood according to the specific circumstances.
[0019] In the description of this invention, unless otherwise explicitly specified and limited, the term "connection" or similar designation indicating a connection between components should be interpreted broadly. For example, it can refer to a fixed connection, a detachable connection, or an integral part; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium; it can refer to the internal communication between two components or the interaction between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0020] Example 1: like Figure 1 The present invention includes a network security operation and management platform, comprising: The intelligent decision-making center is used to receive security events and operational data from the security toolchain, and generate dynamic and adaptive response strategies based on preset operational goals, constraints and reinforcement learning models. The strategy execution engine, which is communicatively connected to the intelligent decision-making center, is used to decompose the response strategy into executable task instructions and schedule and drive the corresponding tools in the underlying security toolchain to execute the task instructions. The digital operations twin module is communicatively connected to the intelligent decision-making center and the strategy execution engine. It is used to build and maintain a virtual simulation environment that is synchronized with the real security operation environment. The intelligent decision-making center uses the digital operations twin module to simulate and evaluate the effects of candidate response strategies. The feedback and evolutionary learning module is used to collect the execution result data of the strategy execution engine in the real environment and the feedback from the operators, and to continuously train and optimize the reinforcement learning model in the intelligent decision-making center. The Natural Language Interaction and Operational Knowledge Generation Module is used to receive natural language instructions or queries from operations personnel and call the corresponding functions of the platform to process or respond. At the same time, it automatically extracts and summarizes knowledge from security events, response actions and final results to generate readable operational reports and strategy summaries.
[0021] The intelligent decision-making center includes: The strategy generation unit has a built-in reinforcement learning algorithm to generate multiple candidate strategies based on the current security situation, operational objectives and historical data; The strategy evaluation unit interacts with the digital operation twin module to perform parallel simulation and deduction of candidate strategies in the digital operation twin module, and outputs the optimal response strategy based on the deduction results.
[0022] The evaluation indicators used by the strategy evaluation unit include the degree of risk reduction, business impact index, resource consumption cost, and response efficiency.
[0023] The digital operations twin module synchronizes with real security toolchains and network topology asset data through API interfaces, and uses simulation technology to simulate attack behaviors, defense actions, and their interactive effects.
[0024] The feedback and evolutionary learning module quantifies the strategy execution effect by establishing a reward function, which integrates multiple dimensions such as risk reduction, operational efficiency improvement, resource consumption cost, and human feedback score.
[0025] Also includes: The human-machine collaboration interface is used to present the strategies recommended by the intelligent decision-making center, their simulation results, and evaluation criteria to the operations personnel, and to receive the final decision instructions, modification opinions, or feedback scores from the operations personnel.
[0026] The human-machine collaboration interface visualizes the attack path, strategy deduction process, and expected impact.
[0027] The natural language interaction and operational knowledge generation module integrates a large language model and generates a comprehensive report based on instructions, including root causes, actions taken, and effect evaluations.
[0028] Also includes: The unified data lake is used to store and govern raw event data collected from the underlying security toolchain, policy data generated by the platform, simulation data, execution result data, and feedback data, providing data support for the intelligent decision-making center, digital operation twin module, and feedback and evolutionary learning module.
[0029] The operation method of the network security operation management platform includes the following steps: Receive security incident and operational data; Based on reinforcement learning models and digital operations twin simulations, dynamic response strategies are generated and evaluated. Specifically, the intelligent decision-making center includes a strategy generation unit and a strategy evaluation unit. The strategy generation unit automatically generates multiple candidate response strategies adapted to the current scenario based on real-time security posture and historical experience. The unit incorporates two types of algorithms to address different security scenarios. The first algorithm is a deep Q-network, suitable for scenarios with simple attack scenarios and small action spaces. By fitting the Q-function through a deep neural network, it can quickly output the short-term optimal action combination. The second algorithm is a nearest neighbor policy optimization algorithm, suitable for scenarios with complex attack scenarios and large action spaces. This algorithm controls the policy update magnitude by pruning proxy targets and supports continuous action spaces, thus generating more refined strategies in complex scenarios. The state space is designed using a multi-dimensional vector encoding method, with the value of each dimension determined by real-time data collection from a unified data lake. The action space adopts a layered action library design. The basic action layer contains various atomic operations, all verified through security tool APIs. The combined action layer is composed of the basic action layer according to logical dependencies, such as backing up data before isolating the host, excluding actions based on business constraints, and finally generating candidate strategies through the PPO iterative algorithm. The strategy evaluation unit establishes a multi-dimensional evaluation model and constructs a 1:1 virtual mirror of the real environment through the digital operation twin module. Candidate strategies are executed in the virtual environment, and four types of data, namely "risk changes, business impact, resource consumption, and execution time", are collected to calculate the comprehensive score of the strategy. Finally, the strategy with the highest score is selected. The formula for the degree of risk reduction is: (Probability of attack success before simulation - Probability of attack success after simulation) / Probability of attack success before simulation × 100%; The business impact index formula is: business interruption time × asset importance weight. The formula for resource consumption cost is: CPU utilization rate × duration + manual review time × average hourly wage per employee / business tolerance threshold × 100% The execution time is the time from policy initiation to complete containment of the attack. The comprehensive score is calculated as follows: Risk reduction level × Risk weight + (1 - Business impact index / 100) × Business weight + (1 - Resource consumption cost) × Resource weight + (1 - Response efficiency / Maximum tolerable duration) × Time consumption weight.
[0030] The strategy execution engine is responsible for transforming the abstract strategies generated by the intelligent decision-making center into executable tasks and driving the execution of underlying tools. The strategy execution engine includes a task decomposition module, a tool scheduling module, and a monitoring module. The task decomposition module has a built-in conflict detection unit and parameter optimization unit. When decomposing tasks, the task decomposition module focuses on logical dependencies and resource conflicts, breaks down the strategy according to the actions of the combined action layer, and sorts the actions based on dependencies to obtain a task sequence. The conflict detection unit detects resource and logical conflicts between tasks. The parameter optimization module optimizes task parameters based on real-time environmental data to ensure that the core CPU is not overloaded under multi-tasking conditions. The tool scheduling module selects tools through a tool capability library and scheduling algorithms. The tool capability library stores a list of all integrated security tools' capabilities, interface information, and status information. Different tools are selected based on different task types within the task sequence. When multiple tools are available, they are filtered based on availability, load rate, and execution efficiency to determine the optimal and backup tools. The monitoring module tracks task execution in real time and presents a visual overview of the task execution, facilitating quick problem identification by operations personnel. When a temporary error is detected, a retry mechanism is automatically triggered. When a tool failure is detected, a backup tool is selected from the tool capability library. When an error is detected where retry is not enabled, the anomaly information is automatically fed back to the intelligent decision-making center, requesting the generation of alternative strategies. Furthermore, the entire task processing log is automatically recorded.
[0031] The digital operations twin module is based on a three-element architecture of physical entity-virtual image-data interaction. It constructs a precise virtual environment through real-time data synchronization, recreates the attack and defense interaction process using simulation algorithms, and ultimately uses quantitative evaluation results to feed back into decision optimization. First, it collects four categories of data from the real security environment: network topology, asset information, security tool status, and attack events. Then, it standardizes the non-standardized data output from different systems through data cleaning, field mapping, and format conversion. The standardized real data is then mapped to "digital entities" in the virtual environment, constructing a 1:1 virtual image. It captures data changes in the real environment in real time through API interfaces and updates them synchronously in the virtual environment, updating every 30 minutes. The virtual environment and the real environment's full data are verified once. If inconsistencies are found, full synchronization is automatically triggered to ensure the accuracy of the virtual image. Then, based on the current security situation input by the intelligent decision center, corresponding attack and defense scenarios are constructed in the virtual environment. The candidate response strategies generated by the intelligent decision center are transformed into a sequence of defensive actions that can be executed in the virtual environment and injected into the virtual security tool. The attack behavior simulation engine and defense action simulation engine in the virtual environment are started to simulate the real-time interaction process between the attacker and defender. During the simulation, four types of key data are collected in real time to form a complete simulation dataset. The degree of risk reduction, the improvement in operational efficiency, and the resource consumption cost are calculated based on the data before and after the strategy is executed.
[0032] Feedback and Evolutionary Learning Module: This module collects task execution results, business system status, and resource consumption data from the strategy execution engine, based on real-world execution data. It also collects scores and feedback submitted by operators through a human-machine collaboration interface. After preprocessing, the data is stored in a unified data lake. A multi-dimensional reward function quantifies the strategy execution effect. The reward function formula is as follows: ,in The weights are used to input the reward value and the corresponding state-data into the policy generation unit, and the model of the policy generation unit is updated by the gradient descent algorithm.
[0033] The Natural Language Interaction and Operational Knowledge Generation Module integrates a large language model for natural voice interaction. This model translates operators' natural voice commands or query commands into executable machine commands for the platform, while simultaneously converting the platform's structured data into easily understandable natural language responses. When operators perform searches, the module extracts relevant information from raw data stored in a unified data lake, including security events, response strategies, and execution results. It then uses knowledge graph technology to transform this information into structured knowledge, which is subsequently output as standardized documents such as operational reports.
[0034] Human-Machine Collaboration Interface: This interface transforms the optimal strategies, simulation data, and evaluation criteria output by the intelligent decision-making center into a visual and easily understandable format, providing sufficient evidence for human decision-making. It also marks the penetration paths of attacks with red arrows and grants operational personnel permissions to ensure that humans can interfere with key decisions. It provides a scoring interface and feedback input box, allowing operational personnel to evaluate the effectiveness of strategy execution. The decision results and strategy evaluation data are then synchronized to the feedback and evolutionary learning module to train the model.
[0035] Unified Data Lake: This is used to store raw data such as IDS alerts, EDR logs, and vulnerability scan reports collected from the underlying security toolchain; policy data such as candidate policies, optimal policies, and policy evaluation data generated by the intelligent decision-making center; simulation data such as virtual environment configuration and simulation logs of the digital operations twin module; execution data such as task status and tool execution results of the policy execution engine; and feedback data such as quantitative scores and human opinions collected by the feedback and evolutionary learning module. It also provides data query and call interfaces for each module.
[0036] Working Principle: The unified data lake collects raw event data, network topology and asset configuration data, and historical operational data from underlying security tools in real time via standardized APIs. After cleaning, deduplication, and format standardization, the data is categorized and stored according to data type, and an associated index is established. The intelligent decision-making center constructs a state and action space based on the real-time security posture of the unified data lake, and determines whether the attack method is known. Based on two different algorithms, it generates strategies and then synchronizes with the real environment through the digital operations twin module to build a 1:1 virtual mirror, reproducing the attack and defense scenario. Strategy simulations are injected, collecting data on "risk, business impact, resource consumption, and time consumption," and applying multi-dimensional formulas... The strategy's overall score is calculated, and the strategy with the highest score is selected as the optimal strategy. The strategy execution engine breaks down the optimal strategy into a sequence of atomic tasks, while simultaneously detecting resource and logical conflicts and optimizing parameters. It searches the tool capability library and matches suitable tools, selecting the optimal tool based on "availability, load, and efficiency" and equipping alternative tools. The tasks are tracked in real time, and the strategy execution status is visualized. Operations personnel can confirm, modify, and reject strategies through a human-machine collaboration interface. The feedback and evolutionary learning module collects execution results and human scores, quantifies the strategy value according to the reward function, updates the reinforcement learning model parameters using gradient descent, and iteratively improves the accuracy of strategy generation.
[0037] The above are merely embodiments of the present invention. The circuits, electronic components, and modules involved are all prior art, fully achievable by those skilled in the art, and require no further explanation. The scope of protection in this application does not involve improvements to the software and methods. Commonly known structures and characteristics in the solutions are not described in detail here. Those skilled in the art are aware of all common technical knowledge in the field prior to the application date or priority date, are aware of all prior art in that field, and have the ability to apply conventional experimental methods prior to that date. Those skilled in the art can, under the guidance of this application, improve and implement this solution in combination with their own capabilities. Some typical known structures or methods should not be obstacles for those skilled in the art to implement this application. It should be noted that those skilled in the art can make several modifications and improvements without departing from the structure of the present invention. These should also be considered within the scope of protection of the present invention, and will not affect the effectiveness of the implementation of the present invention or the practicality of the patent.
Claims
1. A network security operations management platform, characterized by: Comprise: An intelligent decision hub for receiving security events and operational data from a security tool chain and generating dynamic and adaptive response strategies based on preset operational objectives, constraints and reinforcement learning models; A policy execution engine in communication with the intelligent decision hub for decomposing the response strategies into executable task instructions and scheduling and driving corresponding tools in the underlying security tool chain to execute the task instructions; A digital operational twin module in communication with the intelligent decision hub and the policy execution engine for building and maintaining a virtual simulation environment synchronized with the real security operational environment, which is used by the intelligent decision hub to simulate and evaluate the effects of candidate response strategies; A feedback and evolutionary learning module for collecting execution result data of the policy execution engine in the real environment and feedback from operational personnel and continuously training and optimizing the reinforcement learning model in the intelligent decision hub; A natural language interaction and operational knowledge generation module for receiving natural language instructions or queries from operational personnel and invoking corresponding platform functions for processing or responding, while automatically extracting and summarizing knowledge from security events, response actions and final results to generate readable operational reports and strategy summaries.
2. The network security operations management platform of claim 1, wherein: The intelligent decision hub comprises: A strategy generation unit with built-in reinforcement learning algorithms for generating multiple candidate strategies based on current security situation, operational objectives and historical data; A strategy evaluation unit in interaction with the digital operational twin module for parallel simulation and deduction of candidate strategies in the digital operational twin module and output of the optimal response strategy according to the deduction results.
3. The network security operations management platform of claim 2, wherein: The evaluation indicators used by the strategy evaluation unit include risk reduction degree, business impact index, resource consumption cost and response efficiency.
4. The network security operations management platform of claim 1, wherein: The digital operational twin module synchronizes with real security tool chains, network topology asset data through API interfaces and simulates attack behaviors, defense actions and their interactive effects using simulation technology.
5. The network security operations management platform of claim 1, wherein: The feedback and evolutionary learning module quantitatively scores the effectiveness of the strategy execution by establishing a reward function that integrates risk reduction degree, operational efficiency improvement value, resource consumption cost and artificial feedback score in multiple dimensions.
6. The network security operations management platform of claim 1, wherein: Further comprise: A human-machine collaboration interface for presenting the strategy recommended by the intelligent decision hub and its simulation deduction results and evaluation basis to operational personnel and receiving the final decision instructions, modification suggestions or feedback scores of operational personnel.
7. The network security operations management platform of claim 6, wherein: The human-machine collaboration interface visually displays attack paths, strategy deduction processes and expected impacts.
8. The network security operations management platform of claim 1, wherein: The natural language interaction and operational knowledge generation module integrates a large language model to generate comprehensive reports containing root causes, disposal actions and effect evaluations based on instructions.
9. The network security operations management platform of claim 1, wherein: Further comprise: A unified data lake for storing and governing raw event data collected from the underlying security tool chain, strategy data generated by the platform, simulation data, execution result data and feedback data, providing data support for the intelligent decision hub, digital operational twin module and feedback and evolutionary learning module.
10. An operation method based on the network security operation management platform according to any one of claims 1-9, characterized in that: Comprise the following steps: Receiving security events and operational data; Generate and evaluate dynamic response strategies based on reinforcement learning models and digital operations twin simulations; Execute the preferred response strategy or the strategy confirmed by human; Collect execution feedback data for optimizing the reinforcement learning models; Receive queries or instructions through natural language interactions and generate operations knowledge reports.