Decision-making method based on network defense agent
By constructing a network system topology map and generating intelligent agents based on sigma rules, and deploying attack and defense agents for reinforcement learning, the problems of computational complexity and learning instability of agents in complex network environments are solved, and network threat processing with agent adaptability and strategy optimization is achieved.
Patent Information
- Application Number
- CN202510994756.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-18
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-07-18
AI Technical Summary
In a complex network environment, the computational complexity of single-agent reinforcement learning increases with the environmental state and the scale of the action space, making learning impractical. In multi-agent reinforcement learning, agents cannot utilize information from other agents, and the learning process is unstable and has poor convergence.
Build a network system topology diagram, use sigma rules to generate agents to match network threats based on threat intelligence, deploy attack and defense agents, train agents through reinforcement learning, use agent state encoders and evaluation networks to make action decisions, optimize agent strategies, and use the joint network environment status observed by other agents for evaluation.
It realizes that in the process of multi-agent reinforcement learning, the agents can adapt to changes in the environment and strategies, learn better network threat handling strategies, reduce computational complexity, and improve learning stability and convergence.
Smart Images

Figure CN120639464A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security technology, and in particular to a decision-making method based on a network defense intelligent agent. Background Art
[0002] As computer networks continue to evolve and become increasingly complex, the challenges facing cybersecurity are also intensifying. Traditional cyberthreat mitigation methods are typically static and passive, making them ineffective against dynamic and complex threats. Therefore, adaptive and proactive cyberthreat mitigation approaches show great potential. Reinforcement learning, as an adaptive solution, demonstrates great potential in addressing dynamic and complex cybersecurity challenges. Reinforcement learning can independently learn cyberthreat mitigation strategies through trial and error, adapting to evolving attack vectors. Its continuous learning capability enables it to adapt to new attack patterns over time, defending against unknown threats and overcoming the limitations of traditional defense methods. However, in single-agent reinforcement learning, a single agent observes the entire environment and decides actions based on the observed state, interacting with it and learning better cyberthreat mitigation strategies from this interaction. Single-agent reinforcement learning presents challenges in complex network environments: the scale of the environment state and action space in complex networks is large, and the computational complexity of single-agent reinforcement learning increases exponentially with the size of the environment state and action space, making single-agent reinforcement learning impractical in large, complex network environments.
[0003] If multi-agent reinforcement learning is employed, each agent observes a portion of the environment state and performs a portion of the action based on that portion of the environment state. This significantly reduces the size of the environment state and action space for each agent, significantly reducing the computational complexity of the agents on the network devices where they are deployed. However, both during training and executing an agent's policy, an agent's policy cannot utilize information about other agents. Training may be affected by the instability caused by the simultaneous training of all agents, and agents may be unable to distinguish whether changes in the environment state are due to random changes caused by the actions of other agents or to changes in the environment's transition function itself. In fact, as the policies of other agents change, each agent's perceived environment transition, observation, and reward function also change accordingly. These changes can lead to unstable learning and poor convergence. Summary of the Invention
[0004] In order to solve the above technical problems or at least partially solve the above technical problems and to address the above shortcomings, the present invention provides a decision-making method based on a network defense agent.
[0005] In a first aspect, the present invention provides a decision-making method based on a network defense agent, comprising: Build a network system topology diagram based on the connection relationship of network devices and construct node attributes; Utilize Sigma Rules to generate Sigma rules based on collected threat intelligence. Use the generated Sigma rules to match network threats that meet node attributes on nodes in the network system topology map. Automatically construct or supplement the attack action space of any node in the corresponding network system topology map based on the corresponding network threats. Multiple attacking and defending agents are constructed and deployed to corresponding nodes in the network system. Each agent is defined as an agent state encoder, a decision network, and an evaluation network. The decision network of each agent gives an action probability based on the local network environment state it observes and selects an action based on the action probability. The evaluation network of each agent evaluates the value of the strategy or action based on the joint network environment state observed by all attacking agents. The joint network environment state observed by the other agents other than the attacking agent is encoded by the corresponding agent state encoder. Reinforcement learning is used to train attack agents and defense agents, and the trained defense agents are used to make network defense decisions to deal with dynamic network attacks.
[0006] Furthermore, the nodes in the network system topology diagram are described by several attributes related to the process, file system, operating system, and hardware architecture; the attribute is represented by a pair of an identifier and a value, and the attribute identifier is used to represent: file path, operating system type used by the node, process ID, and command line used by the agent; the attribute value is used to represent file content, a complete description of the operating system, process running results, and command line output results.
[0007] Furthermore, for any node Nodei, the node attribute set Describes the node situation, where , is the total number of attributes of node Nodei; For a network system, the set of attribute sets of all nodes in the network system Describes the global network environment, which defines a global network environment state space, where , Indicates that node Nodei belongs to the network system.
[0008] Furthermore, the Sigma rules generated by the intelligent agent based on the collected threat intelligence include: crawling open source network threat intelligence web pages from open source network threat intelligence sources through a web crawling tool; guiding the multimodal language model through image analysis prompt words to convert the image-type open source network threat intelligence in the crawled web page elements related to the open source network threat intelligence into text-type; converting the textual content in the web page elements related to the open source network threat intelligence into a unified text format to obtain initial network threat intelligence; analyzing the keywords representing redundant content in the initial network threat intelligence title through a language model, and for the target title with keywords representing redundant content, excluding the repeated and redundant content in the initial network threat intelligence according to the position of the target title in the text structure hierarchy divided by all titles to obtain filtered network threat intelligence; providing the filtered network threat intelligence to at least one intelligent agent based on the language model, and the intelligent agent uses the semantic analysis of the language model to obtain the filtered network threat intelligence. The ability and multi-agent voting method are used to identify the first and second category entities from the filtered network threat intelligence and establish connections; the first category entities are the entities necessary to form the Sigma rule detection query part, and the first category entities include: API or process calls, API or process call request parameters, intrusion indicators, log sources and event sources; the second category entities provide contextual information of the network threat intelligence, and the second category entities include the title and description in the Sigma rule, threat techniques and tactics, false positives and threat levels; the language model used by the Sigma rule to create the prompt word control agent is based on the filtered network threat intelligence block, and the associated first and second category entities are extracted from the network threat intelligence block to create the Sigma rule; the Sigma rule is used to optimize the language model used by the prompt word control agent to optimize the generated Sigma rule; the Sigma rule is used to verify the language model used by the prompt word control agent to verify the generated and optimized Sigma rule.
[0009] Furthermore, the reinforcement learning training process includes: Pre-trained attack and defense agent state encoders; Initialize the model parameters of the constructed attacking and defending agents; For each iteration of reinforcement learning training, the execution includes: Each attack agent and defense agent observes the latest local network environment state from the network environment. The attack agent shares the observed local network environment state with the attack agent state encoder to generate a first encoding vector corresponding to each defense agent. The defense agent shares the observed local network environment state with the defense agent state encoder running on the high-performance device to generate a second encoding vector corresponding to each defense agent. The decision network of each attacking agent obtains the attack action probability distribution based on the local network environment state it observes and the first coding vector, and makes an attack action decision according to the attack action probability distribution; the attack actions of all attacking agents are obtained as the joint attack action of the attacking agent group; the decision network of each defending agent obtains the defense action probability distribution based on the local network environment state it observes and the second coding vector, and makes a defense action decision according to the defense action probability distribution; the defense actions of all defending agents are obtained as the joint defense action of the defending agent group; Apply joint attack and defense actions in the network system. The preset reward calculation strategy obtains the latest rewards for each attacking and defending agent based on the global network environment state and the joint attack and defense actions. After the joint attack and defense actions are applied to the network system, the attacking agent and the defending agent observe the local network environment states after executing the joint actions. Based on the local network environment state before and after executing the joint action, the attack strategy value loss or attack action value loss of the evaluation network of the attacking agent is calculated, and the defense strategy value loss or defense action value loss of the evaluation network of the defending agent is calculated; Calculate the strategic advantages of the attacking and defending agents, and calculate the decision loss of the decision network based on the strategic advantages; The iterative process adjusts the parameters of each agent with the goal of minimizing decision loss and value loss.
[0010] Furthermore, the attacking agent state encoder generates a first coding vector based on the local network environment state of other attacking agents except any target attacking agent, and provides it to the target attacking agent; the defending agent state encoder generates a second coding vector based on the local network environment state of other defending agents except any target defense agent, and provides it to the target defense agent.
[0011] Furthermore, the evaluation network of any agent performs strategy value evaluation based on the joint network environment state observed by all agents, or the evaluation network of any agent performs action value estimation based on the joint network environment state observed by all agents and the attack action decided by the attacking agent.
[0012] Furthermore, when any target agent uses action value evaluation, the actions taken by other agents are guaranteed to remain unchanged. Actions are randomly sampled from the action space of the target agent, and the expected value of all randomly sampled actions is calculated. The advantage is calculated by using the difference between the value of the current decision action and the expected value of the randomly sampled action. The marginal contribution of the agent's action to the overall reward is evaluated by using the difference between the actual action value and the expected value of the random action. Among them, the expected value of the randomly sampled action is obtained by summing the value of the randomly selected action generated by the decision network and the action generated by the probability-weighted evaluation network.
[0013] Furthermore, the value loss of the evaluation network is the square of the difference between the target value of the decision or action and the output of the corresponding type of evaluation network, where the target value of the decision or action includes the latest reward and the discount of the future evaluation network evaluation value.
[0014] Furthermore, for multiple agents, the latest reward of the defending agent in this application is the negative value of the reward of the attacking agent, all attacking agents share the reward, and all defending agents share the reward.
[0015] In the second aspect, the present invention provides a decision-making system based on a network defense intelligent agent, comprising: a plurality of interconnected network devices, any network device comprising: at least one processing unit, the processing unit being connected to a storage unit via a bus unit, the storage unit storing a computer program, the processing unit implementing the decision-making method based on the network defense intelligent agent by running the computer program stored in the storage unit.
[0016] In a third aspect, the present invention provides a computer-readable storage medium, which stores a computer program. When the computer program is executed, it implements the decision-making method based on the network defense agent.
[0017] The above technical solution provided by the embodiment of the present invention has the following advantages compared with the prior art: This application uses Sigma rules to generate Sigma rules based on collected threat intelligence. The generated Sigma rules are used to match network threats that meet node attributes at nodes in the network system topology map, and automatically construct or supplement the attack action space of any node in the corresponding network system topology map based on the corresponding network threats. It supports directly converting the latest threat intelligence into corresponding attack methods, and more conveniently introduces the latest attack methods into the attack action space. In this way, during the multi-agent reinforcement learning process, multiple attack agents constructed can learn to carry out network attacks with the latest network threats, and defense agents can learn defense strategies to deal with the latest network threats.
[0018] For any agent in this application, it always learns from the latest strategies of other agents. Learning from the latest strategies of other agents enables each agent to adapt to environmental changes or changes in the strategies of other agents. The agent state encoder of any agent extracts network environment information based on a relatively small joint network environment state and uses the extracted network environment information from a larger range for evaluation. This low computational cost enables evaluation based on more extensive and reliable information, resulting in more stable learning and guiding the decision network to learn better decisions. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0020] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0021] Figure 1 A flowchart of a decision-making method based on a network defense agent provided by an embodiment of the present invention; Figure 2 A schematic diagram of an exemplary network system topology diagram provided in an embodiment of the present invention; Figure 3 A schematic diagram of the range of various network environment states provided by an embodiment of the present invention; Figure 4 An architecture diagram of reinforcement learning provided by an embodiment of the present invention; Figure 5 A flowchart of reinforcement learning provided by an embodiment of the present invention; Figure 6 A schematic diagram of network devices in a decision-making system based on a network defense agent provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0023] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0024] Example 1 The present invention provides a decision-making method based on a network defense agent, comprising: S100, construct a network system topology diagram based on the connection relationship of network devices, such as Figure 2 As shown, the network system topology diagram includes the intranet's network entrance, firewall nodes, routing nodes, server nodes, and intranet user device nodes, as well as the extranet's external terminal nodes. The lines between the network system topology diagrams represent the connections between network devices; the extranet's external terminal nodes are used to identify network threats originating from the extranet.
[0025] The nodes in the network system topology are described by several attributes related to processes, file systems, operating systems, and hardware architectures. The format of any attribute related to processes, file systems, operating systems, and hardware architectures is an identifier and value pair. The j-th attribute of any node Nodei is expressed as: ,in, is the identifier of the jth attribute of node Nodei. The attribute identifier is used to represent: file path, operating system type used by the node, process ID, and command line used by the agent; is the value of the j-th attribute of node Nodei. Attribute values are used to represent: file content, a complete description of the operating system, process running results, and command line output results.
[0026] For node Nodei, the attribute set of node Nodei Describes the node situation, where , is the total number of attributes of node Nodei.
[0027] like Figure 3 As shown, for a network system, the set of attribute sets of all nodes in the network system is Describes the global network environment, which defines a global network environment state space, where , Indicates that node Nodei belongs to the network system. For complex network systems, the set of all node attributes is very large, and the global network environment state space based on the global network environment is also very large. The intelligent agents constructed later do not directly observe the entire global network environment state space. Instead, they observe the local network environment state, and the local network environment states observed by all intelligent agents form the joint network environment state.
[0028] S200, using sigma rules to generate sigma rules based on the collected threat intelligence, using the generated sigma rules to match network threats that meet the node attributes on the nodes of the network system topology map, and automatically constructing or supplementing the attack action space of any node in the corresponding network system topology map according to the corresponding network threats.
[0029] To further clarify the objectives, technical solutions, and advantages of the embodiments of the present invention, the following describes Sigma rules. Sigma rules are a universal signature format and a set of security event detection rules used to analyze and identify abnormal network behavior. The structure of a Sigma rule can be divided into three main parts: a header, options, and a detection query. The header contains basic information about the Sigma rule, such as the rule ID, title, description, author, and date. This information is crucial for understanding the context, purpose, and origin of the rule. The options define the rule's contextual requirements, such as process creation time and the use of a specific process ID. The detection query is the core of the rule definition, describing the specific conditions to be detected, such as fields and their values in the log source. It leverages fields and values from various log sources for precise matching. Log sources can be operating systems, applications, or any other log generator. Fields, combined with corresponding values, define the conditions for detecting security threats. Sigma rules provide a range of condition combinations and logical operators to create richer rule expressions. Condition combinations connect different rule selection components using the logical operators "and," "or," and "not" to construct complex detection scenarios. In addition, rule sets can be structured into hierarchies and dependencies, allowing related rules to be organized together to form a hierarchical structure. This structure helps manage a large number of rules and improves readability and maintainability. The relationships between rule sets can be inclusion and dependency.
[0030] The Sigma rules for generating an intelligent agent based on the collected threat intelligence include: crawling open source network threat intelligence web pages from open source network threat intelligence sources through a web crawling tool; guiding the multimodal language model through image analysis prompt words to convert the image-type open source network threat intelligence in the crawled web page elements related to the open source network threat intelligence into text-type; converting the textual content in the web page elements related to the open source network threat intelligence into a unified text format to obtain initial network threat intelligence; analyzing the keywords representing redundant content in the initial network threat intelligence title through a language model, and for the target title with keywords representing redundant content, excluding the repeated and redundant content in the initial network threat intelligence according to the position of the target title in the text structure hierarchy divided by all titles to obtain filtered network threat intelligence; providing the filtered network threat intelligence to at least one intelligent agent based on the language model, and the intelligent agent uses the semantic analysis ability of the language model and A multi-agent voting method identifies first-category entities and second-category entities from filtered cyber threat intelligence and establishes connections. First-category entities are necessary for forming the query portion of a sigma rule detection, including: API or process calls, API or process call request parameters, intrusion indicators, log sources, and event sources. Second-category entities provide contextual information for cyber threat intelligence, including titles and descriptions, threat techniques and tactics, false positives, and threat levels. A language model for creating prompt words in a sigma rule-based control agent is used to create sigma rules based on the associated first-category entities and second-category entities extracted from the filtered cyber threat intelligence block. The generated sigma rules are optimized using the language model used by the sigma rule-based control agent. The generated and optimized sigma rules are verified using the language model used by the sigma rule-based control agent. The aforementioned method for generating sigma rules has been filed under patent application number CN202510648405.4.
[0031] When a cyber threat process indicated by any Sigma rule exists within a device system, the device is likely vulnerable to this cyber threat. The attack vectors that represent this cyber threat are then extracted and added to the attack action space of the device's corresponding node. This process allows the latest threat intelligence to be directly converted into corresponding attack vectors, making it easier to incorporate these new attack vectors into the attack action space. This allows multiple attack agents to learn to attack using the latest cyber threats, and the defense agents to learn defense strategies to handle these threats.
[0032] S300: Construct n attack agents and m defense agents required for multi-agent reinforcement learning, deploy them to the corresponding nodes of the network system, and use reinforcement learning to conduct attack and defense confrontation training. Figure 4As shown, any agent is defined as an agent state encoder, a decision network and an evaluation network; the decision network of any agent gives an action probability according to the local network environment state it observes, and selects an action according to the action probability; the evaluation network of any agent evaluates the joint network environment state observed by all attacking agents to obtain the value of the strategy or execution action; wherein, the joint network environment state observed by other agents other than the agent is encoded by the corresponding agent state encoder. This application utilizes the introduced agent state encoder to integrate a smaller joint network environment state than the global network environment state space, and compresses and encodes the joint network environment state and provides it to the evaluation network as an evaluation basis. While ensuring that the agent state encoder can operate, the largest possible range of environmental state perception is used to evaluate the impact of decisions on the network environment, reduce the instability caused by other agent decisions, and achieve effective generalization of the agent.
[0033] For any attacking agent, the attacking agent observes the properties of the nodes it affects to obtain the local network environment state, selects attack actions from the attack action space of the nodes it affects based on the observed local network environment state, and executes them by the corresponding nodes. The attack actions executed by the attacking agent-controlled nodes will modify the properties of one or more nodes, thereby changing the global network environment state of the network system and the local network environment state that the attacking agent can observe; after the attacking agent-controlled node executes the action, it approaches or moves away from the attack target.
[0034] For any defensive agent, the defensive agent observes the attributes of the nodes it affects to obtain the local network environment status. The defensive agent selects a defensive action from the defensive action space of the nodes it affects based on the local network environment status it observes, and the action is executed by the corresponding node. The defensive action executed by the defensive agent control node will modify the attributes of one or more nodes, thereby changing the global network environment status of the network system and the local network environment status that the defensive agent can observe. After each defensive agent control node executes a defensive action, it approaches or moves away to prevent network attacks.
[0035] The local network environment that any agent can observe and the action space it can perform are determined by the properties of the nodes it affects. For example, reading a specific file or remapping a port may require a certain permission level. If an attacking agent obtains the required permission level in a node, it can read specific files to obtain more local network environment status or perform port remapping actions, further expanding the scope of the network attack.
[0036] In this application, any intelligent agent is defined as an agent state encoder, a decision network, and an evaluation network. During the training phase, the evaluation network is trained with the decision network through reinforcement learning, enabling the decision network to learn the decision-making ability to deal with network threats. Once training is complete, the evaluation network is no longer used. During the application phase of the intelligent agent, only the decision network is responsible for generating the agent's actions.
[0037] The critic network of any agent has more information than the decision network for value estimation, which allows it to more accurately estimate the rewards of the decision network's strategy, bringing benefits. In addition, by utilizing the local network environment state information observed by all other agents, the critic network can more quickly adapt to the non-stationary strategies of other agents.
[0038] like Figure 5 As shown in Figure 2, the entire reinforcement learning training process includes: Pre-train the attack and defense agent state encoders. The agent state encoder uses a VAE encoder, which maps high-dimensional attributes representing the network environment state to a low-dimensional latent space. The VAE decoder then reconstructs the original attributes from the latent vector. Training uses a reconstruction loss and a KL loss between the VAE encoder output and the attribute prior distribution.
[0039] Initialize the model parameters of the attacking and defending agents: In the specific implementation process, the decision network of any attacking agent a_agentj is expressed as: ,in, are the model parameters of the decision network of the attacking agent a_agentj, is the local network environment state observed by the attacking agent a_agentj at time t. The decision network of the attacking agent a_agentj gives the attack action probability distribution based on the local network environment state observed by the attacking agent a_agentj, and selects the attack action to be executed according to the attack action probability distribution. For example, the action probability is expressed as: The probabilities of different attack actions form the attack action probability distribution. The decision network of the attacking agent a_agentj models the conditional probabilities between state and action based only on the local network environment state observed by the attacking agent, rather than the global network environment state. This ensures that the decision network of the attacking agent can be executed decentralized during the application process.
[0040] The evaluation network of an example attack agent a_agentj that evaluates the value of an attack strategy is expressed as: ,in, is the model parameter of the evaluation network of the attacking agent agentj, is the corresponding first encoding vector obtained based on the joint network environment state observed by other attacking agents except the attacking agent a_agentj: , the first encoding vector is passed through the attack agent state encoder The encoding is a compressed representation of the joint network environment state. are the model parameters of the attack agent state encoder; an example uses a VAE encoder; therefore, the evaluation network of the attack agent a_agentj evaluates the attack strategy value made by the decision network of the attack agent a_agentj based on the joint network environment state observed by all attack agents.
[0041] The evaluation network of an example attack agent a_agentj that evaluates the value of an attack action is expressed as: , The attack action selected for the attacking agent a_agentj, The model parameters of the evaluation network for evaluating the value of attack actions; when evaluating the value, the attack action value is estimated not only based on the local and joint network environment states, but also based on the attack actions decided by the attacking agent.
[0042] The decision network of the defense agent d_agentk is expressed as: ,in, are the parameters of the decision network of the defense agent d_agentk, The decision network of the defense agent d_agentk gives the attack action to be executed based on the local network environment state observed by the defense agent d_agentk at time t The conditional probability is expressed as: The decision network of the defensive agent d_agentk models the conditional probabilities between state and action based only on the local network environment state observed by the defensive agent, rather than the global network environment state. This ensures that the decision network of the defensive agent can be executed decentralized during the application process.
[0043] The evaluation network of an example defense agent d_agentk for evaluating the value of a defense strategy is represented as: ,in, is the model parameter of the evaluation network of the defense agent agentk, is the second encoding vector obtained based on the joint network environment state observed by other defense agents except the defense agent d_agentk: , the second encoding vector is passed through the defense agent state encoder The encoding is a compressed representation of the joint network environment state. are the model parameters of the defense agent state encoder; an example uses a VAE encoder; therefore, the evaluation network of the defense agent d_agentk evaluates the value of the defense agent d_agentk strategy based on the joint network environment state observed by all attacking agents.
[0044] The evaluation network of an example defensive agent d_agentk that evaluates the value of defensive actions is represented as: ,in, is the model parameter of the evaluation network of the defense agent agentk, The value of the currently selected defense action for each defending agent d_agentk is evaluated based not only on the local and joint network environment states, but also on the defense action chosen by the defending agent. The method for calculating the value loss is the same.
[0045] After initializing the model parameters, the reinforcement learning iterative training is performed as follows. For each reinforcement learning training iteration, the execution includes: Each attack agent and defense agent observes the latest local network environment status from the environment. The attack agent shares the observed local network environment status with the attack agent state encoder running on the high-performance device, and the defense agent shares the observed local network environment status with the defense agent state encoder running on the high-performance device.
[0046] The attacking agent state encoder provides a first coding vector to the target attacking agent based on the generation of the local network environment state of other attacking agents except any target attacking agent, and generates a corresponding first coding vector for each attacking agent; the defending agent state encoder provides a second coding vector to the target defense agent based on the generation of the local network environment state of other defense agents except any target defense agent, and generates a corresponding second coding vector for each defense agent.
[0047] The decision network of each attacking agent obtains the attack action probability distribution according to its local network environment state and the first coding vector, and makes attack action decisions according to the attack action probability distribution; the attack actions of all attacking agents are obtained as the joint attack action of the attacking agent group; the decision network of each defending agent obtains the defense action probability distribution according to its local network environment state and the second coding vector, and makes attack action decisions according to the attack action probability distribution; the defense actions of all defending agents are obtained as the joint defense action of the defense agent group.
[0048] In a network system, joint attack and defense actions are applied. A preset reward calculation strategy uses the global network environment state and the combined attack and defense actions to obtain the latest rewards for each attacking and defending agent. For multiple agents, the latest reward for the defending agent in this application is the negative of the attacking agent's reward. That is, the greater the attacking agent's reward, the smaller the defending agent's reward. This game is achieved by controlling the attacking and defending agents using rewards. In this application, all attacking agents share the reward, and all defending agents share the reward.
[0049] After the joint attack action and the joint defense action act on the network system, the attacking agent and the defending agent observe the local network environment status after their respective joint actions are executed.
[0050] Based on the local network environment state before and after executing the joint action, the attack strategy value loss or attack action value loss of the evaluation network of the attacking agent is calculated, and the defense strategy value loss or defense action value loss of the evaluation network of the defending agent is calculated. In the specific implementation process: the value loss of the evaluation network that evaluates the value of the attack strategy is: ; in, is the target value of the attack agent's decision in the network system environment, For non-final decisions: ; in, is the latest reward of the current attacking agent a_agentj, is the future (corresponding to the next observed local network environment state , the next first encoding vector ) The discount of the decision value evaluated by the evaluation network of the attacking agent a_agentj, is the discount factor.
[0051] For the final decision: There will be no next action and corresponding state after the final decision.
[0052] The value loss of the evaluation network that evaluates the value of the attack action is: ; in, is the target value of the attack agent's action in the network system environment, For non-final decisions: ; in, is the latest reward of the current attacking agent a_agentj, is the future (corresponding to the next observed local network environment state , the next first encoding vector and the next action ) The discount of the attack agent a_agentj’s evaluation network’s action value, is the discount factor.
[0053] For the final decision: There will be no next action and corresponding state after the final decision.
[0054] The two calculation methods for any defensive agent are the same and will not be repeated here.
[0055] Calculate the advantages of the attacking agent and the defending agent, based on the advantages and calculate the decision loss of the decision network.
[0056] When using decision value evaluation, the strategic advantage of the attacking agent for the final decision is: ; When using decision value evaluation, for non-final decisions, the strategic advantage of the attacking agent is: .
[0057] When using policy value evaluation, the policy advantage of the defending agent for the final decision is: ; When using decision value evaluation, for non-final decisions, the strategic advantage of the defending agent is: .
[0058] When any target agent uses action value evaluation, the actions taken by other agents are guaranteed to remain unchanged. Actions are randomly sampled from the target agent's action space, and the expected value of all randomly sampled actions is calculated. The advantage is calculated by the difference between the value of the current decision action and the expected value of the randomly sampled action. The marginal contribution of the agent's action to the overall reward is evaluated by the difference between the actual action value and the expected value of the random action: The attack agent advantage is obtained based on the action value output by the attack agent evaluation network according to the following formula: ; in, Represents the attacking agent a_agentj in the observed local network environment and the first encoding vector Next select action advantage value.
[0059] When the actions performed by other attacking agents remain unchanged, the attacking agent a_agentj performs the action The action-value function of . It limits the actions performed by other attacking agents to remain unchanged.
[0060] The expected value of the action randomly selected by the attacking agent a_agentj according to the current strategy, while the actions of other attacking agents remain unchanged, is represented. By comparing the value of the current action with the expected value of the random action, the difference between the actual action value and the expected value of the random action is used to evaluate the marginal contribution of the attacking agent a_agentj's action to the overall reward. A positive difference indicates that the current action is better than the random action; a negative difference indicates that the current action is worse.
[0061] When using action value evaluation, the action value output by the defense agent evaluation network is used to obtain the advantage of the defense agent according to the following formula: ; in, Represents the defense agent in the observed local network environment and the second encoding vector Next select action advantage value.
[0062] When the actions performed by other defense agents remain unchanged, the defense agent d_agentk performs random actions The action-value function of . It limits the actions performed by other defensive agents to remain unchanged.
[0063] , represents the expected value of randomly selecting an action according to the current strategy for the defending agent d_agentk, while the actions of other defending agents remain unchanged. By comparing the value of the current action with the expected value of the random action, we evaluate the marginal contribution of the defending agent d_agentk's action to the overall reward. If the difference is positive, the current defense action is better than the random choice; otherwise, it is worse.
[0064] The decision loss of the decision network is calculated based on the policy advantage as follows: For the attacking agent, when using strategy evaluation, its decision loss is: ; in, Indicates local observation Select an attack action The selected attack action is good to make When multiplied by After , the decision loss will tend to increase the probability of the attack action being selected. The difference in the selected attack action makes When multiplied by After that, the decision loss will tend to reduce the probability of the attack action being selected. Guided by the decision loss, the decision network of the attacking agent will gradually strengthen good attack actions and weaken bad attack actions, and eventually learn a better attack decision strategy.
[0065] For the defensive agent, when using strategy evaluation, its decision loss is: ; in, Indicates local observation Strategy selection action The selected defensive action is good to make When multiplied by After , the decision loss will tend to increase the probability of the action being selected. The selected defensive action is poor so that When multiplied by After that, the decision loss will tend to reduce the probability of the defensive action being selected. Guided by the decision loss, the decision network of the defensive agent will gradually strengthen good defensive actions and weaken bad defensive actions, and ultimately learn a better defensive decision strategy.
[0066] The parameters of each agent are adjusted with the goal of minimizing decision loss and value loss.
[0067] In this application, the attacking agent and the defending agent simulate a more complex interaction between the attacker and the defender, and the probability of reaching a new state of the network system (for example, being compromised or secure) depends on the joint actions of both parties. For example, the attacker chooses a specific attack strategy, and the defender chooses a defensive response (for example, blocking an IP address). For any agent in this application, it always learns from the latest strategies of other agents. Learning from the latest strategies of other agents enables each agent to adapt to environmental changes or changes in the strategies of other agents. The agent state encoder of any agent extracts network environment information based on a relatively small joint network environment state, and uses the extracted network environment information in a larger range for evaluation. The evaluation is based on more extensive and reliable information at a low computational cost, resulting in more stable learning to guide the decision network to learn better decisions.
[0068] Example 2 The embodiment of the present invention provides a decision-making system based on a network defense agent, comprising a plurality of interconnected network devices, referring to Figure 6As shown, any network device includes: at least one processing unit, the processing unit connected to a storage unit via a bus unit. The storage unit, as a computer-readable storage medium, can be used to store software programs, computer executable programs, and modules, such as the software programs, computer executable programs, and modules corresponding to the decision-making method based on a network defense agent in an embodiment of the present invention. The processing unit implements the decision-making method based on a network defense agent by running the software programs, computer executable programs, and modules stored in the storage unit, including: Build a network system topology diagram based on the connection relationship of network devices and construct node attributes; Utilize Sigma Rules to generate Sigma rules based on collected threat intelligence. Use the generated Sigma rules to match network threats that meet node attributes on nodes in the network system topology map. Automatically construct or supplement the attack action space of any node in the corresponding network system topology map based on the corresponding network threats. Multiple attacking and defending agents are constructed and deployed to corresponding nodes in the network system. Each agent is defined as an agent state encoder, a decision network, and an evaluation network. The decision network of each agent gives an action probability based on the local network environment state it observes and selects an action based on the action probability. The evaluation network of each agent evaluates the value of the strategy or action based on the joint network environment state observed by all attacking agents. The joint network environment state observed by the other agents other than the attacking agent is encoded by the corresponding agent state encoder. Reinforcement learning is used to train attack agents and defense agents, and the trained defense agents are used to make network defense decisions to deal with dynamic network attacks.
[0069] Of course, the computer program stored in the storage unit of the decision-making system based on the network defense agent provided by an embodiment of the present invention is not limited to the method operations described above, and can also execute related operations in the decision-making method based on the network defense agent provided by any embodiment of the present invention.
[0070] Example 3 An embodiment of the present invention provides a computer-readable storage medium storing a computer program. When the computer program is executed, the decision-making method based on the network defense agent is implemented, including: Construct a network system topology diagram based on the connection relationship of network devices; Utilize Sigma Rules to generate Sigma rules based on collected threat intelligence. Use the generated Sigma rules to match network threats that meet node attributes on nodes in the network system topology map. Automatically construct or supplement the attack action space of any node in the corresponding network system topology map based on the corresponding network threats. Multiple attacking and defending agents are constructed and deployed to corresponding nodes in the network system. Each agent is defined as an agent state encoder, a decision network, and an evaluation network. The decision network of each agent gives an action probability based on the local network environment state it observes and selects an action based on the action probability. The evaluation network of each agent evaluates the value of the strategy or action based on the joint network environment state observed by all attacking agents. The joint network environment state observed by the other agents other than the attacking agent is encoded by the corresponding agent state encoder. Reinforcement learning is used to train attack agents and defense agents, and the trained defense agents are used to make network defense decisions to deal with dynamic network attacks.
[0071] An embodiment of the present invention provides a computer-readable storage medium, in which the computer program stored is not limited to the method operations described above, but can also execute related operations in a decision-making method based on a network defense agent provided by any embodiment of the present invention.
[0072] In the embodiments provided by the present invention, it should be understood that the disclosed structures and methods can be implemented in other ways. For example, the structural embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, structure or unit, which can be electrical, mechanical or other forms.
[0073] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0074] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0075] The foregoing description is intended only to provide specific embodiments of the present invention, which will enable those skilled in the art to understand and implement the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not intended to be limited to the embodiments shown herein, but is intended to be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A decision-making method based on a network defense agent, characterized in that: include: Build a network system topology diagram based on the connection relationship of network devices and construct node attributes; Utilize Sigma Rules to generate Sigma rules based on collected threat intelligence. Use the generated Sigma rules to match network threats that meet node attributes on nodes in the network system topology map. Automatically construct or supplement the attack action space of any node in the corresponding network system topology map based on the corresponding network threats. Multiple attacking and defending agents are constructed and deployed to corresponding nodes in the network system. Each agent is defined as an agent state encoder, a decision network, and an evaluation network. The decision network of each agent gives an action probability based on the local network environment state it observes and selects an action based on the action probability. The evaluation network of each agent evaluates the value of the strategy or action based on the joint network environment state observed by all attacking agents. The joint network environment state observed by the other agents other than the attacking agent is encoded by the corresponding agent state encoder. Reinforcement learning is used to train attack agents and defense agents, and the trained defense agents are used to make network defense decisions to deal with dynamic network attacks.
2. The decision-making method based on network defense agent according to claim 1, characterized in that: The nodes in the network system topology diagram are described by several attributes related to the process, file system, operating system, and hardware architecture; an attribute is represented by a pair of an identifier and a value. The attribute identifier is used to represent: file path, operating system type used by the node, process ID, and command line used by the agent; the attribute value is used to represent file content, a complete description of the operating system, process running results, and command line output results.
3. The decision-making method based on network defense agent according to claim 2, characterized in that: For any node Nodei, the node attribute set D Nodei Describes the node situation, where J Nodei is the total number of attributes of node Nodei; For a network system, the set D of attribute sets of all nodes in the network system describes the global network environment. The global network environment defines a global network environment state space, where D = {D Nodei |Nodei∈Net}, Nodei∈Net indicates that node Nodei belongs to the network system.
4. The decision-making method based on network defense agent according to claim 1, characterized in that: The Sigma rules for generating an intelligent agent based on the collected threat intelligence include: crawling open source network threat intelligence web pages from open source network threat intelligence sources through a web crawling tool; guiding the multimodal language model through image analysis prompt words to convert the image-type open source network threat intelligence in the crawled web page elements related to the open source network threat intelligence into text-type; converting the textual content in the web page elements related to the open source network threat intelligence into a unified text format to obtain initial network threat intelligence; analyzing the keywords representing redundant content in the initial network threat intelligence title through a language model, and for the target title with keywords representing redundant content, excluding the repeated and redundant content in the initial network threat intelligence according to the position of the target title in the text structure hierarchy divided by all titles to obtain filtered network threat intelligence; providing the filtered network threat intelligence to at least one intelligent agent based on the language model, and the intelligent agent uses the semantic analysis ability of the language model and A multi-agent voting method is used to identify first-category entities and second-category entities from filtered network threat intelligence and establish connections; first-category entities are entities necessary to form the query part of sigma rule detection, and first-category entities include: API or process calls, request parameters of API or process calls, intrusion indicators, log sources and event sources; second-category entities provide contextual information of network threat intelligence, and second-category entities include titles and descriptions in sigma rules, threat techniques and tactics, false positives and threat levels; sigma rules are used to create a language model used by prompt word control agents based on the filtered network threat intelligence block, and sigma rules are created based on the associated first-category entities and second-category entities extracted from the network threat intelligence block; sigma rules are used to optimize the language model used by the prompt word control agent to optimize the generated sigma rules; sigma rules are used to verify the language model used by the prompt word control agent to verify the generated and optimized sigma rules.
5. The decision-making method based on network defense agent according to claim 1, characterized in that: The reinforcement learning training process includes: Pre-trained attack and defense agent state encoders; Initialize the model parameters of the constructed attacking and defending agents; For each iteration of reinforcement learning training, the execution includes: Each attack agent and defense agent observes the latest local network environment state from the network environment. The attack agent shares the observed local network environment state with the attack agent state encoder to generate a first encoding vector corresponding to each defense agent. The defense agent shares the observed local network environment state with the defense agent state encoder running on the high-performance device to generate a second encoding vector corresponding to each defense agent. The decision network of each attacking agent obtains the attack action probability distribution based on the local network environment state it observes and the first coding vector, and makes an attack action decision according to the attack action probability distribution; the attack actions of all attacking agents are obtained as the joint attack action of the attacking agent group; the decision network of each defending agent obtains the defense action probability distribution based on the local network environment state it observes and the second coding vector, and makes a defense action decision according to the defense action probability distribution; the defense actions of all defending agents are obtained as the joint defense action of the defending agent group; Apply joint attack and defense actions in the network system. The preset reward calculation strategy obtains the latest rewards for each attacking and defending agent based on the global network environment state and the joint attack and defense actions. After the joint attack and defense actions are applied to the network system, the attacking agent and the defending agent observe the local network environment states after executing the joint actions. Based on the local network environment state before and after executing the joint action, the attack strategy value loss or attack action value loss of the evaluation network of the attacking agent is calculated, and the defense strategy value loss or defense action value loss of the evaluation network of the defending agent is calculated; Calculate the strategic advantages of the attacking and defending agents, and calculate the decision loss of the decision network based on the strategic advantages; The iterative process adjusts the parameters of each agent with the goal of minimizing decision loss and value loss.
6. The decision-making method based on network defense agent according to claim 5, characterized in that: The attacking agent state encoder generates a first coding vector based on the local network environment state of other attacking agents except any target attacking agent, and provides it to the target attacking agent; the defending agent state encoder generates a second coding vector based on the local network environment state of other defending agents except any target defense agent, and provides it to the target defense agent.
7. The decision-making method based on network defense agent according to claim 5, characterized in that: The evaluation network of any agent performs strategy value evaluation based on the joint network environment state observed by all agents, or the evaluation network of any agent performs action value estimation based on the joint network environment state observed by all agents and the attack action decided by the attacking agent.
8. The decision-making method based on network defense agent according to claim 7, characterized in that: When any target agent uses action value evaluation, the actions taken by other agents are guaranteed to remain unchanged. Actions are randomly sampled from the action space of the target agent, and the expected value of all randomly sampled actions is calculated. The advantage is calculated by using the difference between the value of the current decision action and the expected value of the randomly sampled action. The difference between the actual action value and the expected value of the random action is used to evaluate the marginal contribution of the agent's action to the overall reward. Among them, the expected value of the randomly sampled action is obtained by summing the value of the randomly selected action generated by the decision network and the action generated by the probability-weighted evaluation network.
9. The decision-making method based on network defense agent according to claim 5, characterized in that: The value loss of the evaluation network is the square of the difference between the target value of the decision or action and the output of the corresponding type of evaluation network, where the target value of the decision or action includes the latest reward and the discount of the future evaluation network evaluation value.
10. The decision-making method based on network defense agent according to claim 5, characterized in that: For multiple agents, the latest reward of the defending agent in this application is the negative value of the reward of the attacking agent. All attacking agents share the reward, and all defending agents share the reward.
Citation Information
Patent Citations
A method for generating network threat rules based on threat intelligence
CN120185930B
Cloud boundary network active decision defense method based on deep reinforcement learning
CN116599704A
ATTCK knowledge graph-based attack chain generation method and device
CN117978476A
Multi-agent system
WO2025078127A1