Automatic network defense method and system fusing threat intelligence
By integrating threat intelligence and using hierarchical reinforcement learning algorithms to train defense intelligence, the problem of untimely update of defense strategies in the existing technology is solved, real-time and dynamic nature of automated network defense is achieved, and new types of cyber attacks can be effectively defended against.
Patent Information
- Application Number
- CN202510131606.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-05-16
AI Technical Summary
It is difficult for the existing technology to update the defense strategy dynamically based on the iterative cyber attacks without human participation, resulting in a lack of real-time and dynamic nature of the defense strategy, making it difficult to effectively defend against new cyber attacks.
By integrating threat intelligence, measuring the relevance of threat intelligence, screening threat intelligence, and generating defense actions based on the action guidelines and observation indicators in the screened threat intelligence, expanding the defense action space of the defense agent, training the defense agent and generating defense strategies. Design a layered reinforcement learning algorithm, deploy the trained defense intelligence into the network system, conduct network offensive and defense tests, and realize automated network defense.
It realizes that without human participation, the defense strategy is updated dynamically according to the system environment, mitigate network attacks in real time, improve the real-time and dynamic nature of automated defense, and can effectively defend against new types of cyberattacks.
Smart Images

Figure CN120017343A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of automated network defense, and in particular to an automated network defense method and system integrating threat intelligence. Background Art
[0002] In recent years, cyber attacks have shown an increasing trend. With the rapid development of science and technology, cyber attackers continue to develop new methods and tools to invade systems, steal data and destroy infrastructure, posing serious risks to individuals, businesses and governments. According to the Cybersecurity Technology Report, during 2023, attackers will be discovered on average 19 days after launching attacks in a network environment. During this period, attackers have enough time to penetrate deeply, launch attacks, and perform other malicious activities. However, network defense is a complex task, and currently this task is usually still completed by security technicians. Manual participation increases operating costs and response time, putting network systems at risk.
[0003] As large enterprise network systems become more and more automated, security personnel can respond to network threats more quickly and accurately using network security orchestration technology. However, these security tools can only respond to known attacks, such as isolating known malware and blocking known malicious traffic. Researchers consider using artificial intelligence technology for network defense. Machine learning technology has been used to protect systems and reduce the damage caused by attacks. Many supervised and unsupervised technologies have been used to develop defensive network applications, such as attack classification and intrusion detection systems. However, traditional technologies cannot generate effective defense strategies when faced with unrecognized attacks.
[0004] Deep reinforcement learning technology can train intelligent agents to identify whether there are attackers in the current scene and generate active defense strategies to minimize the losses caused by attacks and achieve automated defense. However, when deep reinforcement learning technology is currently used to achieve automated defense, the defense actions used by the defense agent are set in advance by security personnel. Human participation makes it difficult to update and adjust the defense actions in a timely manner. The generated defense strategies lack pertinence and real-time performance, which limits the dynamic and scalability of automated defense. When new threats emerge, the generated defense strategies cannot protect system security well.
[0005] Therefore, how to dynamically update defense strategies based on the ever-increasing network attacks without human intervention is an urgent problem that needs to be solved. Summary of the invention
[0006] The present invention aims to solve the problem in the prior art of how to dynamically update the defense strategy according to the continuously iterative network attacks without human intervention.
[0007] To solve the above technical problems, the present invention is achieved through the following technical solutions:
[0008] Solution 1: The present invention proposes an automated network defense method integrating threat intelligence, the method comprising:
[0009] Step 1: Measure the relevance of threat intelligence and filter threat intelligence;
[0010] Step 2: Generate defense actions based on the action guidelines and observation indicators in the threat intelligence screened in step 1, expand the defense action space of the defense agent, train the defense agent and generate defense strategies;
[0011] Step 3: Based on the defense strategy generated in step 2, a hierarchical reinforcement learning algorithm is designed, the trained hierarchical reinforcement learning defense agent is deployed into the network system, and network attack and defense tests are performed to complete the construction of an automated network defense method that integrates threat intelligence.
[0012] Further, a preferred implementation is provided, in which the method for screening threat intelligence relevance metrics in step 1 is:
[0013] Step 101: Collect threat intelligence data in a structured threat information expression format, and collect network and asset information in a network scenario;
[0014] Step 102, extract all action plan statements in the STIX threat intelligence and convert them into vectors; extract network and asset information from the network system and convert them into vectors;
[0015] Step 103: Calculate the cosine similarity between the action policy vector and the asset information vector, use the cosine similarity index to sort the threat intelligence, and screen out the threat intelligence that is highly relevant to the network scenario.
[0016] Further, a preferred implementation is provided, in which the method for generating a defense strategy in step 2 is:
[0017] Step 201, generating a general defense action according to an open command and control instruction, for commanding and controlling a network defense component;
[0018] Step 202: dynamically generate CTI defense actions based on STIX format threat intelligence. The generation of CTI defense actions uses the action guidelines in STIX threat intelligence. The action guidelines include specific measures for preventing, detecting, and responding to network attacks. The action guidelines and system information are extracted to obtain operations, targets, and parameters. Operations include shutdown, update, and installation. The target is the host IP address. The parameters include any parameters required to execute defense commands, such as IP addresses, ports, and patches.
[0019] Step 203: Construct a defense action space by combining the general defense action and the CTI defense action.
[0020] Furthermore, a preferred implementation scheme is provided, wherein the actions in the general defense action space include general defense actions and CTI defense actions, the general defense actions are generated according to OpenC2 defense commands, including five types: monitoring, analysis, induction, deletion and recovery, and the CTI defense actions are generated according to STIX threat intelligence.
[0021] Furthermore, a preferred implementation is provided, in which a hierarchical reinforcement learning algorithm is designed based on the defense strategy generated in step 3, and a method for training a defense agent to generate a defense strategy is implemented by adopting a hierarchical proximal strategy optimization algorithm.
[0022] Furthermore, a preferred implementation is provided, wherein the defense agent includes a control agent, two defense agents, and a CTI defense action is added to the agent defense action space.
[0023] Furthermore, a preferred implementation is provided, wherein the threat intelligence relevance metrics screened in step 1 include IP addresses, action guidelines, and network topology information.
[0024] Solution 2: An automated network defense system integrating threat intelligence, the system comprising:
[0025] A filtering module for filtering threat intelligence relevance metrics;
[0026] Defense action generation module, where users generate defense strategies based on the action guidelines and observation indicators in the threat intelligence relevance metrics selected by the screening module, train defense agents and generate defense strategies;
[0027] The automated defense agent design and training module is used to design a hierarchical reinforcement learning algorithm based on the defense strategy generated by the defense action generation module, deploy the trained hierarchical reinforcement learning defense agent into the network system, and conduct network attack and defense tests to complete the construction of an automated network defense method that integrates threat intelligence.
[0028] Solution three: A computer device includes a memory and a processor, wherein the memory stores a computer program, and when the processor runs the computer program stored in the memory, the processor executes any one of the methods described in Solution one.
[0029] Solution 4: A computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of the method described in any one of Solution 1.
[0030] The present invention is beneficial in that:
[0031] The method described in the present invention aims at automated network defense, proposes defense action generation based on STIX threat intelligence, designs a hierarchical PPO reinforcement learning algorithm, and allows the intelligent agent to learn automated defense network attacks. The automated network defense method that integrates threat intelligence constructed by the present invention can generate active defense strategies according to system environment conditions without human intervention, mitigate network attacks in real time, and make automated defense real-time and dynamic.
[0032] Compared with the existing methods, the present invention has better real-time defense effect against network attacks, can defend against new network attacks, and realize automatic network defense that can dynamically update defense strategies.
[0033] The present invention is also applicable to the field of defense strategies generated by artificial intelligence technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 This is a flow chart of an automated network defense method integrating threat intelligence as described in Implementation Method 1.
[0035] Figure 2 This is a schematic diagram of the overall framework of an automated network defense method that integrates threat intelligence as described in Implementation Method 1.
[0036] Figure 3 This is a typical network attack and defense network topology diagram described in Implementation Mode 11.
[0037] Figure 4 This is a schematic diagram of the box plot of the attack reward value distribution of different automated defense methods described in Implementation Method 11.
[0038] Figure 4 Among them, (a) is a box plot of the average reward value distribution when the different automated defense methods defend against network attacks with known attack paths for 30 rounds, (b) is a box plot of the average reward value distribution when the different automated defense methods defend against network attacks with known attack paths for 50 rounds, (c) is a box plot of the average reward value distribution when the different automated defense methods defend against network attacks with known attack paths for 100 rounds, (d) is a box plot of the average reward value distribution when the different automated defense methods defend against network attacks with unknown attack paths for 30 rounds, (e) is a box plot of the average reward value distribution when the different automated defense methods defend against network attacks with unknown attack paths for 50 rounds, and (f) is a box plot of the average reward value distribution when the different automated defense methods defend against network attacks with unknown attack paths for 100 rounds.
[0039] Figure 5 A schematic diagram of reward values for defending against new network attacks with unknown attack paths using different automated defense methods described in Implementation Example 11.
[0040] Figure 6 A schematic diagram of the reward values for defending against new network attacks with known attack paths using different automated defense methods described in Implementation 11. DETAILED DESCRIPTION
[0041] In order to make the purpose, technical solutions and advantages of the implementation methods of the present application clearer, the technical solutions in the implementation methods of the present application will be clearly and completely described below in conjunction with the drawings in the implementation methods of the present application. Obviously, the described implementation methods are only part of the implementation methods of the present application, not all of the implementation methods.
[0042] Implementation 1: This implementation provides an automated network defense method integrating threat intelligence, the method comprising:
[0043] Step 1: Measure the relevance of threat intelligence and filter threat intelligence;
[0044] Step 2: Generate defense actions based on the action guidelines and observation indicators in the threat intelligence (CTI) screened in step 1, expand the defense action space of the defense agent, train the defense agent and generate defense strategies;
[0045] Step 3: Based on the defense strategy generated in step 2, a hierarchical reinforcement learning algorithm is designed, the trained hierarchical reinforcement learning defense agent is deployed into the network system, and network attack and defense tests are performed to complete the construction of an automated network defense method that integrates threat intelligence.
[0046] Implementation method 2: This implementation method further limits the automated network defense method for integrating threat intelligence described in implementation method 1. The method for screening the threat intelligence correlation metric in step 1 is:
[0047] Step 101: Collect threat intelligence data in the Structured Threat Information Expression (STIX) format, and collect network and asset information in the network scenario;
[0048] Step 102, extract all action plan statements in the STIX threat intelligence and convert them into vectors; extract network and asset information from the network system and convert them into vectors;
[0049] Step 103: Calculate the cosine similarity between the action policy vector and the asset information vector, use the cosine similarity index to sort the threat intelligence, and screen out the threat intelligence that is highly relevant to the network scenario.
[0050] Implementation method 3: This implementation method further limits the automated network defense method integrating threat intelligence described in implementation method 1. The method for generating a defense strategy in step 2 is:
[0051] Step 201: Generate a general defense action according to an Open Command and Control (OpenC2) instruction for commanding and controlling network defense components;
[0052] Step 202: dynamically generate CTI defense actions based on STIX format threat intelligence. The generation of CTI defense actions uses the action guidelines in STIX threat intelligence. The action guidelines include specific measures for preventing, detecting, and responding to network attacks. The action guidelines and system information are extracted to obtain operations, targets, and parameters. Operations include shutdown, update, and installation. The target is the host IP address. The parameters include any parameters required to execute defense commands, such as IP addresses, ports, and patches.
[0053] Step 203: Construct a defense action space by combining the general defense action and the CTI defense action.
[0054] Implementation method 4. This implementation method further limits the automated network defense method integrating threat intelligence described in implementation method 3. The actions in the general defense action space include general defense actions and CTI defense actions. The general defense actions are generated according to OpenC2 defense commands, including five types: monitoring, analysis, induction, deletion and recovery. The CTI defense actions are generated according to STIX threat intelligence.
[0055] Implementation method five: This implementation method is a further limitation of the automated network defense method integrating threat intelligence described in implementation method one. A hierarchical reinforcement learning algorithm is designed based on the defense strategy generated in step 3. The method for training the defense agent to generate the defense strategy is implemented by adopting a hierarchical proximal strategy optimization algorithm.
[0056] Implementation method six: This implementation method further limits the automated network defense method integrating threat intelligence described in implementation method five. The defense agent includes a control agent and two defense agents. CTI defense actions are added to the agent defense action space.
[0057] Implementation method seven: This implementation method further limits the automated network defense method integrating threat intelligence described in implementation method one. In step 1, the threat intelligence relevance metrics screened include IP addresses, action guidelines, and network topology information.
[0058] Embodiment 8: This embodiment proposes an automated network defense system integrating threat intelligence, the system comprising:
[0059] The screening module is used to measure the relevance of threat intelligence and screen relevant threat intelligence;
[0060] Defense action generation module: users generate defense strategies based on the action guidelines and observation indicators in the relevant threat intelligence selected by the screening module, train defense agents and generate defense strategies;
[0061] The automated defense agent design and training module is used to design a hierarchical reinforcement learning algorithm based on the defense strategy generated by the defense action generation module, deploy the trained hierarchical reinforcement learning defense agent into the network system, and conduct network attack and defense tests to complete the construction of an automated network defense system that integrates threat intelligence.
[0062] Embodiment 9. This embodiment proposes a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the processor runs the computer program stored in the memory, the processor executes the method described in any one of embodiments 1 to 7.
[0063] Embodiment 10: This embodiment proposes a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of the method described in any one of embodiments 1 to 7 are implemented.
[0064] Implementation eleven: This implementation provides an example, which is used to explain the above implementations one to eight. The specific example is as follows:
[0065] See also Figures 1 to 6 To illustrate this embodiment, the method described in this embodiment includes the following steps:
[0066] This embodiment is an automated network defense method that can dynamically update defense strategies.
[0067] First, filter out threat intelligence related to system information, and generate defense strategies based on the action guidelines and observable indicators in the threat intelligence to specifically solve the problem of untimely defense strategy updates. Design a hierarchical reinforcement learning algorithm to train defense agents to generate defense strategies to deal with different attackers.
[0068] The automated network defense method and system integrating threat intelligence described in the present invention dynamically generates defense actions by introducing threat intelligence, trains network defense agents to automatically generate defense strategies, and alleviates network attacks in real time.
[0069] Combination Figure 1 and Figure 2 The present invention proposes an automated network defense method integrating threat intelligence, the method comprising the following steps:
[0070] Step 1: Measure the relevance of threat intelligence and filter relevant threat intelligence;
[0071] The threat intelligence relevance metric is specifically:
[0072] Step 101, data collection, collects STIX threat intelligence data and collects network and asset information in network scenarios;
[0073] Step 102, extracting the action policy statements, extracting all the action policy statements in the STIX report and storing them in coa_set, where coa_set = {coa1, coa2, ..., coa n}, the action plan coa n Convert to vector c n ; Extract network information, and store the network and asset information in the network scene into ass_set, where ass_set = {ass1, ass2, ..., ass n}, and pass the information ass n Convert to vector a n .
[0074] Step 103, STIX threat intelligence correlation measurement, calculate the cosine similarity between the action policy vector and the asset information vector, use the cosine similarity index to sort the threat intelligence, and screen out the threat intelligence that is highly relevant to the network scenario. The specific calculation formula is as follows;
[0075]
[0076] Step 2: Defensive action generation based on OpenC2 and STIX threat intelligence;
[0077] The defense action generation based on OpenC2 and STIX threat intelligence is specifically as follows:
[0078] Step 201, Generate a generic defense action, Generate a generic defense action based on Open Command and Control (OpenC2), OpenC2 defines a set of standard semantic specifications for commanding and controlling network defense components, and provides a standardized interface for introducing network defense systems. Generic defense actions are generated based on OpenC2 defense commands, including five types: monitoring, analysis, induction, deletion, and recovery.
[0079] Step 202, CTI defense action generation, a targeted CTI defense action is dynamically generated based on the STIX format threat intelligence. The generation of CTI defense actions mainly uses the action guidelines in the STIX threat intelligence. The action guidelines include specific measures for preventing, detecting and responding to network attacks. The action guidelines and system information are extracted to obtain operations, targets and parameters. Operations include shutdown, update, installation, etc. The target is the host IP address; the parameters include any parameters required to execute the defense command such as IP address, port, patch, etc.; CTI defense actions can effectively defend against new network attacks.
[0080] Step 203, constructing a defense action space, combining the general defense action and the CTI defense action to construct a defense action space;
[0081] Step 3: Automated Defense Agent Design and Training
[0082] The design and training of the automated defense agent are specifically as follows:
[0083] Step 301, automated defense agent training, the reinforcement learning algorithm uses the hierarchical proximal policy optimization (PPO) algorithm, and trains three reinforcement learning agents, two of which are defense agents and one is a control agent. One of the defense agents is trained against an attack agent with a known system network topology, and the other is trained against an attack agent with an unknown system network topology. Finally, an agent is trained to control which defense agent to choose when facing an attack. The reward function design during training is as follows.
[0084] r t =-0.1(host Attack -1)-server Accack -Defense restore -10(Attack impact )
[0085]
[0086] Step 302, automated network defense agent deployment, deploying the trained hierarchical reinforcement learning defense agent into the network system, and conducting network attack and defense tests.
[0087] Aiming at typical network attack and defense scenarios, the present invention deploys a hierarchical reinforcement learning defense agent to automatically defend against network attacks, and verifies the effectiveness of the automated network defense method that integrates threat intelligence.
[0088] 1. Implementation methods
[0089] According to the present invention Figure 1The automated network defense process shown mainly includes threat intelligence relevance measurement, defense action generation based on OpenC2 and STIX threat intelligence, and automated defense agent training and deployment.
[0090] In step 101, STIX threat intelligence data is collected from ATT&CK, unit42, and MISP Threat Sharing threat intelligence websites, while collecting network and asset information in network scenarios;
[0091] In step 102, the data is extracted and stored in a quantitative manner, and the action plan statements collected from the STIX report are extracted and stored in coa_set, where coa_set = {coa1, coa2, ..., coa n}, the action plan coa n Convert to vector c n ; The network and asset information in the network scene is stored in ass_set, where ass_set = {ass1, ass2, ..., ass n}, and pass the information ass n Convert to vector a n .
[0092] In step 103, the cosine similarity between the two vectors is calculated to measure their correlation. The value range of cosine similarity is between [0,1]. This similarity index is used to sort the threat intelligence and filter out the threat intelligence that is highly relevant to the network scenario. This step provides preparation for the subsequent generation of defense actions based on threat intelligence.
[0093] In step 201, in order to ensure the universality of defense actions, general defense actions are generated according to OpenC2 defense commands. They include five types: monitoring, analysis, induction, deletion, and recovery. Monitoring collects information about malicious activities on the system; analysis collects more information to determine whether there is an attack; induction allows only malicious users to access, and then monitors and analyzes them; deletion removes malicious processes, files, and services, and removes attacks from the host; recovery restores the system to a normal state, but may affect system availability. General defense actions can defend against various known attacks.
[0094] In step 202, targeted CTI defense actions are generated based on the STIX threat intelligence that has been screened and has high relevance to the system. The generation of CTI defense actions mainly utilizes the action guidelines in the STIX threat intelligence. The action guidelines include specific measures to prevent, detect, and respond to network attacks. The action guidelines and system information are extracted to obtain operations, targets, and parameters. Operations include shutdown, update, installation, etc. The target is the host IP address; parameters include any parameters required to execute defense commands such as IP addresses, ports, patches, etc.; CTI defense actions can effectively defend against new network attacks.
[0095] In step 203, the general defense actions generated according to the OpenC2 defense commands and the CTI defense actions generated according to the STIX threat intelligence are combined to form the defense action space of the automated defense agent.
[0096] The network topology of the defense agent training in step 301 is as follows Figure 4 As shown in the figure, during the training process, the defense agent faces the attacker and selects a series of defense actions according to the current environment state to generate an active defense strategy. During the entire training process, the agent conducts strategy exploration, and its exploration rate is high in the initial stage of training, but gradually decreases as the number of training times increases. In the exploration state, the agent will randomly select defense actions in order to more comprehensively explore various defense strategy combinations. In the normal state, the agent will select the best strategy based on previous experience.
[0097] After all the training stages in step 302, the automated defense agent is deployed into the system, an optimized defense strategy is generated and implemented, thereby achieving automated network defense.
[0098] (II) Experimental verification
[0099] In order to verify the effectiveness of the present invention, the method of the present invention is compared with the transfer learning plus word embedding method (TL-Embedding), the integrated reinforcement learning method (Ensemble), the PPO method with transfer learning (TL-PPO), and the hierarchical PPO method with curiosity mechanism (HIPPO). Figure 4 The box plots showing the distribution of reward values for different methods of defending against network attacks show that the average reward value of the method of the present invention is higher, and the upper and lower boundaries of the box plot are closer. After adding threat intelligence to the automated defense, the generated defense strategy has better stability under different test conditions. Figure 5 , Figure 6The reward values of different automated defense methods for defending against new network attacks are shown. For the agent with added CTI defense action, the reward value of the new network attack drops steadily, indicating that after adding threat intelligence, the agent can take proactive measures to prevent the attack. However, for the agent without using CTI defense action, the reward value of the new network attack drops sharply after 100 steps, and the defense effect is not good. By adding threat intelligence, the automated defense system can update malicious activity patterns and attack characteristics in real time, and take proactive measures for defense, reduce dependence on manual intervention, respond to new network attack methods in a timely manner, and improve the dynamic and accuracy of defense.
[0100] Figure 1 Any process or method description in the flowchart described in or otherwise described herein can be understood as representing a module, fragment or portion of a code including one or more executable instructions for implementing the steps of a custom logic function or process, and the scope of the preferred embodiment of the present invention includes other implementations, in which the functions may be performed in a substantially simultaneous manner or in a reverse order according to the functions involved, which should be understood by a person skilled in the art of the present invention. The logic and / or steps represented in the flowchart or otherwise described herein illustrate the possible implementation architecture, functions and operations of the apparatus and methods according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, a program segment, or a part of a code, and the above-mentioned module, program segment, or a part of a code contains one or more executable instructions for implementing the specified logic function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in reverse order, depending on the functions involved. It is also noted that each block in the block diagram or flowchart, and combinations of blocks in the block diagram or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions. For example, what may be considered as an ordered list of executable instructions for implementing a logical function may be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other system that can fetch instructions from and execute instructions in an instruction execution system, apparatus, or device).
[0101] Those skilled in the art will appreciate that the above are only preferred embodiments of the present invention, and the various embodiments of the present disclosure and / or the features described in the claims may be combined or combined in various ways, even if such combinations or combinations are not explicitly described in the present disclosure. It is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art may still modify the technical solutions described in the aforementioned embodiments, or perform equivalent substitutions on some of the technical features therein. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.
[0102] Although preferred embodiments of the present invention have been described, additional changes and modifications may be made to these embodiments by those skilled in the art once the basic inventive concepts are known. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention. Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.
Claims
1. An automated network defense method integrating threat intelligence, characterized in that: The method comprises: Step 1: Measure the relevance of threat intelligence and filter threat intelligence; Step 2: Generate defense actions based on the action guidelines and observation indicators in the threat intelligence screened in step 1, expand the defense action space of the defense agent, train the defense agent and generate defense strategies; Step 3: Based on the defense strategy generated in step 2, a hierarchical reinforcement learning algorithm is designed, the trained hierarchical reinforcement learning defense agent is deployed into the network system, and network attack and defense tests are performed to complete the construction of an automated network defense method that integrates threat intelligence.
2. The automated network defense method integrating threat intelligence according to claim 1, characterized in that: The method for screening threat intelligence relevance metrics in step 1 is: Step 101: Collect threat intelligence data in a structured threat information expression format, and collect network and asset information in a network scenario; Step 102, extract all action plan statements in the STIX threat intelligence and convert them into vectors; extract network and asset information from the network system and convert them into vectors; Step 103: Calculate the cosine similarity between the action policy vector and the asset information vector, use the cosine similarity index to sort the threat intelligence, and screen out the threat intelligence that is highly relevant to the network scenario.
3. The automated network defense method integrating threat intelligence according to claim 1, characterized in that: The method for generating a defense strategy in step 2 is: Step 201, generating a general defense action according to an open command and control instruction, for commanding and controlling a network defense component; Step 202: dynamically generate CTI defense actions based on STIX format threat intelligence. The generation of CTI defense actions uses the action guidelines in STIX threat intelligence. The action guidelines include specific measures for preventing, detecting, and responding to network attacks. The action guidelines and system information are extracted to obtain operations, targets, and parameters. Operations include shutdown, update, and installation. The target is the host IP address. The parameters include any parameters required to execute defense commands, such as IP addresses, ports, and patches. Step 203: Construct a defense action space by combining the general defense action and the CTI defense action.
4. The automated network defense method integrating threat intelligence according to claim 3 is characterized in that: The actions in the general defense action space include general defense actions and CTI defense actions. The general defense actions are generated according to OpenC2 defense commands, including five types: monitoring, analysis, induction, deletion and recovery. The CTI defense actions are generated according to STIX threat intelligence.
5. The automated network defense method integrating threat intelligence according to claim 1, characterized in that: Based on the defense strategy generated in step 3, a hierarchical reinforcement learning algorithm is designed. The method for training the defense agent to generate the defense strategy is implemented by using a hierarchical proximal strategy optimization algorithm.
6. The automated network defense method integrating threat intelligence according to claim 5, characterized in that: The defense agent includes a control agent and two defense agents, and a CTI defense action is added to the agent defense action space.
7. The automated network defense method integrating threat intelligence according to claim 1, characterized in that: The threat intelligence relevance metrics screened in step 1 include IP addresses, action plans, and network topology information.
8. An automated network defense system integrating threat intelligence, characterized in that: The system comprises: A filtering module for filtering threat intelligence relevance metrics; Defense action generation module, where users generate defense strategies based on the action guidelines and observation indicators in the threat intelligence relevance metrics selected by the screening module, train defense agents and generate defense strategies; The automated defense agent design and training module is used to design a hierarchical reinforcement learning algorithm based on the defense strategy generated by the defense action generation module, deploy the trained hierarchical reinforcement learning defense agent into the network system, and conduct network attack and defense tests to complete the construction of an automated network defense method that integrates threat intelligence.
9. A computer device comprising a memory and a processor, characterized in that A computer program is stored in the memory. When the processor runs the computer program stored in the memory, the processor executes the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.