Active security defense system and method based on reinforcement learning and attack intention inference
By using a proactive security defense system based on reinforcement learning and attack intent inference, dynamic attack graphs are generated and defense instruction sets are optimized, solving the problems of lagging attack situation awareness and resource waste in the power Internet of Things, and realizing the accuracy and adaptability of multi-device collaborative defense.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- 四川省大数据技术服务中心
- Filing Date
- 2026-02-13
- Publication Date
- 2026-05-12
AI Technical Summary
Existing technologies cannot accurately perceive real-time changes in the attack situation in the power Internet of Things, resulting in a lag in the prediction of new atomic attacks and cross-regional and cross-device combined attack paths. High-risk attack paths are easily overlooked, and the formulation of defense strategies lacks clear guidance, leading to indiscriminate defense, wasted resources, frequent execution conflicts, chaotic equipment configuration, and inability to adapt to the cybersecurity defense needs of multiple industries.
An active security defense system based on reinforcement learning and attack intent inference is adopted. It generates a dynamic attack graph by collecting and standardizing multi-source data, filters out effective paths, optimizes the defense instruction set using a conflict rule base, and performs dynamic evaluation and feedback during execution to achieve collaborative defense across multiple devices and paths.
It achieves precise positioning and targeted defense for network security across multiple industries, solves the problems of indiscriminate defense and resource waste, ensures the effectiveness of multi-device collaborative defense, reduces the probability of conflict, and adapts to new attacks and changes in the network environment.
Smart Images

Figure CN121727864B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of network security technology, and specifically to a proactive security defense system and method based on reinforcement learning and attack intent inference. Background Technology
[0002] As the core asset carrier and business operation hub of an enterprise or organization, the intranet is vulnerable to attacks that affect its security and normal operation. Proactive defense can identify risks and handle them automatically in the early stages of an attack, preventing the attack from spreading from a single point of infection to a network-wide paralysis, and significantly reducing the scope of losses and repair costs.
[0003] Existing technologies, such as the security defense method, system, storage medium, and server disclosed in patent application CN114143348B, include: establishing a security protection system for the power Internet of Things (IoT); simplifying the security protection system from a collaborative perspective; monitoring and collecting security data within the power IoT security protection system; blocking the attack path upon detecting an intrusion attack; and conducting collaborative defense at the device, region, and global levels. This invention addresses the new and complex attacks facing the power IoT by breaking through traditional single-point and boundary protection models. It introduces the concept of cloud-network-device collaborative defense, formulating a global security protection strategy based on multi-point data collection, and conducting collaborative protection at multiple levels: device, region, and global. The collaborative protection process includes attack graph generation, strategy generation, conflict resolution, and strategy deployment and execution, achieving a better protection effect for the power IoT.
[0004] Existing technologies, such as the one disclosed in application CN119135408A, are a dynamic defense decision-making method and system for phishing emails based on reinforcement learning, which relates to the field of network security technology. The method includes: acquiring historical phishing email data, extracting features, generating multi-dimensional features and constructing a high-quality phishing email feature dataset, constructing and training a dynamic defense decision-making model for phishing emails, encapsulating the trained dynamic defense decision-making model into a standardized interface for integration, acquiring newly received emails, extracting initial email features, making initial judgments, performing deep feature extraction, performing secondary screening, generating accurate identification results, and if the email is phishing, generating a phishing email attack profile and proactive defense strategy; outputting defense action instructions, sending them to defense execution nodes, acquiring user feedback and behavioral data in real time, evaluating the overall phishing email attack situation and defense effectiveness, and updating the dynamic defense decision-making model for phishing emails based on the analysis of feedback.
[0005] Existing technologies generate strategies using static attack graphs and game theory Nash equilibria, and resolve conflicts using attribute labels, DAG graphs, and simple priority rules. While these technologies offer features such as scenario customization, static strategies, and simplified conflicts, static attack graphs cannot accurately perceive real-time changes in the attack landscape within the power IoT. They lag behind in predicting new atomic attacks and cross-regional, cross-device combined attack paths, easily overlooking high-risk attack paths and exposing core power assets to intrusion risks. Furthermore, weightless graphs cannot quantify the risk level of attack paths, resulting in a lack of clear guidance for defense strategy formulation and indiscriminate defense, wasting the protection resources of the power IoT.
[0006] The execution of defense commands is prone to problems such as equipment resource overload, command execution sequence reversal, and failure of cross-device command linkage, which can lead to errors when power-specific security protection equipment (such as power firewalls and encryption gateways) execute defense commands, greatly reducing the effectiveness of defense. Furthermore, priority rules for newly loaded policies can easily lead to global policies being overridden by local policies, disrupting the collaborative defense system of the power Internet of Things cloud network terminal, and even causing confusion in the configuration of equipment at the plant and substation level, affecting power production scheduling.
[0007] Existing technologies revolve around text feature extraction for phishing emails, reinforcement learning decision models, and multi-level progressive defenses. Reinforcement learning is only applied to training decision models for phishing emails. Overall, these technologies are characterized by a single attack type, single-dimensional reinforcement learning, and a lack of multi-device collaborative conflict handling. The entire technical architecture, feature extraction, and model training are all customized for phishing emails, relying on the text, image, or attachment features of emails and the specific data of email sending and receiving systems. There is no generalized design, making it impossible to migrate to other network security scenarios such as server protection, industrial control security, and IoT defense. Even for spam and email bomb attacks, which are also related to email security, it is necessary to rebuild the model and feature system. Furthermore, these technologies cannot collaborate with existing enterprise network security defense systems (such as firewalls, WAFs, or IDS), forming information silos.
[0008] All phishing emails employ the same multi-layered defense strategy, lacking resource allocation and priority execution mechanisms. High-risk phishing emails cannot be prioritized for blocking and tracing, potentially causing significant economic losses for the enterprise; while low-risk phishing emails consume substantial email gateway and sandbox resources, resulting in wasted defense resources and a dilution of core business defense capabilities. Summary of the Invention
[0009] To address the aforementioned technical shortcomings, the present invention aims to provide an active security defense system and method based on reinforcement learning and attack intent inference.
[0010] To solve the above-mentioned technical problems, the present invention adopts the following technical solution: In the first aspect, the present invention provides an active security defense system based on reinforcement learning and attack intent inference, including: an attack inference module, which is used to connect to all existing devices in the intranet, perform full collection and standardization processing of multi-source data, and then generate a dynamic attack graph, which includes several effective paths, the cumulative success probability of each effective path, the path weight, and path information.
[0011] The defense analysis module is used to retrieve defense records from the database, filter out several candidate instruction sets for each effective path, analyze the conflict of each candidate instruction set, and then optimize the instructions. The optimization result serves as the proactive security defense instruction set for each effective path.
[0012] The defense execution module is used to distribute the active security defense instruction set of each valid path to the execution device in the corresponding path to execute the active security defense. During execution, the defense effect and instruction conflict status are evaluated. When the defense effect or instruction conflict status is output as an error, security defense feedback is provided.
[0013] Secondly, the present invention provides an active security defense method based on reinforcement learning and attack intent inference, including: S1, connecting to all existing devices in the intranet, performing full collection and standardization processing of multi-source data, and then generating a dynamic attack graph, which includes several effective paths, the cumulative success probability of each effective path, path weight, and path information.
[0014] S2. Retrieve defense records from the database, filter out several candidate instruction sets for each effective path, analyze the conflict of each candidate instruction set, and then optimize the instructions. The optimization result serves as the active security defense instruction set for each effective path.
[0015] S3. Distribute the active security defense instruction set of each valid path to the execution device in the corresponding path, execute the active security defense, evaluate the defense effect and instruction conflict status during execution, and provide security defense feedback when the defense effect or instruction conflict status outputs an error.
[0016] The beneficial effects of this invention are as follows: 1. The proactive security defense system and method based on reinforcement learning and attack intent inference provided by this invention firstly identifies multiple possible effective attack paths through multi-source data standardization processing, reinforcement learning, and attack intent inference, generates a dynamic attack graph, and constructs and utilizes a conflict rule base to select the best and optimize the defense instruction set for different types of effective paths without conflict. During execution, dynamic evaluation and anomaly feedback are performed, realizing collaborative defense of multiple devices and multiple paths, as well as an upgrade to proactive perception, precise defense, and dynamic optimization of security defense. It can adapt to the network security defense needs of multiple industries, without being bound to specific scenarios or attack types, thus improving the defense effect.
[0017] 2. By optimizing dynamic attack graphs based on reinforcement learning, not only can attack paths be accurately predicted, but the attacker's key intentions can also be quantitatively inferred, providing clear guidance for defense strategy formulation. This completely solves the problems of indiscriminate defense and resource waste in existing technologies, enabling precise risk identification and targeted defense, thus improving the pertinence of defense.
[0018] 3. By building a conflict rule base through conflict testing, it enables accurate identification, quantitative assessment, and targeted resolution of conflicts. It also supports global conflict iterative replacement for high-risk paths and pre-processing adaptation to avoid conflicts for medium- and low-risk paths, reducing the probability of conflicts at the source, ensuring the effectiveness of multi-device, multi-path collaborative defense, and resolving issues such as command failures, device resource overload, and business anomalies in collaborative defense.
[0019] 4. By classifying effective paths into high, medium, and low risk levels, targeted execution strategies are designed to avoid cross-path conflicts. This ensures both rapid and effective defense of high-risk attack paths and avoids medium- and low-risk paths consuming too many core resources, thus guaranteeing the rational allocation of defense resources and the timeliness of defense against high-risk targets.
[0020] 5. By deeply integrating reinforcement learning technology into multiple processes such as data collection optimization and dynamic attack graph construction, strategies can be dynamically adjusted according to actual defense scenarios, defense logic can be continuously optimized, and new attack methods and network environment changes can be quickly adapted to solve the problems of static traditional defense solutions and poor adaptability to new attacks. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a schematic diagram of the system structure connection of the present invention.
[0023] Figure 2 This is a schematic diagram of the implementation steps of the method of the present invention. Detailed Implementation
[0024] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0025] Please see Figure 1 As shown, the proactive security defense system based on reinforcement learning and attack intent inference includes the following modules: an attack inference module, which is used to connect to all existing devices in the intranet, collect and standardize multi-source data, and then generate a dynamic attack graph. The dynamic attack graph includes several effective paths, the cumulative success probability of each effective path, the path weight, and path information.
[0026] In a specific embodiment, the full collection and standardization process of the multi-source data is as follows: the multi-source data includes network layer data, host layer data, asset and vulnerability data, and threat intelligence data. The multi-source data is uniformly converted into a fixed JSON format, and unified basic fields are defined. A unique identifier is generated based on the data source, timestamp, and core attributes in the unified fields. The unique identifier is compared to eliminate duplicate data. Invalid data is eliminated by using preset unified field filtering rules and a noise reduction model optimized by reinforcement learning, thus completing the data processing.
[0027] In the above context, network layer data refers to various types of data generated in the network transmission link, reflecting the network connection and data transmission status between devices in the intranet and between the intranet and the extranet. It is the core data for detecting network layer attacks (such as port scanning, lateral movement, DDoS, and initial confidence scoring), and originates from network security and transmission equipment such as firewalls, WAFs, IDS / IPS, core switches, routers, and network traffic probes; including abnormal packet characteristics, device interface bandwidth utilization, network latency, and packet loss rate.
[0028] It should be noted that the initial confidence score for multi-source data has been specifically disclosed in the invention patent with publication number CN112019519B, and will not be traced back here.
[0029] Host-level data refers to various data generated by a single host, server, or terminal on the intranet during operation. It reflects the local operating status of a single device and is the core data for detecting host-level attacks (such as Trojan horse execution, privilege escalation, and process injection). It originates from operating system logs of servers and terminals, host intrusion detection systems, terminal detection and response systems, operation and maintenance audit systems, edge computing terminal local collection modules, and other systems or modules; including logged-in users, login IPs, and virus scanning records.
[0030] Asset and vulnerability data refers to the basic ledger information of all network assets and devices on the internal network, as well as data related to security vulnerabilities and configuration defects existing in these assets. It is the core foundational data for constructing attack graph nodes and deducing attack paths, reflecting the internal network's security baseline and attack entry points. This data originates from the internal network asset management system, vulnerability scanning tools (such as Nessus and AWVS), or patch management systems, and includes unique asset IDs, asset IPs, asset types (servers, terminals, network devices, databases, etc.), associated business systems, operating systems, application service versions, vulnerability risk levels, and deployment locations. The vulnerability risk level can be obtained using CVSS (Common Vulnerability Scoring System), a standard technique in this field, and will not be elaborated upon here.
[0031] Threat intelligence data refers to structured data related to cybersecurity threats from both internal and external sources. It is core data for early detection of external attacks, matching abnormal behavior within the internal network, and accurately locating the source of attacks, providing external references for inferring attack intent and formulating defense strategies.
[0032] It should be noted that the unified basic fields include data source, timestamp, core attributes, and additional information. Among them, the core attributes are defined differently according to the data type. For example, network layer data is "source IP + target IP + port + protocol", host layer data is "asset IP + operation type + process / log identifier", asset and vulnerability data is "asset IP + CVE number + vulnerability risk level", and threat intelligence data is "malicious identifier + threat type + intelligence source".
[0033] In the above, the process of removing invalid data using a preset unified field filtering rule and a reinforcement learning-optimized noise reduction model is as follows: The preset unified field filtering rule is used to remove low-value or worthless data such as data that does not meet compliance standards or is outdated. This rule is set by the operations and maintenance personnel according to their operational needs. For example, it can be used to remove data such as missing required fields in core attributes, data generated outside the valid window for timestamps, and invalid timestamp formats.
[0034] The reinforcement learning-optimized noise reduction model involves acquiring at least 10,000 historical fuzzy data points from the backend. These data are manually labeled as valid or invalid, and a pre-defined, uniform field filtering rule is used to initially screen and retain / remove data, marking them as valid and invalid samples, respectively. The data comprises 30% network layer data, 30% host layer data, 20% asset vulnerabilities data, and 20% threat intelligence data. The labeled samples are used to train a LightGBM basic classification model, optimizing the cross-entropy loss. Training stops once baseline accuracy (e.g., 85%) is achieved, completing model pre-training. The pre-trained model is then used as the initial strategy and integrated into the business system's closed loop: the model's output "valid data" enters the attack graph module, generating reward signals based on actual utilization effects; every 24 hours, the PPO algorithm's policy network is updated with new reward signals to optimize model parameters. Validation Phase: Input: Feature vector of fuzzy data; Inference: Model output label + confidence score; Judgment rules: Confidence score ≥ 0.8 → Remove / retain by label; Confidence score < 0.8 → Manual review. Randomly select 5% of the finally retained valid data for manual review. If the false positive rate is > 5%, immediately pause processing, update the screening rules (such as adding new false positive rules) or retrain the model; until the false positive rate is < 5%.
[0035] It should be noted that the training and validation process of the reinforcement learning-optimized denoising model is a common technical means in the fields of network security data processing and proactive security defense. The specific implementation process can be found on the Internet.
[0036] In another specific embodiment, the specific process of generating the dynamic attack graph is as follows: S1001, the processed multi-source data is precisely mapped according to the preset data type-graph element type mapping relationship. Based on the mapped graph elements, the core elements of the dynamic attack graph are defined. The core elements include nodes and edges. At the same time, the attributes of the core elements are initialized according to the preset node initialization rules and edge initialization weight rules to complete the logical association of graph elements.
[0037] It should be noted that the preset data type-graph element type mapping relationship clearly defines which node or edge each type of data can be transformed into in the dynamic attack graph, while defining the belonging dimension of graph elements to ensure the uniqueness and accuracy of the mapping and avoid mismatch between data and graph elements. This setting and adjustment is made by operations and maintenance personnel according to network security monitoring needs. For example: the graph element types for asset-vulnerability data mapping are asset nodes or vulnerability nodes, etc. One piece of asset-vulnerability data corresponds to one asset node and one vulnerability node attached to that node. The asset is the attack target, and the vulnerability is the attack entry point. The graph element types for host layer data mapping are host status nodes or attack behavior nodes, etc. Host IP is associated with asset nodes, abnormal behavior / privilege escalation operations are transformed into attack behavior nodes, and host protection status is transformed into host status nodes. The graph element types for threat intelligence data mapping are threat intelligence nodes or attack group nodes, etc., which are external threat source nodes in the attack chain and are associated with corresponding attack behavior nodes / asset nodes. The graph element types for network layer data mapping are asset access edges, vulnerability exploitation edges, or attack behavior progression edges, etc., which represent the attack relationship between nodes. For example, the access behavior from source IP to target IP is transformed into asset access edges, and port scanning and vulnerability detection are transformed into vulnerability exploitation edges.
[0038] The core elements of the dynamic attack graph are defined based on the mapping results: Nodes are the vertices of the directed weighted attack graph, carrying core entity information such as assets, vulnerabilities, attack behaviors, threat sources, and device status in the attack chain. Each node has a unique type identifier and attribute set, divided into basic nodes (required, included in all scenarios) and extended nodes (optional, supplemented according to scenario requirements). The core identifier of a node is "unique data identifier + graph element type code". For example, the core identifier of an asset node (various asset entities attacked on the internal network are the core targets of the attack, and the basic node) is: asset IP + CVE number + vulnerability risk level (unique data identifier) + ZC (graph element type code).
[0039] An edge is a directed edge connecting two nodes, representing the attack path, behavioral progression, vulnerability exploitation, and other relationships of an attacker from one node to another. It is the core link in the attack chain deduction. Each edge has a unique type identifier, weight attribute, and direction (from the attack source node to the attack target node). All edges are weighted and computable. The core weight represents "the probability of the attack behavior corresponding to this edge being successfully implemented". For example, the asset access edge (the attacker's basic network access behavior to the target asset, which is the pre-attack link) connects the node pair (source → target): attack behavior node → asset node / threat intelligence node → asset node. The core information carried includes access protocol, source / target IP, port number, access frequency, and traffic characteristics (from network layer data). Element type encoding: FW.
[0040] All edges are directed, and their directions strictly follow the "from source to target" logic of the attack chain (e.g., if an attacker initiates vulnerability exploitation from an attack behavior node, the edge direction is attack behavior node → vulnerability node); a single edge connects only two nodes, and the association of multiple nodes is achieved through the combination of multiple edges; the edge type is strongly bound to the node type, and only preset node pairs are allowed to establish edges of the corresponding type to avoid invalid associations.
[0041] The definition of the core elements of a dynamic attack graph can be adaptively set and adjusted according to the specific needs of the scenario. The above examples are only illustrative and not the only limitation.
[0042] Core element attribute initialization: Based on preset node initialization rules and edge initialization weight rules, the attributes of nodes and edges are initialized in batches. The initialized attributes are quantifiable values that can be calculated and dynamically updated.
[0043] Node initialization rules: Based on the node type, extract fields from the core attributes of multi-source data to complete the initialization of exclusive business attributes. For example: Asset node unique ID: Initialization value basis: unique identifier + node type code + timestamp suffix, quantification / formatting requirements: string, example value: ZC_1921681100_20260126100530 (asset node_core database IP_timestamp).
[0044] Edge initialization weight rules: These include general basic attributes and specific business attributes. The general basic attributes are the same as the node initialization rules and will not be elaborated upon here. The specific business attributes are the weight attributes, and the specific calculation process is as follows: W base =S cre ×K threat Among them, W base S represents the weight. cre表示 Initial confidence score of multi-source data inherited from the edge, K threat This represents the threat level coefficient of edge association.
[0045] It should be noted that the value is determined based on the vulnerability risk level corresponding to the edge, ranging from 0.3 to 1.0. High risk = 1.0, medium risk = 0.7, low risk = 0.3.
[0046] S1002. Based on the logical association of graph elements, an improved Dijkstra algorithm optimized by reinforcement learning is used to deduce all paths from potential entry points to core assets by an attacker. All reachable edges are traversed, and the initial cumulative success probability and initial path weight of each path are calculated. The cumulative success probability and path weight of each path are obtained through dynamic optimization by the reinforcement learning model.
[0047] Preferably, the specific deduction process is as follows: according to the preset entry node filtering rules, multiple entry nodes are selected; according to the preset target node filtering rules, multiple target nodes are selected; then, the improved Dijkstra algorithm is used to solve all reachable paths, output all paths from the entry point to the core asset, and output the initial cumulative success probability and initial path weight of each path.
[0048] It should be noted that both the preset entry point node filtering rules and target node filtering rules are set by operations and maintenance personnel according to scenario requirements. For example: Entry point node filtering rules: 1. Threat intelligence nodes associated with malicious IPs / domains on the external network and APT attack sources; 2. Nodes representing initial attacks launched from the external network into the internal network (such as port scanning, external SSH attempts, SQL injection probing, etc.); 3. High-risk vulnerability nodes exposed to the external network or without protection; 4. Business asset nodes directly connected to the external network without boundary protection (such as web server nodes published on the public network). Target node filtering rules: 1. Belonging to core databases, scheduling center servers, and financial system servers; 2. Asset nodes carrying core internal network business, such as business middleware, identity authentication servers, and data storage servers; 3. Hub nodes associated with other internal network assets, such as asset nodes corresponding to core switches and edge computing gateways.
[0049] The improved Dijkstra algorithm is a standard technique in the field of network security and will not be elaborated upon here.
[0050] The process of constructing the reinforcement learning model is the same as that of optimizing the denoising model using reinforcement learning, and will not be repeated here.
[0051] S1003. Select each path with a cumulative success probability ≥ 10% and a path weight ≥ 5 as each valid path, obtain the path information of each valid path, and output a dynamic attack graph. Each valid path in the dynamic attack graph contains the complete link of attack entry point - intermediate asset - core target, initial cumulative success probability, initial path weight, and path information.
[0052] The path information includes the path generation time, the multiple data sources upon which the path was generated, the unique IDs of all nodes in the path, node types, and core attributes. This information can be viewed during the visualization of the dynamic attack graph.
[0053] The defense analysis module is used to retrieve defense records from the database, filter out several candidate instruction sets for each effective path, analyze the conflict of each candidate instruction set, and then optimize the instructions. The optimization result serves as the proactive security defense instruction set for each effective path.
[0054] In a specific embodiment, the process of selecting several candidate instruction sets for each valid path is as follows: using each valid path as an index, several defense instruction sets strongly associated with each valid path are selected from the defense records as candidate instruction sets, and a path ID is set for each valid path to form an association list of path ID and candidate instruction sets.
[0055] A strong association is determined if any two of the following conditions are met: (1) the defense target of the defense instruction set is consistent with the core target of the effective path; (2) the defense instruction set covers any node of the effective path; and (3) the attack behavior of the defense instruction set is consistent with the attack behavior of the effective path.
[0056] It should be noted that an attack refers to a single atomic attack or a series of related operations launched by an attacker against a network, host, asset, or vulnerability with malicious intent.
[0057] Information about attack behaviors is scattered across network layer data, host layer data, asset and vulnerability data, and threat intelligence data. Network layer data and host layer data are the core sources for directly extracting attack behaviors, while asset and vulnerability data and threat intelligence data are supplementary sources for inferring attack behaviors. All extractions are based on unified basic fields in a standardized JSON format of multi-source data. For example, the unified basic fields for network layer data include alarm type, source IP, target IP, communication protocol, access behavior, traffic characteristics, and port number. The types of attack behaviors that can be extracted or inferred include: probing (port scanning, asset probing, vulnerability scanning), injection and exploitation (…). The extraction or derivation logic for SQL injection, XSS cross-site scripting, remote code execution detection, access-related attacks (illegal access across network segments, malicious external connections, brute-force attacks on default ports), and traffic attack-related attacks (DDoS / DoS attacks, abnormal traffic packet sending, fragmented packet attacks) is as follows: Using the "alarm type" as the core identifier, combined with fields such as source or target IP, port, and protocol, it is directly mapped to the corresponding atomic attack behavior; for example: an alarm type of "SQL injection" → extracted as an atomic attack behavior of "SQL injection detection"; a source IP of a malicious external IP, a target IP of a core internal asset, and port 3306 → extracted as an attack behavior of "remote database access attempt".
[0058] As mentioned above, the path ID must be unique across the entire network, traceable, and contain core characteristics. The generation rules must balance uniqueness, readability, and manageability. The path ID generation rules are set and adjusted by the operations and maintenance personnel according to the environment requirements, and are not limited here.
[0059] In another specific embodiment, the specific analysis process of the proactive security defense instruction set is as follows: S2001, firstly, each effective path is classified into high-risk paths and medium-low-risk paths; high-risk paths are analyzed using parallel defense execution rules, and medium-low-risk paths are analyzed using sequential defense execution rules;
[0060] It should be noted that the classification of high-risk and medium-to-low-risk paths is as follows: a cumulative success rate of ≥30% is considered a high-risk path, and otherwise it is considered a medium-to-low-risk path.
[0061] Parallel defense execution rules: S2011, combine all candidate instructions of all high-risk paths into a global instruction pool, based on the conflict rule base, detect the logical conflict score, resource conflict score and cross-device linkage conflict score of each candidate instruction set, and denot them as L, R and C respectively, then calculate the comprehensive conflict score, denoted as Q, and retain the 3 candidate instruction sets with the smallest Q in each effective path.
[0062] It should be noted that Q = L × 0.4 + R × 0.3 + C × 0.3.
[0063] S2012. Extract the L, R, C, execution device, execution object, action direction, and core instruction of each candidate instruction set retained in each valid path. Select the candidate instruction set with the smallest Q from each candidate instruction set retained in each valid path as the initial optimal candidate set. Merge the core instructions of all initial optimal candidate sets. Extract global instruction features according to execution device, execution object, and action direction to form an initial global instruction pool. Based on the conflict rule base, only detect instruction conflicts between different valid paths. Output the incremental values of logical conflict score, resource conflict score, and cross-device linkage conflict score, denoted as ΔL, ΔR, and ΔC, respectively. L, R, and C remain unchanged. Then, add them together to obtain the cross-path corrected L, R, and C scores of each initial optimal candidate set, denoted as L1, R1, and C1. Calculate the global comprehensive score G.
[0064] In the above, when L1, R1 or C1 > 1, the value is 1, and when L1, R1 or C1 < 0, the value is 0.
[0065] First, L1, R1, and C1 are calculated according to the calculation method of Q to obtain the cross-path corrected comprehensive conflict score of each initial optimal candidate set. Then, the mean is calculated to obtain the global comprehensive score G.
[0066] If S2013 and G < 0.3, make slight optimizations to the initial global instruction pool; if 0.3 ≤ G, iteratively replace the candidate set and rebuild the global instruction pool. After optimizing the global instruction pool, obtain the active security defense instruction set for each candidate path.
[0067] In the above, the process of S2013 is as follows: S2131, minor optimization: According to L1, R1, and C1, the instructions are divided into minor logical conflicts, minor resource conflicts, and minor cross-device linkage conflicts. Minor logical conflicts: reduce the effective range of adjustment instructions; minor resource conflicts: execute instructions from the same device in two batches with an interval of 1-2 seconds, with no resource overlap; minor cross-device linkage conflicts: merge similar instructions; delay the timing of non-core instructions.
[0068] It should be noted that minor logical conflicts are defined as L1 < 0.4, minor resource conflicts as R1 < 0.3, and minor cross-device linkage conflicts as C1 < 0.3.
[0069] S2132. Iterative replacement of candidate sets for some paths: Based on L1, R1, and C1, select each effective path that needs to be iterated, and iteratively replace its initial optimal candidate set with the second smallest candidate set of Q for the corresponding path. Reconstruct the global instruction pool and recalculate the G value until a global instruction pool with global G < 0.3 is obtained.
[0070] In the above, if the cross-path correction of the overall conflict score of the initial optimal candidate set in a certain effective path is >0.3, the effective path is determined to be an effective path that needs to be iterated.
[0071] Preferably, the sequential defense execution rule is as follows: S2021, sort all low- and medium-risk paths in descending order of cumulative success probability, and use the sorting result as the execution order. Only retain candidate instruction sets whose attack behavior is consistent with the attack behavior of the effective path. Then, based on the conflict rule base, detect the conflict within the candidate instruction set, obtain L, R, and C in each candidate instruction set, calculate the comprehensive conflict score of each candidate instruction set, and remove candidate instruction sets with a comprehensive conflict score ≥ 0.6.
[0072] S2022. Select the one with the smallest overall conflict score as the main selection set; before executing defense on each path, collect the full execution status of the previous executed path, and then perform pre-adaptation optimization on the main selection set to obtain the proactive security defense instruction set.
[0073] It should be noted that the full execution status is all the data that reflects the effect of the defense command execution, including the real-time status of the execution object (asset, node, IP) (whether it has been blocked, hardened or patched), the CPU, memory, concurrent connection count and real-time utilization of the execution device, etc.
[0074] In the above, the process of pre-adaptation optimization is as follows: First, remove the execution devices and instructions that are the same as those of the previous executed path from the main selection set. Then, remove the instructions that conflict with the execution instructions of the previous executed path. Adjust the execution scope. Using the conflict rule table, output the logical conflict score, resource conflict score, and cross-device linkage conflict score of the adjusted main selection set again. Calculate the comprehensive conflict score of the main selection set according to the calculation method of Q. If it is less than 0.3, optimize according to the slight optimization method in S2131 to obtain the proactive security defense instruction set. If it is greater than 0.3, select the candidate instruction set with the second smallest comprehensive conflict score. Then, optimize according to the adjustment method of the main selection set and the iterative replacement method of the candidate sets of some paths in S2131 until the comprehensive conflict score is less than 0.3, and the optimization is completed.
[0075] It should be noted that the instruction conflict determination is as follows: the operation and maintenance personnel set the instruction set corresponding to each instruction for conflict. If the instruction set that should conflict with an instruction contains the execution instruction of the previous executed path, it can be determined that the instruction is the instruction that conflicts with the execution instruction of the previous executed path.
[0076] In a specific embodiment, a conflict rule base construction module is also included. The construction process is as follows: S4001, replicate all devices in the intranet, build a 1:1 conflict test environment, set up several defense scenarios, each defense scenario includes multiple high-risk paths executed globally in parallel or multiple medium- and low-risk paths executed sequentially. Among them, each high-risk path and each medium- and low-risk path in each defense scenario are not completely the same, and obtain several defense instruction sets for each type of conflict in each defense scenario from the intranet.
[0077] S4002. Simulate each defense scenario in the conflict test environment, and then issue the preset defense command sets in each defense scenario through the command issuing tool in sequence. Detect whether a conflict is triggered in real time through the conflict monitoring tool. After a conflict is triggered, collect the core test data.
[0078] It should be noted that the defense scenarios and preset defense command sets during the testing process were all set and adjusted by professional testers according to the testing requirements, and no specific limitations are made here.
[0079] It should also be noted that, in order to avoid errors from a single test, each defense scenario needs to be tested at least 5 times, and then the average of the results from multiple tests is taken.
[0080] S4003. Extract logical conflict data, resource conflict data, and cross-device linkage conflict data from the core test data, quantify them into logical conflict scores, resource conflict scores, and cross-device linkage conflict scores in each conflict test group, and then store them in association with the corresponding defense scenarios and defense instruction set features, thereby forming a conflict rule base.
[0081] The core test data mentioned above consists of all data obtained from the execution device, monitoring platform, and command issuance tool, including command execution status data, device operation monitoring data, cross-device linkage data, and business impact data. Command execution status data includes command issuance success or failure, command failure points, and command execution deviations. Device operation monitoring data includes CPU usage, concurrent connections, device overload duration, and resource contention points. Cross-device linkage data includes command synchronization duration of linked devices, command execution deviations, number of coordination failures, and linkage interruption points. Business impact data includes whether conflicts affect business operations, the scope of impact, and the duration of impact.
[0082] Logical conflict data includes conflict trigger rate, command failure rate, and business impact scope ratio. The conflict trigger rate is the number of times a conflict is triggered in a defense scenario divided by the number of repeated tests. The command failure rate is the number of command failures in a defense scenario divided by the total number of defense commands issued. The business impact scope ratio is the number of asset nodes actually affected after a conflict is triggered in a defense scenario divided by the total number of asset nodes. The conflict trigger rate, command failure rate, and business impact scope ratio are averaged, and the result is used as the logical conflict score. These parameters can be obtained from command issuance tool logs, conflict monitoring tool logs, and the internal network basic asset ledger.
[0083] Resource conflict data includes peak resource utilization, duration of resource contention, average device overload, and number of affected devices. Peak resource utilization refers to the highest actual utilization rate of core resources (CPU, memory, concurrent connections) of devices affected by the resource conflict within the collection window after the conflict is triggered. Duration of resource contention refers to the cumulative duration for which the resource utilization rate of multiple devices continuously exceeds the preset resource contention alarm threshold within the collection window after the conflict is triggered; the maximum value is taken as the duration of resource contention. Average device overload represents the total resource overload of all devices that experience resource overload (resource utilization rate continuously exceeding the preset threshold) after the conflict is triggered. The resource overload is the average value of the resource contention alarm threshold (devices with resource contention alarms), where the resource overload is the difference between the continuous resource occupancy rate and the preset resource contention alarm threshold. The number of affected devices is the total number of execution devices whose abnormalities are directly caused by resource conflicts. A device is considered affected if any of the following conditions are met: 1. Resource occupancy rate continuously exceeds the resource contention threshold for ≥30 seconds; 2. Resource overload occurs (single resource peak value ≥ overload threshold); 3. Defense command execution is stalled, times out, or fails due to resource contention; 4. Device performance significantly degrades (e.g., service response timeout, sudden drop in data transmission rate) and is directly caused by resource contention. All of the above parameters can be extracted from the conflict monitoring tool logs and the device operation monitoring platform.
[0084] It should be noted that all the thresholds mentioned above were set and adjusted by the operations and maintenance personnel according to environmental requirements.
[0085] The peak resource utilization rate, duration of resource contention, average equipment overload, and number of affected equipment are normalized, and then the average value is used to calculate the resource conflict data.
[0086] Cross-device linkage conflict data includes the number of coordination failures, the linkage defense failure ratio, the linkage device anomaly ratio, and the command synchronization deviation duration. The number of coordination failures refers to the effective number of linkage execution failures caused by chaotic coordination logic or abnormal command interaction when multiple devices execute linkage defense commands within the collection window after a conflict is triggered. The linkage defense failure ratio refers to the proportion of failed linkage defense links in a defense scenario to the total number of linkage defense links in that scenario after a conflict is triggered. The linkage device anomaly ratio refers to the proportion of devices experiencing coordination execution anomalies among the participating devices after a conflict is triggered, out of the total number of linkage devices. The command synchronization deviation duration refers to the average deviation duration between the actual execution time of the same linkage defense command on each linkage device and the preset command baseline execution time after a conflict is triggered. All of these parameters can be extracted from the cross-device linkage monitoring platform and device execution logs.
[0087] The number of collaborative failures, the failure ratio of linked defenses, the ratio of linked equipment anomalies, and the duration of command synchronization deviations are normalized, and then the cross-device linkage conflict score is obtained by averaging.
[0088] It should be noted that there are linkage instructions in the instruction set. Linkage instructions require two or more devices to coordinate and execute synchronously according to preset logic. The device that executes the linkage instruction is the linkage device.
[0089] The defense execution module is used to distribute the active security defense instruction set of each valid path to the execution device in the corresponding path to execute the active security defense. During execution, the defense effect and instruction conflict status are evaluated. When the defense effect or instruction conflict status is output as an error, security defense feedback is provided.
[0090] In a specific embodiment, the specific process of the security defense feedback is as follows: real-time monitoring of the status data of the execution device, logical conflict data in the intranet, asset conflict data and cross-device linkage conflict data; based on the status data of the execution device, judging the defense effect during the execution process; and using the logical conflict data, asset conflict data and cross-device linkage conflict data in the intranet to analyze the instruction conflict status.
[0091] In the above, the defense effectiveness is judged during the execution process as follows: the status data of the execution device is obtained from the real-time monitoring logs and status interface of the execution device, including: attack behavior monitoring data, defense action execution data, node protection status data, and business operation status data; among them, attack behavior monitoring data includes the number of times the attack behavior is triggered, the trigger frequency, and the access success rate of the source IP, etc.; defense action execution data includes the actual execution status of the defense command (executed, executing, or not executed, corresponding to values 1, 0, and -1, respectively), the execution coverage, and the duration of the action's effectiveness, etc.; node protection status data includes whether the access is blocked, such as whether the node port is closed and whether the vulnerability has been patched; business operation status data includes the operation status of the corresponding business after the defense command is executed (normal, stuck, interrupted, corresponding to values 1, 0, and -1, respectively).
[0092] The attack behavior monitoring data range is set by the operation and maintenance personnel. If the attack behavior monitoring data is within the attack behavior monitoring data range, it indicates that the attack behavior is qualified; otherwise, it is judged as unqualified. The attack behavior analysis method is used to obtain whether the defense action, node protection, and business operation are qualified. If any three are unqualified, the defense effect outputs an error; if all are qualified, the defense effect outputs a normal; if only two are unqualified, the defense effect outputs a warning.
[0093] Command Conflict Status: Logical conflict data, asset conflict data, and cross-device linkage conflict data in the intranet are analyzed according to the analysis methods of logical conflict score, resource conflict score, and cross-device linkage conflict score to obtain the logical conflict score, resource conflict score, and cross-device linkage conflict score after execution. Then, the comprehensive conflict score after execution is calculated according to the Q calculation method. If at least one of the comprehensive conflict score, logical conflict score, resource conflict score, or cross-device linkage conflict score after execution is greater than 0.3, the command conflict status is output as error; otherwise, it is output as normal.
[0094] When the defense outputs an error, it is immediately fed back to the defense analysis module, which regenerates the proactive security defense instruction set and reissues it for execution. If the new instruction set is invalid, the manual intervention process is triggered, and the operation and maintenance personnel are notified to handle the emergency.
[0095] When an error is output in the command conflict state, the preset rollback command is immediately executed to restore the device configuration, and feedback is sent to the defense analysis module to regenerate the proactive security defense command set and re-execute it; if the rollback fails or the command conflict state is still output as an error after re-execution, the defense of the corresponding path is suspended and the operation and maintenance are notified for emergency handling.
[0096] A database used to store defense records.
[0097] Please see Figure 2As shown, the proactive security defense method based on reinforcement learning and attack intent inference includes: S1, connecting to all existing devices in the internal network, collecting and standardizing multi-source data, and then generating a dynamic attack graph. The dynamic attack graph includes several effective paths, the cumulative success probability of each effective path, the path weight, and path information.
[0098] S2. Retrieve defense records from the database, filter out several candidate instruction sets for each effective path, analyze the conflict of each candidate instruction set, and then optimize the instructions. The optimization result serves as the active security defense instruction set for each effective path.
[0099] S3. Distribute the active security defense instruction set of each valid path to the execution device in the corresponding path, execute the active security defense, evaluate the defense effect and instruction conflict status during execution, and provide security defense feedback when the defense effect or instruction conflict status outputs an error.
[0100] The examples described in this invention are not limited to the specific embodiments listed above. The examples are merely illustrative to facilitate understanding of the invention and do not constitute a limitation on the scope of protection of this invention. Any modifications, equivalent substitutions, etc., made within the spirit and principles of this invention should be included within the scope of protection.
[0101] The above description is merely an example and illustration of the concept of the present invention. Those skilled in the art can make various modifications or additions to the specific embodiments described or use similar methods to replace them, as long as they do not deviate from the concept of the invention or exceed the scope defined in this specification, they should all fall within the protection scope of the present invention.
Claims
1. A proactive security defense system based on reinforcement learning and attack intent inference, characterized in that, Includes the following modules: The attack inference module is used to connect to all existing devices on the intranet, collect and standardize multi-source data, and then generate a dynamic attack graph. The dynamic attack graph includes several effective paths, the cumulative success probability of each effective path, the path weight, and path information. The specific process of generating the dynamic attack graph is as follows: S1001, the processed multi-source data is precisely mapped according to the preset data type-graph element type mapping relationship. Based on the mapped graph elements, the core elements of the dynamic attack graph are defined. The core elements include nodes and edges. At the same time, the attributes of the core elements are initialized according to the preset node initialization rules and edge initialization weight rules to complete the logical association of graph elements. S1002. Based on the logical association of graph elements, using the improved Dijkstra algorithm optimized by reinforcement learning, we deduce all paths from potential entry points to core assets by attackers, traverse all reachable edges, and calculate the initial cumulative success probability and initial path weight for each path. The cumulative success probability and path weight of each path are obtained through dynamic optimization using a reinforcement learning model. S1003. Select each path with a cumulative success probability ≥ 10% and a path weight ≥ 5 as each valid path, obtain the path information of each valid path, and output a dynamic attack graph. Each valid path in the dynamic attack graph contains the complete link of attack entry point - intermediate asset - core target, initial cumulative success probability, initial path weight, and path information. The defense analysis module is used to retrieve defense records from the database, filter out several candidate instruction sets for each effective path, analyze the conflict of each candidate instruction set, and then optimize the instructions. The optimization result serves as the active security defense instruction set for each effective path. The specific process of selecting several candidate instruction sets for each valid path is as follows: using each valid path as an index, select several defense instruction sets that are strongly associated with each valid path from the defense records as candidate instruction sets, set a path ID for each valid path, and form an association list of path ID-candidate instruction sets. Among them, any two of the following conditions must be met to be considered as a strong association: (1) the defense target of the defense instruction set is consistent with the core target of the effective path; (2) the defense instruction set covers any node of the effective path; (3) the attack behavior of the defense instruction set is consistent with the attack behavior of the effective path. The specific analysis process of the proactive security defense instruction set is as follows: S2001, First, each effective path is classified into high-risk paths and medium-low-risk paths; high-risk paths are analyzed using parallel defense execution rules, and medium-low-risk paths are analyzed using sequential defense execution rules; Parallel defense execution rules: S2011, collect all candidate instructions of all high-risk paths and merge them into a global instruction pool. Based on the conflict rule base, detect the logical conflict score, resource conflict score and cross-device linkage conflict score of each candidate instruction set, which are denoted as L, R and C respectively. Then calculate the comprehensive conflict score, which is denoted as Q. Keep the 3 candidate instruction sets with the smallest Q in each effective path. S2012. Extract the L, R, C, execution device, execution object, action direction, and core instruction of each candidate instruction set retained in each valid path. Select the candidate instruction set with the smallest Q from each candidate instruction set retained in each valid path as the initial optimal candidate set. Merge the core instructions of all initial optimal candidate sets. Extract global instruction features according to execution device, execution object, and action direction to form an initial global instruction pool. Based on the conflict rule base, only detect instruction conflicts between different valid paths. Output the logical conflict score increment, resource conflict score increment, and cross-device linkage conflict score increment, denoted as ΔL, ΔR, and ΔC, respectively. L, R, and C remain unchanged. Then, add them accordingly to obtain the cross-path corrected L, R, and C scores of each initial optimal candidate set, denoted as L1, R1, and C1. Calculate the global comprehensive score G. If S2013 and G < 0.3, make slight optimizations to the initial global instruction pool; if 0.3 ≤ G, iteratively replace the candidate set and rebuild the global instruction pool. After optimizing the global instruction pool, obtain the active security defense instruction set for each candidate path. The defense execution module is used to distribute the active security defense instruction set of each valid path to the execution device in the corresponding path to execute the active security defense. During execution, the defense effect and instruction conflict status are evaluated. When the defense effect or instruction conflict status is output as an error, security defense feedback is provided.
2. The proactive security defense system based on reinforcement learning and attack intent inference according to claim 1, characterized in that, The specific process of full collection and standardization of the multi-source data is as follows: Multi-source data includes network layer data, host layer data, asset and vulnerability data, and threat intelligence data. The multi-source data is uniformly converted into a fixed JSON format, and unified basic fields are defined. A unique identifier is generated based on the data source, timestamp, and core attributes in the unified fields. The unique identifier is compared to eliminate duplicate data. By using pre-defined unified field filtering rules and a noise reduction model optimized by reinforcement learning, invalid data is removed, and data processing is completed.
3. The proactive security defense system based on reinforcement learning and attack intent inference according to claim 1, characterized in that, The process in S2013 is as follows: S2131, Minor Optimization: Based on L1, R1, and C1, instructions are divided into minor logical conflicts, minor resource conflicts, and minor cross-device linkage conflicts. Minor logical conflicts: reduce the scope of adjustment instructions; minor resource conflicts: execute instructions from the same device in two batches with an interval of 1-2 seconds, with no resource overlap; minor cross-device linkage conflicts: merge similar instructions; delay the timing of non-core instructions. S2132. Iterative replacement of candidate sets for some paths: Based on L1, R1, and C1, select each effective path that needs to be iterated, and iteratively replace its initial optimal candidate set with the second smallest candidate set of Q for the corresponding path. Reconstruct the global instruction pool and recalculate the G value until a global instruction pool with global G < 0.3 is obtained.
4. The proactive security defense system based on reinforcement learning and attack intent inference according to claim 1, characterized in that, The sequential defense execution rule is as follows: S2021. Sort all medium and low risk paths in descending order of cumulative success probability. Use the sorting result as the execution order. Only retain candidate instruction sets whose attack behavior is consistent with that of the effective path. Then, based on the conflict rule base, detect conflicts within the candidate instruction sets, obtain L, R, and C in each candidate instruction set, calculate the comprehensive conflict score of each candidate instruction set, and remove candidate instruction sets with a comprehensive conflict score ≥ 0.
6. S2022. If there are candidate instruction sets with a comprehensive conflict score < 0.3, select the two smallest ones as the initial optimal candidate set, with the smallest one as the main selection set and the second smallest one as the alternative set. If the comprehensive conflict scores are all within [0.3, 0.6], select the smallest one as the main selection set. Before executing defense on each path, collect the full execution status of the previous executed path, and then perform pre-order adaptation optimization on the main selection set to obtain the proactive security defense instruction set.
5. The proactive security defense system based on reinforcement learning and attack intent inference according to claim 1, characterized in that, It also includes a conflict rule base building module, the building process of which is as follows: S4001 replicates all devices on the intranet, builds a 1:1 conflict test environment, sets up several defense scenarios, each defense scenario includes multiple high-risk paths executed globally in parallel or multiple medium- and low-risk paths executed sequentially. Among them, each high-risk path and each medium- and low-risk path in each defense scenario are not completely the same, and obtains several defense instruction sets for each type of conflict in each defense scenario from the intranet. S4002. Simulate each defense scenario in the conflict test environment, and then issue the preset defense command sets in each defense scenario through the command issuing tool in sequence. Detect whether the conflict is triggered in real time through the conflict monitoring tool. After the conflict is triggered, collect the core test data. S4003. Extract logical conflict data, resource conflict data, and cross-device linkage conflict data from the core test data, quantify them into logical conflict scores, resource conflict scores, and cross-device linkage conflict scores in each conflict test group, and then store them in association with the corresponding defense scenarios and defense instruction set features, thereby forming a conflict rule base.
6. The proactive security defense system based on reinforcement learning and attack intent inference according to claim 1, characterized in that, The specific process of the security defense feedback is as follows: Real-time monitoring of the status data of the execution device, logical conflict data in the intranet, asset conflict data, and cross-device linkage conflict data; based on the status data of the execution device, judge the defense effect during the execution process; and use the logical conflict data in the intranet, asset conflict data, and cross-device linkage conflict data to analyze the command conflict status. When the defense effect outputs an error, it is immediately fed back to the defense analysis module, which regenerates the proactive security defense instruction set and reissues it for execution. If the new instruction set is invalid, the manual intervention process is triggered, and the operation and maintenance department is notified to handle the emergency. When an error is output in the command conflict state, the preset rollback command is immediately executed to restore the device configuration, and feedback is sent to the defense analysis module to regenerate the proactive security defense command set and re-execute it; if the rollback fails or the command conflict state is still output as an error after re-execution, the defense of the corresponding path is suspended and the operation and maintenance are notified for emergency handling.
7. The method executed using the proactive security defense system based on reinforcement learning and attack intent inference as described in any one of claims 1-6, characterized in that, include: S1. Connect to all existing devices on the intranet, collect and standardize multi-source data, and then generate a dynamic attack graph. The dynamic attack graph includes several effective paths, the cumulative success probability of each effective path, the path weight, and path information. S2. Retrieve defense records from the database, filter out several candidate instruction sets for each effective path, analyze the conflict of each candidate instruction set, and then optimize the instructions. The optimization result serves as the active security defense instruction set for each effective path. S3. Distribute the active security defense instruction set of each valid path to the execution device in the corresponding path, execute the active security defense, evaluate the defense effect and instruction conflict status during execution, and provide security defense feedback when the defense effect or instruction conflict status outputs an error.