An APT attack and defense data generation method and system based on intelligent game confrontation

CN122802275APending Publication Date: 2026-09-22北京中关村实验室
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611272856.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-21
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0006]本发明的目的在于提出一种基于智能博弈对抗的APT攻防数据生成方法及系统,以解决现有APT攻防数据真实性不足、全过程覆盖不完整及难以表征攻防策略动态演进的技术问题,达到生成与目标环境、攻击链阶段时序和攻防反馈一致且可版本化、可复现、可扩展的APT攻防全过程数据集的效果

Benefits of technology

[0023]1.本发明基于可配置阶段数的半马尔可夫攻防博弈配置攻击链阶段,以目标网络拓扑、关键资产、漏洞集合和攻防动作作为环境输入,并为不同攻击链阶段分别配置相位型阶段驻留时间分布,能够表征长尾潜伏、快速渗透和阶段性停滞的攻击时间剖面,使生成的攻防轨迹与具体目标环境、攻击链阶段推进过程和驻留时序相对应,改善现有APT攻防数据真实性不足的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122802275A_ABST
    Figure CN122802275A_ABST
Patent Text Reader

Abstract

The application discloses an APT attack and defense data generation method and system based on intelligent game confrontation, and belongs to the technical field of network security. In order to solve the technical problems of insufficient authenticity of existing APT attack and defense data, incomplete whole-process coverage and difficulty in representing dynamic evolution of attack and defense strategies, the application configures an attack and defense game environment, arranges a multi-stage attack chain script, drives an attack intelligent agent and a defense intelligent agent to perform multiple rounds of attack and defense games and update attack and defense strategies, records whole-process trajectory data according to a seven-tuple composed of an attack stage, environment observation, attack action, defense action, attack reward, defense reward and stage time sequence, and obtains an APT attack and defense data set through quality checking and versioning organization. The application can generate APT whole-process data which is traceable, reproducible and has dynamic confrontation, and is used for APT detection, defense intelligent agent training and network security large model training and evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of network security technology, specifically relating to a method and system for generating APT (Advanced Persistent Threat) attack and defense data based on intelligent game-based adversarial competition. Background Technology

[0002] APTs (Aggressive Persistent Threats) are among the most severe security threats facing critical information infrastructure today. Unlike traditional attacks, APT attackers possess ample resources, a clear target orientation, and high stealth capabilities, enabling them to penetrate target systems in stages over timescales of months or even years. Based on the Cyber ​​Kill Chain and MITRE ATT&CK framework, researchers typically characterize APTs as a continuous process consisting of several attack phases, actions within each phase, dwell time, and defense feedback. Research on APT detection and defense methods heavily relies on high-quality attack and defense datasets. Current mainstream network attack and defense datasets, such as CTU-13, UNSW-NB15, CIC-IDS2017, CIC-APT, and DARPA TC, share common limitations in the following three dimensions: (1) Insufficient authenticity. Existing datasets are mostly derived from offline collection, scripted generation, or fixed scene replay. Once released, they remain unchanged and cannot reflect the strategic changes under the continuous evolution of technologies such as 0-day exploitation, new C2 channels, and LLM-assisted attacks. Their attack time profiles, dwell times, and defense feedback are often simplified, which makes it easy for detection or defense models to overfit historical samples and have insufficient transfer and generalization ability to real APTs.

[0003] (2) Insufficient comprehensiveness. Existing datasets often only record single attack-side traffic, malicious labels, or attack types, lacking key fields such as defense-side observations, decision-making actions, rewards for both attackers and defenders, phase sequence, and strategy versions. At the same time, most public datasets focus on a single or a few phases, lacking continuous trajectory collection of the complete attack chain. Therefore, downstream tasks find it difficult to conduct full-process detection, full-process defense strategy training, and end-to-end network linkage evaluation.

[0004] (3) Insufficient adversarial nature. Existing data augmentation methods mainly include adversarial sample generation based on GANs, text or payload synthesis based on LLMs, or data collection based on network test ranges and scripted execution tools. These methods usually lack continuous feedback, strategy updates, and cross-adversarial interaction between attackers and defenders, and the generated data cannot fully reflect the key characteristics of mutual evolution and adaptation of strategies between the two sides in the attack-defense game.

[0005] Therefore, there is an urgent need to construct a method and system: on the one hand, by deploying attack and defense agents with learning capabilities to conduct game-based adversarial activities, the data can reflect the real strategic attack and defense process; on the other hand, by using generalizable multi-stage game modeling to cover attack stages of different numbers and granularities, and adopting a five-stage APT attack chain as the engineering orchestration in the system design; at the same time, by uniformly collecting bilateral actions, rewards, observations, stage time sequences and policy versions, a high-quality baseline dataset can be formed that can be used for APT detection algorithm training, defense agent reinforcement learning training, and attack-defense adversarial robustness evaluation. Summary of the Invention

[0006] The purpose of this invention is to propose an APT attack and defense data generation method and system based on intelligent game-theoretic adversarial approach, in order to solve the technical problems of insufficient authenticity, incomplete full-process coverage, and difficulty in representing the dynamic evolution of attack and defense strategies in existing APT attack and defense data. The goal is to generate an APT attack and defense full-process dataset that is consistent with the target environment, attack chain stage timing, and attack and defense feedback, and is versionable, reproducible, and scalable.

[0007] To achieve the above objectives, the present invention adopts the following technical solution.

[0008] A method for generating APT attack and defense data based on intelligent game theory includes the following steps: Based on target environment information, attack behavior profile, phase dwell time distribution and attack and defense actions, configure attack and defense tasks, and arrange multi-stage attack chain scripts to obtain the attack and defense game environment and phased tasks. Based on the aforementioned attack and defense game environment and phased tasks, the attacking agent executes attack actions, and the defending agent detects, deceives, and blocks responses. The attack and defense strategy is updated according to the attack and defense results, resulting in a multi-round attack and defense game process. The multi-round attack and defense game process is recorded and labeled according to a seven-tuple consisting of attack phase, environmental observation, attack action, defense action, attack reward, defense reward and phase sequence to obtain the full process trajectory data; The entire process trajectory data is subjected to quality verification, and the attack and defense tasks, attack and defense strategies or abnormal trajectories are adjusted or filtered according to the verification results to obtain trajectory data that passes the quality verification. The trajectory data that has passed quality verification is versioned, organized, and packaged to obtain the APT attack and defense dataset.

[0009] Furthermore, based on target environment information, attack behavior profiles, phase dwell time distribution, and attack and defense actions, attack and defense tasks are configured, and multi-stage attack chain scripts are arranged to obtain the attack and defense game environment and phased tasks, including: Based on the target environment information and attack and defense actions, configure the network topology, key assets, vulnerability information, and attack and defense action space to obtain environmental constraint information; Based on the attack behavior profile and the stage dwell time distribution, the attack chain stages, the dwell rules of each attack chain stage, and the stage transition rules are determined to obtain the attack chain configuration information. Based on the environmental constraints and attack chain configuration information, attack and defense tasks are configured, and multi-stage attack chain scripts are arranged according to the attack chain stages to obtain the attack and defense game environment and staged tasks.

[0010] Furthermore, based on the attack behavior profile and stage dwell time distribution, the attack chain stages, dwell rules for each attack chain stage, and stage transition rules are determined to obtain attack chain configuration information, including: Determine the attack chain phases and the behavioral configurations for each attack chain phase based on the attack behavior profile; The dwell rules for each attack chain stage are determined based on the dwell time distribution of the phase-type stage corresponding to each attack chain stage. Based on the offensive and defensive actions, the behavior configuration, and the dwell rules, the phase transition rules are determined, and the attack chain configuration information is obtained.

[0011] Furthermore, attack and defense tasks are configured based on the environmental constraint information and the attack chain configuration information, and multi-stage attack chain scripts are arranged according to the attack chain stages to obtain the attack and defense game environment and staged tasks, including: The environmental constraint information and the attack chain configuration information are loaded into the game environment to obtain the attack and defense game environment. Based on the attack chain configuration information, determine the stage objectives, execution order, completion conditions, and advancement rules of each attack chain stage, and arrange multi-stage attack chain scripts to obtain staged tasks.

[0012] Furthermore, based on the aforementioned attack-defense game environment and phased tasks, an attacking agent executes attack actions, while a defending agent detects, deceives, and blocks responses. The attack-defense strategy is updated based on the attack-defense results, resulting in a multi-round attack-defense game process, including: The attacking agent generates attack actions based on the phased tasks, the current attack phase, and defense feedback. The defensive agent determines the attack phase confidence distribution based on environmental observations and action history, and determines detection actions, deception actions, and blocking actions based on the attack phase confidence distribution to obtain defensive actions; In the attack and defense game environment, the attack action and the defense action are executed. The phase transition, residence time, attack reward and defense reward are determined according to the phase residence time distribution to obtain the attack and defense result. The attack strategy model and defense strategy model are trained based on the attack and defense results to obtain updated attack and defense strategies. The attack and defense game is then repeatedly executed based on the updated attack and defense strategies to obtain a multi-round attack and defense game process.

[0013] Furthermore, the defensive agent determines the attack phase confidence distribution based on environmental observations and action history, and determines detection actions, deception actions, and blocking actions based on the attack phase confidence distribution to obtain defensive actions, including: By temporally correlating environmental observations and action history, stage observation characteristics can be obtained; Update the attack phase confidence distribution based on the observed phase characteristics; Based on the confidence distribution of the attack phase, detection actions, deception actions, and blocking actions are determined respectively, and the detection actions, deception actions, and blocking actions are combined into defensive actions.

[0014] Furthermore, the attack strategy model and defense strategy model are trained based on the attack and defense results to obtain updated attack and defense strategies. The attack and defense game is then repeatedly executed based on these updated strategies to obtain a multi-round attack and defense game process, including: The multiple attack and defense results are aggregated to obtain a batch of attack and defense results; Based on the batches of attack and defense results, the cumulative reward, stage coverage and strategy output distribution are determined, and when the strategy update conditions are met, the attack strategy model and defense strategy model are trained by multi-agent reinforcement learning to obtain the updated attack strategy and defense strategy. Generate policy version identifiers and policy generations for the updated attack and defense strategies, and store the corresponding policy versions and historical policy versions in the policy pool; The sampling mode is determined from the training sampling mode, the evolution sampling mode and the steady-state sampling mode, and the latest policy version, the historical policy version or a combination of different policy versions are selected from the policy pool according to the sampling mode to obtain candidate attack policies and candidate defense policies. When the strategy output distribution meets the degradation condition or the stage coverage decreases, the candidate attack strategy or candidate defense strategy is replaced with a historical strategy version that meets the stability condition, and the strategy exploration rate is increased. The attack and defense game is repeatedly executed based on the candidate attack strategy and candidate defense strategy to obtain a multi-round attack and defense game process.

[0015] Furthermore, the multi-round attack-defense game process is recorded and labeled according to a seven-tuple consisting of attack phase, environmental observation, attack action, defense action, attack reward, defense reward, and phase sequence, to obtain the full-process trajectory data, including: Before selecting attack and defense actions at each time step, record the attack phase and environmental observations. After simultaneously selecting attack and defense actions at each time step, the attack and defense actions are recorded. After the phase transition is completed in the attack and defense game environment, the attack reward, defense reward and phase sequence are recorded, and the attack phase, environment observation, attack action, defense action, attack reward, defense reward and phase sequence are combined into a time step seven-tuple. By associating the seven-tuples of each time step in chronological order and then labeling the trajectory of each associated seven-tuple, the trajectory data of the entire process can be obtained.

[0016] Furthermore, the seven-tuples of each time step are associated in chronological order, and trajectory annotations are performed on the associated seven-tuples of each time step to obtain the trajectory data of the entire process, including: Obtain attack behavior profile parameters, attack and defense strategy generation, strategy version identifier, random seed and field masking rules, and generate trajectory metadata; The seven-tuples of each time step are associated with the trajectory identifier and time sequence, and the associated seven-tuples of each time step are stored in correspondence with the trajectory metadata to obtain the labeled trajectory data. The publication visibility of each field in the labeled trajectory data is marked according to the field masking rules to obtain the full-process trajectory data.

[0017] Furthermore, the entire trajectory data undergoes quality verification, and the attack and defense tasks, strategies, or abnormal trajectories are adjusted based on the verification results to obtain trajectory data that passes the quality verification, including: Based on the full-process trajectory data, determine the stage coverage rate, the consistency index of reward decay, the trajectory diversity index, and the abnormal trajectory status; The stage coverage, reward decay consistency index, trajectory diversity index, and abnormal trajectory status are compared with the corresponding quality verification conditions to obtain the quality verification results. Based on the quality verification results, adjust the attack and defense tasks and strategies or filter abnormal trajectories, and supplement the entire process trajectory data based on the adjusted attack and defense tasks or strategies to obtain trajectory data that has passed the quality verification.

[0018] Furthermore, based on the quality verification results, the attack and defense tasks, attack and defense strategies, or abnormal trajectories are adjusted, and the entire process trajectory data is supplemented based on the adjusted attack and defense tasks or strategies to obtain trajectory data that has passed the quality verification, including: The type of failure to pass the verification is determined based on the quality verification results; When the failure verification type is insufficient stage coverage, the strategy exploration rate of the corresponding attack chain stage is increased; When the failure type is inconsistent reward decay, adjust the attack behavior profile, reward parameters, or attack and defense strategies. When the failure type is insufficient trajectory diversity, historical attack and defense strategies are selected for attack and defense game. Abnormal trajectories are filtered out, and full-process trajectory data is supplemented based on the adjusted attack and defense tasks, attack and defense strategies, or the historical attack and defense strategies. The quality of the supplemented full-process trajectory data is then re-verified to obtain trajectory data that passes the quality verification.

[0019] Furthermore, the trajectory data that has passed quality verification is versioned, organized, and packaged for release to obtain an APT attack and defense dataset, including: The attack and defense strategy hierarchy is obtained, and the trajectory data that has passed the quality verification is organized according to the attack behavior profile and the attack and defense strategy hierarchy to obtain the trajectory data packet to be released. The field masking rules are determined based on the published task, and the fields in the trajectory data packet to be published are masked based on the field masking rules to obtain a strongly supervised data view, a semi-supervised data view, and an observation data view. The strongly supervised data view, semi-supervised data view, and observation data view are associated with the attack and defense task configuration, stage dwell time configuration, quality verification records, and reproduction information to generate a dataset version and release and encapsulate it to obtain the APT attack and defense dataset.

[0020] Further, field masking rules are determined based on the publishing task, and fields in the trajectory data packet to be published are masked based on the field masking rules to obtain a strongly supervised data view, a semi-supervised data view, and an observation data view, including: Retain all fields of the seven-tuple in the trajectory data packet to be published, and generate a strongly supervised data view; The attack phase of a portion of the trajectory in the trajectory data packet to be published is masked to generate a semi-supervised data view; The attack phase, attack action, defense action, attack reward, defense reward, and phase sequence in the trajectory data packet to be published are masked to generate an observation data view.

[0021] An APT attack and defense data generation system based on intelligent game-theoretic confrontation includes an attack and defense task preparation module, a director scheduling module, an attack execution module, a defense response module, a Cangjie data module, a data quality verification module, and a data asset management module. The attack and defense task preparation module is used to configure attack and defense tasks based on target environment information, attack behavior profile and stage dwell time distribution to obtain the attack and defense game environment. The director scheduling module is used to schedule multi-stage attack chain scripts according to the attack and defense tasks to obtain staged tasks. The attack execution module and the defense response module are used to execute multiple rounds of attack and defense games based on the phased tasks in the attack and defense game environment, advance the attack chain stages according to the phase dwell time distribution, and train the attack strategy model and defense strategy model according to the attack and defense results to obtain the attack and defense game process. The Cangjie data module is used to record and label the attack and defense game process according to a seven-tuple including stage information, environmental observation, attack action, defense action, attack reward, defense reward and stage sequence, to obtain the trajectory data of the whole process; The data quality verification module is used to verify the quality of the entire process trajectory data, filter out valid trajectory data, and generate quality feedback based on the verification results for adjusting attack and defense tasks or updating attack and defense strategies. The data asset management module is used to version the valid trajectory data to obtain an APT attack and defense dataset.

[0022] The present invention has achieved the following beneficial effects.

[0023] 1. This invention configures attack chain stages based on a configurable number of stages in a semi-Markov attack-defense game. It uses the target network topology, key assets, vulnerability sets, and attack and defense actions as environmental inputs, and configures phase-type stage dwell time distributions for different attack chain stages. This can characterize the attack time profiles of long-tail infiltration, rapid penetration, and staged stagnation, so that the generated attack and defense trajectories correspond to the specific target environment, the attack chain stage progression process, and the dwell time sequence, thus improving the problem of insufficient authenticity of existing APT attack and defense data.

[0024] 2. This invention employs trainable attack and defense agents to execute multiple rounds of attack and defense games, and uses multi-agent reinforcement learning to update attack and defense strategies. Trajectory collection is organized through training sampling mode, evolution sampling mode, and steady-state sampling mode. Different generations of attack and defense strategy versions are stored in a policy pool. When strategy collapses or stage coverage decreases, strategy rollback, exploration rate improvement, or introduction of historical opponents are performed, so that the generated data can reflect the dynamic adversarial relationship of mutual adaptation and continuous evolution between the attacker and defender.

[0025] 3. This invention synchronously records the attack and defense game process according to a seven-tuple consisting of attack phase, environmental observation, attack action, defense action, attack reward, defense reward, and phase sequence, and associates attack behavior profile parameters, agent generation, policy version hash, random seed, and field masking rules, so that each trajectory simultaneously covers the attack side, defense side, environmental state, phase sequence, and reproducibility information, thus improving the problem of incomplete labeling of the entire process in existing datasets.

[0026] 4. Based on field masking rules, this invention generates strongly supervised data views, semi-supervised data views, and observation data views for the same batch of full-process trajectory data, enabling the underlying trajectory to adapt to data requirements with different field visibility and annotation granularity. Furthermore, the trajectory generation conditions are preserved through strategy version, random seed, and metadata, facilitating versioned management, process traceability, and experiment reproduction of attack and defense data.

[0027] 5. This invention filters valid trajectory data through phase coverage verification, reward decay consistency verification, trajectory diversity verification, and abnormal trajectory verification. Based on the verification results, it adjusts attack and defense tasks, attack behavior profiles, attack and defense strategies, or sampling methods to ensure that the published APT attack and defense dataset is consistent with environmental constraints, phase time series, and reward feedback, thereby reducing the impact of single-phase trajectories, strategy degradation trajectories, and abnormal trajectories on data quality.

[0028] 6. This invention categorizes defensive actions into three types: detection, non-intrusive deception, and perceptible blocking. It also ensures that each type of defensive action covers three operational levels: endpoint-side, network-side, and endpoint-network linkage. This creates a defensive action space that is distinct in terms of attacker perception and scope of action. The generated data simultaneously includes endpoint-side detection responses, network-side detection responses, and endpoint-network collaborative processing data. This data can be used for APT detection, defensive agent training, endpoint-network linkage effect evaluation, attack and defense strategy review, and large-scale network security model training and evaluation. Attached Figure Description

[0029] Figure 1 This is a flowchart of the APT attack and defense data generation method based on intelligent game-theoretic confrontation in the embodiment. Figure 2 This is a schematic diagram illustrating the state transitions of the evolutionary game scheduling and the three sampling modes in the embodiment. Figure 3 This is a system architecture diagram of an APT attack and defense data generation system based on intelligent game theory in the embodiment. Detailed Implementation

[0030] To make the various technical features, advantages, or effects of the present invention more apparent and understandable, detailed descriptions are provided below through embodiments.

[0031] This invention provides a method for generating APT attack and defense data based on intelligent game adversarial competition. The method is based on a semi-Markov attack and defense game with configurable number of stages. In a configurable attack and defense game environment, attack agents and defense agents perform multiple rounds of attack and defense games. The attack and defense strategies are updated according to the attack and defense results. The attack and defense game process is recorded according to a seven-tuple consisting of attack stage, environmental observation, attack action, defense action, attack reward, defense reward, and stage sequence. After quality verification, versioning organization, and release encapsulation, the APT attack and defense dataset is obtained.

[0032] Before running this method, target environment information, attack behavior profile, phase dwell time distribution, attack and defense actions, and reward parameters are configured, and a configurable multi-stage attack chain is set. The attack process state set is represented as follows: in, Represents the set of states during the attack process; This indicates the configurable number of attack chain stages, and , ; Indicates either the cleared state or the initial safe state; Indicates the target is in a trapped state; Indicates the first Each attack chain stage, and .

[0033] The number of attack chain stages, stage names, and attack chain granularity can all be configured according to the target scenario. For example, the number of attack chain stages can be configured to 3, 4, 5, 6, or more to adapt to CKC, ATT&CK, industry-specific attack chains, or enterprise-customized attack flows. In this embodiment, the default is adopted. The CKC five-stage orchestration consists of five attack chain stages: reconnaissance, foothold, command and control, lateral movement, and data leakage, with two absorption states: a cleanup state and a target compromise state. This five-stage orchestration is used throughout the configuration of offensive and defensive missions, the phased execution of missions, and the organization of trajectory data fields.

[0034] At time step The attacking agent and the defending agent select actions based on the attack strategy and defense strategy, respectively, which are expressed as follows: in, Indicate the attack strategy; Indicates time step The attack action; This indicates that the attacking agent is at time step Historical information that can be obtained; Indicates a defensive strategy; Indicates time step Defensive actions; Indicates time step Environmental observation; Indicates time step The confidence distribution of the attack phase.

[0035] The confidence distribution during the attack phase satisfies: in, Represents the set of attack process states The probability simplex on the attack phase represents the posterior probability of the defensive agent judging that the attack process is in the corresponding attack chain phase.

[0036] The offensive and defensive game environment generates the next attack chain stage, dwell time, attack reward, and defense reward based on the current attack chain stage, attack action, defense action, and stage dwell time distribution, forming a semi-Markov transition relationship: in, This represents the semi-Markov transition probability; Indicates time step The attack chain phase; Indicates time step The attack chain phase; Indicates time step Corresponding length of stay; Indicates time step Instant rewards for attacks; Indicates time step Instant rewards for defense; Indicates time step The attack action; Indicates time step Defensive actions.

[0037] For any stage of the attack chain The corresponding dwell time is parameterized according to the phase distribution: in, Indicates the first The dwell time of each stage of the attack chain; Indicates a phase-type distribution; Indicates the first Initial phase probability vectors for each stage of the attack chain; Indicates the first The sub-random matrix of each attack chain stage. The phase-type distribution can approximate the distribution of non-negative random variables with arbitrary precision, and the time profile of different attack chain stages can be represented by Erlang distribution, Coxian distribution, super-exponential distribution or mixed phase-type distribution.

[0038] The main process of this method is as follows: Figure 1 As shown, the specific implementation involves the following steps.

[0039] Step S1: Configure attack and defense tasks based on target environment information, attack behavior profile, phase dwell time distribution and attack and defense actions, and arrange multi-stage attack chain scripts to obtain the attack and defense game environment and phased tasks.

[0040] Specifically, target environment information, attack behavior profile, stage dwell time distribution, attack and defense actions, reward parameters and attack chain stages are configured through a structured configuration file in YAML or JSON format. The configuration data is then loaded into the attack and defense game environment to complete the attack and defense task initialization.

[0041] In an optional embodiment of the present invention, step S1 may include: Step S11: Based on the target environment information and attack and defense actions, configure the network topology, key assets, vulnerability information, and attack and defense action space to obtain environmental constraint information.

[0042] Specifically, the target environment information includes the target system's network topology, critical assets, and vulnerability information. The network topology includes the number of nodes, the link relationships between nodes, and subnetting; critical assets include databases, file servers, domain controllers, and endpoints; and vulnerability information includes CVE numbers and vulnerability severity.

[0043] Allocate importance weights to key assets The importance weight of key assets is used to determine the stage advancement benefits generated by the corresponding asset during the advancement of the attack chain. .

[0044] The attack action space is represented as follows: in, This represents the attack action space. Each attack action can be configured with corresponding technical sub-prototypes. For example, exploit-type attack actions can include vulnerability exploitation, credential theft, and social engineering attack sub-prototypes.

[0045] The defensive action space is represented as follows: in, Indicates the space for defensive actions; Represents the set of detection actions; This represents a set of deceptive actions that are imperceptible to the user. This represents a set of actions that can be felt as blocking the flow of energy.

[0046] Defense actions are configured according to action category, level of action, and intensity level. Action categories include detection, non-intrusive deception, and sensory blocking; level of action includes end-side, network-side, and end-network coordination; intensity levels include light, medium, and heavy, thus forming... There are 27 defensive action prototypes, numbered D-001 to D-027, as shown in Table 1. During deployment, these 27 defensive action prototypes can be subsetted or expanded based on the target system's capability list.

[0047] Table 1. Detailed Table of Defensive Action Prototypes This table lists 27 specific action prototypes of the present invention within the three-dimensional action space of "detection × imperceptible deception × sensory blocking". The three dimensions are: action category (3 items) × level of action (end-side / network-side / end-network linkage, 3 items) × intensity level (light / medium / heavy). The final action of the defensive agent... This can be represented as a probability distribution over these 27 prototypes. The three-dimensional decomposition guarantees: first, that the category dimensions are mutually exclusive (detection / imperceptible deception / perceptible blocking are orthogonal in terms of attacker perceptibility); and second, that the intensity dimension is ordered (light < medium < heavy, monotonic in terms of resource consumption and business impact). This structure helps reduce the sample complexity of policy learning.

[0048] Step S12: Based on the attack behavior profile and stage dwell time distribution, determine the attack chain stages, the dwell rules of each attack chain stage, and the stage transition rules to obtain the attack chain configuration information.

[0049] In an optional embodiment of the present invention, step S12 may include: Step S121: Determine the attack chain stages and the behavior configuration of each attack chain stage based on the attack behavior profile.

[0050] Specifically, each attack chain phase is configured with executable attack behavior sub-prototypes and events that can trigger phase transitions. In the default CKC five-phase orchestration, Indicates the reconnaissance phase. Indicates the initial stage. This indicates the command and control phase. Indicates the lateral movement phase. This indicates the stage of data leakage.

[0051] The reconnaissance phase can be configured with scanning and target information collection behaviors; the foothold phase can be configured with vulnerability exploitation, credential theft, privilege acquisition, and residency establishment behaviors; the command and control phase can be configured with command and control communication and residency maintenance behaviors; the lateral movement phase can be configured with credential usage, remote access, and cross-node movement behaviors; and the data leakage phase can be configured with data collection and data transmission behaviors.

[0052] Attack behavior profiles are used to configure the dwell characteristics and progression rhythm of different attack chain stages. In this embodiment, the following three attack behavior profiles can be configured.

[0053] (1) Heavy reconnaissance attack behavior profile, with long tail persistence during the reconnaissance phase, using Coxian-3 heavy tail configuration, typical APTs include APT29 or Cozy Bear.

[0054] (2) Rapid penetration attack behavior profile, each stage of the attack chain progresses rapidly, and a full-stage Erlang-2 configuration is adopted. A typical comparison APT includes the Lazarus financial targeted attack.

[0055] (3) Long-latent attack behavior profile: The command and control phase and the lateral movement phase exhibit long-tail persistence. It adopts a hybrid Erlang configuration. Typical APTs include APT1 or Comment Crew.

[0056] The phase distribution configurations of the above three attack behavior profiles are shown in Table 2.

[0057] Table 2. pH distribution configuration of three typical APT behavior profiles In addition to the attack behavior profiles mentioned above, other attack behavior profiles can be configured according to the target scenario.

[0058] Step S122: Determine the dwell rules for each attack chain stage based on the dwell time distribution of the phase-type stage corresponding to each attack chain stage.

[0059] Specifically, an initial phase probability vector is configured for each stage of the attack chain. and sub-random matrices ,according to and Determine the dwell time distribution for the corresponding attack chain stages. Different phase distributions can be configured for different attack chain stages to represent attack time profiles of long-tailed lurking, rapid advancement, and phased stagnation.

[0060] Step S123: Determine the phase transition rules based on attack and defense actions, behavior configurations, and dwell rules to obtain attack chain configuration information.

[0061] Specifically, the phase transition rules are used to determine the next attack chain phase based on the current attack chain phase, attack actions, defense actions, attack behavior configuration, and the dwell rules of the current attack chain phase. Each attack chain phase is configured with transition events that can trigger phase advancement, phase dwell, phase rollback, entry into a clearing state, or entry into a target compromised state.

[0062] Step S13: Configure attack and defense tasks based on environmental constraint information and attack chain configuration information, and arrange multi-stage attack chain scripts according to the attack chain stages to obtain the attack and defense game environment and staged tasks.

[0063] In an optional embodiment of the present invention, step S13 may include: Step S131: Load the environmental constraint information and attack chain configuration information into the game environment to obtain the attack and defense game environment.

[0064] Specifically, network topology, key assets, vulnerability information, attack and defense action space, attack chain stages, dwell rules, stage transition rules, and reward parameters are injected into the game environment kernel, enabling the game environment kernel to perform semi-Markov state transitions based on the current attack chain stage, attack action, and defense action.

[0065] Attack and defense rewards are determined based on the phase advancement gains, attack costs, deception losses, defense costs, and business losses. The weights for phase advancement gains, attack costs, deception losses, defense costs, and business losses are respectively expressed as follows: , , , and In this embodiment, , , , and The default value is 0.2. Other embodiments can be configured according to the target scenario.

[0066] Step S132: Determine the stage objectives, execution order, completion conditions and advancement rules of each attack chain stage based on the attack chain configuration information, and arrange multi-stage attack chain scripts to obtain staged tasks.

[0067] Specifically, the phased tasks include the phase objectives, executable actions, execution order, success conditions, failure conditions, phase progression rules, and review rules corresponding to each phase of the attack chain. When an attack action meets the success conditions of the current attack chain phase, it proceeds to the next attack chain phase; when an attack action is detected, deceived, or blocked and meets the failure conditions, attack strategy feedback or defense strategy feedback is generated based on the reason for the failure.

[0068] Step S2: Based on the attack and defense game environment and phased tasks, the attacking agent executes attack actions, and the defending agent detects, deceives, and blocks responses. The attack and defense strategy is updated according to the attack and defense results, resulting in a multi-round attack and defense game process.

[0069] Specifically, instantiating attack agents and defensive intelligent agents and respectively attacking intelligent agents and defensive intelligent agents Configure the policy network (actor) and the value network (critic). The policy network generates the action probability distribution, and the value network evaluates the state value under the corresponding policy. Attack agent. and defensive intelligent agents Trainable adversarial training takes place in an offensive and defensive game environment.

[0070] In an optional embodiment of the present invention, step S2 may include: Step S21: The attacking agent generates an attack action based on the phased task, the current attack phase, and the defense feedback.

[0071] Specifically, attack strategies At the current stage of the attack chain and some observable defense history The input is the probability distribution of attack actions, and the output is the attack action probability distribution. The attacking agent can fully observe its current stage in the attack chain, but its observation of defensive actions is limited to the defensive feedback returned by the attack-defense game environment. The defensive feedback includes honeypot responses and delayed signals of defensive measures being triggered.

[0072] The attacking agent selects attack actions based on the probability distribution of attack actions. attack actions This includes attack action categories and corresponding technical sub-prototypes.

[0073] In step S22, the defensive agent determines the attack phase confidence distribution based on environmental observations and action history, and determines the detection action, deception action, and blocking action based on the attack phase confidence distribution to obtain the defensive action.

[0074] In an optional embodiment of the present invention, step S22 may include: Step S221: Perform time-series correlation between environmental observations and action history to obtain stage observation characteristics.

[0075] Specifically, environmental observation includes traffic, logs, endpoint behavior, alarm events, service status, and visible traces on the attack side. The partially observable sequences obtainable by the defensive agent are represented as follows: in, Indicates time step Environmental observation; Indicates an alarm event; Represents log data; This indicates the data on the decoy's hit rate.

[0076] The defensive agent performs temporal correlation on current environmental observations, historical environmental observations, and historical defensive actions to obtain stage observation features used to determine the stage of the attack chain.

[0077] Step S222: Update the attack phase confidence distribution based on the phase observation characteristics.

[0078] Specifically, the confidence distribution during the attack phase The corresponding part of the observable Markov decision process POMDP is the belief state, which represents the posterior probability distribution of the defensive agent for the current stage of the attack chain.

[0079] In one alternative implementation, a Bayesian filter is employed to perform precise Bayesian filtering based on stage observation features and action history, thereby updating the attack stage confidence distribution. .

[0080] In another alternative implementation, a recurrent neural network is used to encode the phase observation features and action history, and the attack phase confidence distribution is approximately updated based on the encoding results. .

[0081] Step S223: Determine the detection action, deception action, and blocking action according to the confidence distribution of the attack phase, and combine the detection action, deception action, and blocking action into a defense action.

[0082] Specifically, defense strategies It includes three parallel policy sub-networks, namely the detection dimension policy. Seamless deception strategy and sensory blocking strategy The three policy subnetworks share a single observation encoder, which receives observation features from the attack phase and outputs the confidence distribution for the attack phase. The three strategy subnetworks are based on the confidence distribution of the attack phase. Output the probability distributions of detection actions, deception actions, and blocking actions, respectively. Defensive actions. This can be represented as the probability distribution over the 27 defensive action prototypes in Table 1.

[0083] Step S23: Execute attack and defense actions in the attack and defense game environment, determine the phase transition, dwell time, attack reward and defense reward based on the phase dwell time distribution, and obtain the attack and defense result.

[0084] Specifically, within the same time step, the attacking agent and the defending agent simultaneously select attack actions. and defensive actions The offensive and defensive game environment is based on the current stage of the attack chain. Attack actions Defensive actions And the distribution of stage dwell time determines the next attack chain stage. and length of stay And calculate the immediate reward for the attack. and defense instant rewards .

[0085] In the seven-tuple record and reward configuration, attack rewards and defense rewards can also be recorded as follows: and ,in, In the corresponding semi-Markov transition relation , In the corresponding semi-Markov transition relation .

[0086] Step S24: Train the attack strategy model and defense strategy model based on the attack and defense results to obtain updated attack and defense strategies, and repeatedly execute the attack and defense game based on the updated attack and defense strategies to obtain a multi-round attack and defense game process.

[0087] In an optional embodiment of the present invention, step S24 may include: Step S241: Collect multiple attack and defense results to obtain a batch of attack and defense results.

[0088] Specifically, each round of attack and defense includes Each time step is used to aggregate the phase transitions, dwell times, attack actions, defense actions, attack rewards, and defense rewards generated from one or more rounds of attack and defense game into a batch of attack and defense results. In this embodiment, The default value is 200.

[0089] Step S242: Determine the cumulative reward, stage coverage, and strategy output distribution based on the batch of attack and defense results. When the strategy update conditions are met, train the attack strategy model and defense strategy model through multi-agent reinforcement learning to obtain the updated attack strategy and defense strategy.

[0090] Specifically, in this embodiment, the Multi-Agent Proximity Policy Optimization (MAPPO) algorithm is used as the default training algorithm. In other embodiments, MADDPG, QMIX, or Nash-Q multi-agent reinforcement learning algorithms can be used to train the attack strategy model and the defense strategy model.

[0091] During training, a fixed attack strategy is employed. And update defense strategies Continue this process until the value network critic loss of the defense strategy model meets the convergence condition; then fix the defense strategy. And update attack strategies The above alternating training process is repeated cyclically. In this embodiment, one epoch. The default value is 500.

[0092] When training with MAPPO, the magnitude of a single policy update is limited by the clip term, and the generalized advantage estimation (GAE) is used to determine the advantage value used for policy updates.

[0093] To maintain the evolution of attack and defense strategies, within each epoch, probability is used... Inject perturbations into the policy output and periodically apply them to the parameters of the attack and defense policy models. Regular expressions. Among them, Indicates the probability of strategy exploration; This represents a regularization constraint based on the sum of squares of the model parameters.

[0094] Step S243: Generate policy version identifiers and policy generations for the updated attack and defense policies, and store the corresponding policy versions and historical policy versions in the policy pool.

[0095] Specifically, the policy pool stores attack policy versions and defense policy versions separately, and retains the most recent version. Each attack agent version and recent Each defense agent has a version. Each policy version is associated with a corresponding policy version hash, policy generation, and generation time. During sampling, adversary policies can be randomly selected from the policy pool according to time decay weights.

[0096] Step S244: Determine the sampling mode from the training sampling mode, the evolution sampling mode, and the steady-state sampling mode, and select the latest policy version, the historical policy version, or a combination of different policy versions from the policy pool according to the sampling mode to obtain candidate attack policies and candidate defense policies.

[0097] Specifically, sampling modes include training sampling mode, evolution sampling mode, and steady-state sampling mode, such as... Figure 2 As shown. When generating each attack and defense trajectory, attack and defense strategy versions are selected from the strategy pool, and the attack and defense game environment is reset. After the environment reset, the attack agent and the defense agent simultaneously select attack and defense actions. The attack and defense game environment performs PH transition based on the current attack chain stage, attack action, defense action, and phase-type stage dwell time distribution to determine the next attack chain stage and dwell time, and calculates the immediate attack reward and immediate defense reward.

[0098] Before entering the cleanup state or target compromise state during the attack-defense game, the attacking agent and the defending agent continue to synchronously select attack and defense actions based on the attack chain stage after the transition, and repeatedly execute PH transition, dwell time calculation, and real-time attack-defense reward calculation. When the attack-defense game enters the cleanup state or target compromise state, the current attack-defense trajectory ends, the seven-tuples corresponding to each time step in the current attack-defense trajectory are recorded, and the seven-tuples are appended to the trajectory cache. After the current attack-defense trajectory is recorded, the corresponding sampling mode is selected from the training sampling mode, evolution sampling mode, and steady-state sampling mode according to the convergence state of the attack strategy model and the defense strategy model and the preset sampling configuration.

[0099] (1) Training Sampling Mode. Before the attack strategy model and defense strategy model meet the convergence condition, the attack and defense game is continuously executed based on the attack strategy version and defense strategy version selected from the strategy pool, and the generated attack and defense trajectories are input into the strategy evolution controller. The strategy evolution controller trains the attack strategy model and defense strategy model according to the attack and defense trajectories, updates the trained attack strategy version and defense strategy version to the strategy pool, and outputs a batch of trajectories with training labels. Each attack and defense trajectory corresponds to the attack strategy version and defense strategy version during the training and evolution process.

[0100] (2) Evolutionary Sampling Mode. After generating M attack and defense trajectories, the corresponding attack and defense trajectories are input into the strategy evolution controller. The strategy evolution controller triggers incremental training of the attack strategy model or defense strategy model, and updates the attack strategy version or defense strategy version formed after incremental training to the strategy pool, so that the subsequent M attack and defense trajectories correspond to the updated strategy version. The generated data is organized according to the strategy generation. The strategy within the same strategy generation remains stable, while the strategy between different strategy generations evolves, and trajectory batches with evolution labels are output. In this embodiment, the evolutionary sampling mode is the default sampling mode.

[0101] (3) Steady-state sampling mode. After the attack strategy model and defense strategy model meet the convergence condition, the attack strategy and defense strategy are frozen to obtain the frozen strategy version, and the attack and defense game is performed based on the frozen strategy version. The frozen strategy version is not updated during the steady-state sampling process, and the trajectory batch with steady-state label is output.

[0102] In both training and evolution sampling modes, the policy evolution controller writes the updated policy version into the policy pool. Subsequent attack and defense trajectory generation then reselects the attack and defense policy versions from this updated pool. The trajectory batches generated in training, evolution, and steady-state sampling modes are configured with training, evolution, and steady-state labels, respectively, to identify the policy state and sampling method of the corresponding trajectory batch.

[0103] A single data generation task is managed using a ternary index consisting of policy version, trajectory batch, and sampling mode. A unique generation number is generated for each trajectory batch, and this generation number is written into the trajectory metadata as the policy generation.

[0104] Step S245: When the strategy output distribution meets the degradation condition or the stage coverage decreases, the candidate attack strategy or candidate defense strategy is replaced by the historical strategy version that meets the stability condition and the strategy exploration rate is increased. The attack and defense game is repeatedly executed based on the candidate attack strategy and candidate defense strategy to obtain a multi-round attack and defense game process.

[0105] Specifically, after each trajectory batch sampling is completed, the policy entropy, cumulative reward fluctuation, and stage coverage are checked. When the number of samples reaches a preset threshold and the attack-defense game still retains exploration space, incremental training of the attack strategy model or defense strategy model is triggered.

[0106] When the strategy entropy is close to 0, the strategy output is concentrated on a single action, or the phase coverage decreases, the candidate attack strategy or candidate defense strategy is rolled back to the previous historical strategy version that meets the stability conditions, and the strategy exploration rate is increased; historical attack strategy versions or historical defense strategy versions can also be introduced from the strategy pool for cross-counterattack.

[0107] Each attack and defense trajectory is generated according to the following process: sampling the initial state from the configured initial state distribution. The attacking and defending agents synchronously select actions based on the current attack and defense strategies; the attack-defense game environment determines the next attack chain stage and dwell time based on the phase-type distribution transition kernel; and the immediate reward for the attack is calculated. and defense instant rewards The attack process enters the cleanup state. Or target trapped state The current attack / defense trajectory ends upon completion; after the current attack / defense trajectory ends, the system returns to the cleared state according to the configured restart probability. Re-enter the offensive and defensive game and start the next round.

[0108] Step S3: Record and label the multi-round attack and defense game process according to a seven-tuple consisting of attack phase, environmental observation, attack action, defense action, attack reward, defense reward and phase sequence to obtain the trajectory data of the whole process.

[0109] Specifically, during the multi-round attack and defense game, each time step is recorded synchronously, and the recording results are stored in JSON or JSONL format. The attack phase field in the 7-tuple includes the actual attack phase and the attack phase confidence distribution. All types of data in the 7-tuple jointly record the state, observations, actions, rewards, and phase sequence of both the attacker and defender.

[0110] The data fields of the seven-tuple include true_stage, observation, belief, or The functions are: attacker_action, defender_action, attacker_reward, defender_reward, and phase_dwell, where true_stage and belief are either... Together, they constitute the attack phase information. The type, meaning, and visibility of this seven-tuple field are shown in Table 3.

[0111] Table 3. Definition of the 7-tuple data trajectory field In an optional embodiment of the present invention, step S3 may include: Step S31: Before selecting attack and defense actions at each time step, record the attack phase and environmental observations.

[0112] Specifically, the attack phase includes the real attack phase (true_stage) visible during the generation process and the attack phase confidence distribution (belief) within the defense intelligence. Environmental observation includes alerts for the current time step, log data, and decoy hit data (decoy_hits).

[0113] Step S32: After selecting the attack action and defense action synchronously at each time step, record the attack action and defense action.

[0114] Specifically, the attack action `attacker_action` records the attack action category and its corresponding technical sub-prototype, while the defense action `defender_action` records detection actions, non-intrusive deception actions, and perceptible blocking actions. Attack and defense actions are associated with the attack phase and environmental observations recorded before execution at the same time step.

[0115] Step S33: After the phase transition is completed in the attack and defense game environment, record the attack reward, defense reward and phase sequence, and combine the attack phase, environment observation, attack action, defense action, attack reward, defense reward and phase sequence into a time step seven-tuple.

[0116] Specifically, the attack reward, attacker_reward, records the attack reward. The defense reward (defender_reward) records the defense reward. The phase_dwell records the dwell time of the current attack chain phase and the remaining dwell time of the phase within the phase distribution.

[0117] The recording order is bound to the step sequence of the attack and defense game environment. First, the attack chain stage and the environmental observations visible to the defending agent are recorded before the step. Then, the actions selected synchronously by the attacking agent and the defending agent are recorded. Subsequently, the stage transition, dwell time and immediate reward are determined by the attack and defense game environment. Finally, the attack chain stage and trajectory termination flag are recorded after the step.

[0118] Step S34: Associate the seven-tuples of each time step in chronological order, and then label the trajectory of each associated seven-tuple to obtain the trajectory data of the entire process.

[0119] In an optional embodiment of the present invention, step S34 may include: Step S341: Obtain attack behavior profile parameters, attack and defense strategy generation, strategy version identifier, random seed and field masking rules, and generate trajectory metadata.

[0120] Specifically, each attack and defense trajectory is accompanied by a metadata header, which includes attack behavior profile label, phase distribution parameters of the attack behavior profile, generation number of the attack agent, generation number of the defense agent, attack policy version hash, defense policy version hash, UTC timestamp of the generation time, random seed and field masking rules.

[0121] Step S342: Associate the seven-tuples of each time step according to the trajectory identifier and time sequence, and store the associated seven-tuples of each time step with the trajectory metadata to obtain the labeled trajectory data.

[0122] Specifically, the same trajectory identifier is configured for each time step seven-tuple generated in the same round of attack and defense game, and they are arranged in the order of time steps. Traffic, logs, alarms and action records across attack chain stages, across target nodes and across log sources are correlated, so that the labeled trajectory data retains the correspondence between attack stages, attack actions, defense actions, environmental observations, attack and defense rewards and stage time sequence.

[0123] Step S343: Mark the publication visibility of each field in the trajectory data according to the field masking rules to obtain the full-process trajectory data.

[0124] Specifically, during data generation, true_stage, belief, or The fields are observation, attacher_action, defender_action, attacher_reward, defender_reward, and phase_dwell. When publishing externally, the reserved or masked status of each field is marked according to the field masking rules.

[0125] The `true_stage` field can be used as a target for strongly supervised training. In partial annotation publishing, the `true_stage` field of some trajectories can be masked, retaining only the `observation` field for the corresponding trajectory. The trajectory metadata header, policy version hash, random seed, and field masking rules are not deleted, allowing different data views to be organized based on the same batch of underlying trajectories.

[0126] Step S4: Perform quality verification on the entire process trajectory data, and adjust the attack and defense tasks, attack and defense strategies or filter abnormal trajectories based on the verification results to obtain trajectory data that passes the quality verification.

[0127] Specifically, batch verification is performed on the entire process trajectory data. Quality verification includes stage coverage verification, reward decay consistency verification, trajectory diversity verification, and abnormal trajectory verification. Based on the quality verification results, it is determined whether to release the entire process trajectory data.

[0128] In an optional embodiment of the present invention, step S4 may include: Step S41: Determine the stage coverage, reward decay consistency index, trajectory diversity index, and abnormal trajectory status based on the full-process trajectory data.

[0129] Specifically, the phase coverage is determined based on the actual attack chain phases experienced by each attack and defense trajectory. Trajectories that actually experience two or more CKC attack chain phases are included in the normal sample statistics; trajectories that are too short or only stay at a single attack chain phase are marked as boundary samples.

[0130] The consistency index of return decay is obtained by analyzing a batch of... The empirical decay rate of the expected cumulative reward of the attacking agent is estimated by the attack and defense trajectory. and the empirical decay rate Theoretical decay rate corresponding to the attack behavior profile This was obtained through comparison. This indicates the number of trajectories involved in the consistency check of reward decay.

[0131] The trajectory diversity index is determined based on the action sequence, reward curve, and phase sequence among different attack and defense trajectories within the same trajectory batch. In one optional implementation, the Shannon entropy corresponding to the action sequence, reward curve, and phase sequence is calculated separately, and the trajectory diversity index is determined based on the Shannon entropy of each dimension.

[0132] Abnormal trajectory status is used to mark attack and defense trajectories where the attack reward or defense reward is NaN, the attack reward or defense reward exceeds the configured range, the state dwell time is 0, or the policy output distribution is a degenerate distribution. When the entropy of the policy output distribution is close to 0, the corresponding policy output distribution is determined to be a degenerate distribution.

[0133] Step S42: Compare the stage coverage, reward decay consistency index, trajectory diversity index, and abnormal trajectory status with the corresponding quality verification conditions to obtain the quality verification results.

[0134] Specifically, the phase coverage rate is compared with the phase coverage conditions, and the empirical decay rate is used. Compared with theoretical attenuation rate The deviation between the two is compared with the attenuation deviation threshold, the trajectory diversity index is compared with the diversity threshold, and abnormal trajectories are statistically analyzed based on the abnormal trajectory status.

[0135] In this embodiment, empirical attenuation rate Compared with theoretical attenuation rate The default threshold for deviation is 10%. When the deviation is greater than 10%, the corresponding trajectory batch is determined to have failed the reward decay consistency check.

[0136] Step S43: Adjust the attack and defense tasks, attack and defense strategies or filter abnormal trajectories based on the quality verification results, and supplement the entire process trajectory data based on the adjusted attack and defense tasks or attack and defense strategies to obtain trajectory data that has passed the quality verification.

[0137] Specifically, if any quality verification item fails, the corresponding trajectory batch is not released directly. Instead, the corresponding adjustment branch is entered according to the type of failure, and the adjustment operation is written to the quality verification log and the manifest is released.

[0138] In an optional embodiment of the present invention, step S43 may include: Step S431: Determine the type of failure based on the quality verification results.

[0139] Specifically, the types of failed verifications include insufficient stage coverage, inconsistent reward decay, insufficient trajectory diversity, and the presence of abnormal trajectories. A trajectory batch may correspond to one or more types of failed verifications simultaneously.

[0140] Step S432: When the failure to pass the verification type is insufficient stage coverage, increase the strategy exploration rate of the corresponding attack chain stage.

[0141] Specifically, the phased tasks that need to be adjusted are determined based on the attack chain stages that are not fully covered, and the attack strategy exploration rate corresponding to the early attack chain stages or the attack chain stages that are not fully covered is increased. Then, attack and defense trajectories are generated based on the adjusted attack and defense tasks and attack strategies.

[0142] Step S433: If the failure to pass the verification is due to inconsistent reward decay, adjust the attack behavior profile, reward parameters, or attack and defense strategy.

[0143] Specifically, the phase distribution parameters of the corresponding attack behavior profile are reloaded. and reward weight , , , , The attack and defense strategy models are then trained based on the reloaded phase distribution parameters and reward weights. (Empirical decay rate) Compared with theoretical attenuation rate If the deviation between the two still does not meet the attenuation deviation threshold, the strategy will revert to the attack strategy version or defense strategy version that meets the stability conditions in the previous strategy generation.

[0144] Step S434: If the verification type fails due to insufficient trajectory diversity, select the historical attack and defense strategy to conduct the attack and defense game.

[0145] Specifically, historical attack strategy versions or historical defense strategy versions are selected from the strategy pool to enable the historical attack strategy version to counter the current defense strategy version, or vice versa, in order to supplement the attack and defense trajectories corresponding to different strategy version combinations.

[0146] Step S435: Filter out abnormal trajectories, supplement and generate full-process trajectory data based on the adjusted attack and defense tasks, attack and defense strategies or historical attack and defense strategies, and re-verify the quality of the supplemented full-process trajectory data to obtain trajectory data that passes the quality verification.

[0147] Specifically, trajectories with NaN attack or defense rewards, attack or defense rewards exceeding limits, state dwell time of 0, and policy output distribution degradation are removed, while the corresponding error logs are retained. The adjusted attack and defense tasks, attack and defense policies, policy versions, and reasons for supplementary sampling are written to the quality verification log. Steps S41 and S42 are re-executed on the supplemented full-process trajectory data to filter out the trajectory data that passes the quality verification.

[0148] Step S5: The trajectory data that has passed the quality verification is versioned, organized, and packaged for release to obtain the APT attack and defense dataset.

[0149] In an optional embodiment of the present invention, step S5 may include: Step S51: Obtain the attack and defense strategy generation, and organize the trajectory data that has passed the quality verification according to the attack behavior profile and the attack and defense strategy generation to obtain the trajectory data packet to be released.

[0150] Specifically, trajectory data with the same attack behavior profile and attack / defense strategy generation are organized into the same publishing unit, and a dataset version number is configured for each publishing unit. The dataset version number is incremented each time trajectory data is regenerated and published.

[0151] Step S52: Determine the field masking rules based on the release task, and mask the fields in the trajectory data packet to be released based on the field masking rules to obtain a strongly supervised data view, a semi-supervised data view, and an observation data view.

[0152] In an optional embodiment of the present invention, step S52 may include: Step S521: Retain all fields of the seven-tuple in the trajectory data packet to be published and generate a strongly supervised data view.

[0153] Specifically, strongly supervised data views retain true_stage, belief, or The fields are observation, attacher_action, defender_action, attacher_reward, defender_reward, and phase_dwell, with the true_stage field used as the supervision annotation field.

[0154] Step S522: Mask the attack phase of a portion of the trajectory in the data packet to be published, and generate a semi-supervised data view.

[0155] Specifically, according to the field masking rules, select some trajectories, mask the true_stage field of the selected trajectories, and retain the true_stage field of the unselected trajectories and the observation field of each trajectory to form a semi-supervised data view that includes both labeled and unlabeled trajectories.

[0156] Step S523: Mask the attack phase, attack action, defense action, attack reward, defense reward, and phase sequence in the data packet to be published to generate an observation data view.

[0157] Specifically, the observation data view retains the observation field as well as other visible fields specified by the published task, and provides information on true_stage, belief, or Fields specified by field masking rules in `attacker_action`, `defender_action`, `attacker_reward`, `defender_reward`, and `phase_dwell` are masked. When generating strongly supervised data views, semi-supervised data views, and observation data views, the trajectory metadata header, policy version hash, random seed, and field masking rules are not deleted.

[0158] Step S53: Associate the strongly supervised data view, the semi-supervised data view, and the observation data view with the attack and defense task configuration, the stage dwell time configuration, the quality verification record, and the reproduction information to generate a dataset version and release and encapsulate it to obtain the APT attack and defense dataset.

[0159] Specifically, the packaged content includes strongly supervised data views, semi-supervised data views, observation data views, hyperparameter files for attack strategy models and defense strategy models, phase distribution configuration files, attack and defense task configuration files, quality verification logs, release manifests, and reproduction examples.

[0160] Publishing a manifest records the original configuration, attack behavior profile, phase distribution parameters, attack strategy version, defense strategy version, trajectory batch, strategy generation, field masking rules, trajectory filtering reasons, quality adjustment operations, and dataset version number. By publishing a manifest, the published data can be traced back to the corresponding original configuration, strategy version, and processing reasons.

[0161] The APT attack and defense dataset includes attack traffic, log data, alarm events, defense response records, attack phases, attack strategies, defense strategies, attack and defense results, phase sequence, and strategy evolution information. Different data views generated based on the same batch of full-process trajectory data can be used for APT detection model training, defense agent reinforcement learning, end-to-end network linkage effect evaluation, and full-process strategy review.

[0162] In an optional embodiment of the present invention, the APT attack and defense dataset is further organized into a security large-scale model fine-tuning corpus and a defense agent upgrade sample library. The security large-scale model fine-tuning corpus includes attack chain stages, attack actions, defense actions, environmental observations, attack and defense results, and replay information; the defense agent upgrade sample library includes attack strategy changes, detection results, deception results, blocking results, and defense strategy update records.

[0163] This invention also provides an APT attack and defense data generation system based on intelligent game theory, such as... Figure 3 As shown, the system adopts a three-layer collaborative, four-body driven hierarchical architecture, including an attack and defense task preparation module, a director orchestration module, an attack execution module, a defense response module, a Cangjie data module, a data quality verification module, and a data asset management module. The director orchestration module, attack execution module, defense response module, and Cangjie data module are located in the core module layer, with the director agent, attack agent, defense agent, and Cangjie recorder respectively serving as their functional entities. The attack and defense task preparation module and data quality verification module are located in the support service layer; the data asset management module is located in the data asset layer. Each layer represents the functional affiliation and collaborative relationships between modules. The system's data generation process is achieved through data transfer and function calls between modules.

[0164] The attack and defense task preparation module is used to configure attack and defense tasks based on target environment information, attack behavior profiles, and phase dwell time distributions to obtain the attack and defense game environment. Before the attack and defense game begins, the attack and defense task preparation module completes task initialization and runtime preparation, generates requirement configurations based on APT attack and defense data, including target targets, network topology, business scenarios, key assets, vulnerability information, security devices, log sources, attack phases, phase dwell time distributions, attack action space, defense action space, and reward parameters, and loads the configured data into the attack and defense game environment.

[0165] The attack and defense task preparation module is also used to organize the basic operational flow of attack and defense tasks and provide a unified task context to the director scheduling module, attack execution module, defense response module, and Cangjie data module. The task context enables the director scheduling module to schedule multi-stage attack chain scripts based on the attack and defense tasks, allows the attack execution module and defense response module to compete in the same attack and defense game environment, and enables the Cangjie data module to collect data during the attack and defense game process according to preset data fields and recording granularity. The basic operational flow includes script distribution, attack execution, defense response, and data recording. The attack execution module and defense response module interact bidirectionally through the attack and defense game environment: the attack execution module outputs attack actions, and the defense response module outputs detection actions, deception actions, and blocking actions in response to the attack actions. The attack and defense game environment feeds back the defense response results to the attack execution module and the attack execution results to the defense response module.

[0166] The director orchestration module, located in the core module layer and corresponding to the director agent, is used to orchestrate multi-stage attack chain scripts based on attack and defense tasks, resulting in staged tasks. The director orchestration module determines the attack chain stages based on attack behavior profiles, target environment constraints, attack and defense task objectives, and data generation requirements, and breaks down the multi-stage attack chain into executable and progressive staged tasks.

[0167] The phased tasks include the phase objectives, execution order, success conditions, failure conditions, phase progression rules, and replay rules for each stage of the attack chain. The director scheduling module sends the phased tasks to the attack execution module, defense response module, and Cangjie data module, enabling the attack execution module to execute attack actions according to the current phase objective, the defense response module to detect, deceive, and block responses to the current attack phase, and the Cangjie data module to associate and record the corresponding attack and defense data for the current phase.

[0168] During multiple rounds of attack and defense, the director and scheduling module schedules attack and defense tasks based on the recorded status of attack execution results, defense response results, attack chain phase progression status, and the entire process trajectory data. The director and scheduling module also dynamically scores the task execution status of the current attack chain phase and the performance of attack and defense strategies based on the recorded status of attack execution results, defense response results, attack and defense rewards, attack chain phase progression status, and the entire process trajectory data. Based on the dynamic scoring results, it determines phase advancement, task adjustments, or failure replays. When the attack execution result meets the success conditions of the current phase, the director and scheduling module triggers the phased task corresponding to the next attack chain phase. When the attack action is detected, deceived, or blocked by the defense response module and meets the failure conditions of the current phase, the director and scheduling module replays the failure based on the reasons for failure, generating attack strategy feedback or defense strategy feedback.

[0169] The director and choreography module sends defense strategy feedback to the attack execution module, allowing it to adjust its attack strategy model based on the detection, deception, and blocking strategies implemented by the defense response module. It also sends attack strategy feedback to the defense response module, enabling it to adjust its defense strategy model based on the attack actions, variants, results, and reasons for failure generated by the attack execution module. Furthermore, it sends debriefing results to the attack and defense task preparation module, allowing it to adjust its attack and defense tasks. Thus, the director and choreography module completes the distribution of attack chain scripts, phase progression, dynamic scoring, task scheduling, and failure debriefing during multiple rounds of attack and defense game.

[0170] The attack execution module, located in the core module layer, corresponds to the attack agent. It is used to execute attack actions based on phased tasks in an attack-defense game environment and to train the attack strategy model based on the attack and defense results. The attack execution module generates attack actions based on the current attack chain stage, the current phased task, defense feedback, and the attack strategy model, and sends the attack actions to the attack-defense game environment for execution.

[0171] The attack execution module is capable of executing corresponding attack actions for multi-stage attack chains. In one optional embodiment, the multi-stage attack chain includes reconnaissance, footholding, command and control, lateral movement, and data leakage stages. The attack execution module can invoke corresponding attack actions from phishing, scanning and probing, vulnerability exploitation, credential theft, privilege acquisition, persistence, command and control, lateral movement, and data leakage. Attack actions may include attack action categories, attack technique sub-prototypes, and attack parameters.

[0172] The attack execution module outputs data including attack actions, attack parameters, execution status, attack results, failure reasons, and attack-side sample data. The attack-defense game environment determines the attack chain phase transitions, phase dwell times, and attack-defense rewards based on the attack actions and the defense actions generated by the defense response module, and feeds back the corresponding attack-defense results to the attack execution module.

[0173] The attack execution module trains the attack strategy model based on the attack and defense results and the attack strategy feedback provided by the director and orchestration module. When an attack is detected, deceived, or blocked, the attack execution module can adjust the attack payload, identity spoofing, communication frequency, traffic characteristics, execution path, target selection, or execution time window based on the defense response results and reasons for failure, thereby obtaining an updated attack strategy or attack variant. In subsequent attack and defense games, the attack action is generated based on the updated attack strategy model.

[0174] The defense response module, located in the core module layer, corresponds to the defense agent. It is used to detect, deceive, and block attack actions based on phased tasks in an attack-defense game environment, and to train a defense strategy model based on the attack and defense results. The defense response module receives environmental observations from the attack-defense game environment, including network traffic, log data, terminal behavior, alarm events, service status, and visible traces of attack behavior.

[0175] The defense response module determines the attack phase confidence distribution based on environmental observations and historical actions, and generates defense actions based on the attack phase confidence distribution using a defense strategy model. Defense actions include detection actions, deception actions, and blocking actions, which can be applied at the endpoint, network side, or in endpoint-network coordinated scenarios. The defense response module sends the generated defense actions to the attack-defense game environment to detect, deceive, block, isolate, or handle the attack actions generated by the attack execution module.

[0176] The output data of the defense response module includes detection alarms, spoofing responses, blocking actions, isolation measures, evidence collection records, defense results, and defense-side sample data. The attack-defense game environment generates defense rewards, attack rewards, phase transition results, and phase dwell times based on attack and defense actions, and feeds the corresponding attack and defense results back to the defense response module.

[0177] During multiple rounds of attack and defense, the defense response module continuously trains its defense strategy model based on the strategy changes of the attack execution module. When the attack execution module generates attack variants based on feedback, the defense response module obtains the corresponding attack samples, detection results, handling results, and defense strategy feedback, and updates its detection rules, response strategies, or handling priorities accordingly. The defense response module sends the defense results and defense strategy feedback to the director / arrangement module, and receives the debriefing results and strategy adjustment information generated by the director / arrangement module.

[0178] The attack execution module and the defense response module execute multiple rounds of attack and defense games in the attack and defense game environment. The attack and defense game environment advances the attack chain stages according to the stage dwell time distribution, and trains the attack strategy model and defense strategy model according to the attack and defense results generated by each round of attack and defense games, thereby obtaining the attack and defense game process including attack actions, defense actions, stage transitions, stage dwell time, and attack and defense rewards.

[0179] In one optional embodiment, the attack strategy model and the defense strategy model each include a policy network and a value network. The attack strategy model outputs an attack action probability distribution based on the current attack stage and defense feedback, while the defense strategy model outputs action probability distributions corresponding to detection actions, deception actions, and blocking actions based on environmental observations and attack stage confidence distributions. The attack execution module and the defense response module can alternately train the corresponding attack strategy model and defense strategy model, and save strategy versions from different training generations.

[0180] The Cangjie data module is located in the core module layer and corresponds to the Cangjie record body. It is used to record and label the attack and defense game process according to a seven-tuple including stage information, environmental observation, attack action, defense action, attack reward, defense reward and stage time sequence, so as to obtain the trajectory data of the whole process.

[0181] The Cangjie data module records multimodal behavior throughout the entire multi-round attack-defense game process. It collects multi-source logs and multimodal data on network traffic, log data, terminal behavior, alarm events, environmental observations, attack actions, defense actions, attack-defense rewards, and phase sequence data. It also unifies and aggregates data from the director / arrangement module, attack execution module, defense response module, and the attack-defense game environment. Phase information includes the actual attack phase and attack phase confidence distribution; environmental observations include network traffic, log data, terminal behavior, alarm events, and decoy hit information; and phase sequence data includes the dwell time of each stage of the attack chain.

[0182] The Cangjie data module records data at each time step according to the runtime sequence of the attack-defense game environment. Before the attack execution module and defense response module select actions, it records the current stage information and environmental observations; after the attack execution module and defense response module generate actions, it records the attack actions and defense actions; after the attack-defense game environment completes the stage transition, it records the attack reward, defense reward, and stage sequence, and associates the corresponding data into a seven-tuple for the current time step.

[0183] The Cangjie data module associates seven-tuples across multiple time steps according to trajectory identifiers and time sequence, and organizes and correlates data across attack phases, target hosts, and log sources to form full-process trajectory data. This full-process trajectory data preserves the correspondence between attack phases, environmental observations, attack actions, defense actions, attack and defense results, and the time sequence of each phase.

[0184] The Cangjie data module is also used for structured annotation of the entire process trajectory data. Annotation information may include attack phases, attack techniques, attack actions, defense actions, detection results, deception results, blocking results, successful paths, reasons for failure, attack strategy variations, environmental context, and data source information. The Cangjie data module is also used to organize the completed data association and structured annotation of the entire process trajectory data according to a preset corpus format, forming a structured attack and defense corpus that includes attack chain phases, environmental observations, attack actions, defense actions, attack and defense results, and recap information. This achieves the corpus accumulation of attack and defense process data and provides the structured attack and defense corpus to the data asset management module.

[0185] Each complete trajectory data point can also be associated with trajectory metadata, which includes attack behavior profile labels, phase dwell time distribution parameters, attack strategy generation, defense strategy generation, attack strategy version identifier, defense strategy version identifier, generation time, random seed, and field masking rules. The trajectory metadata allows for tracing the attack and defense tasks, strategy versions, and generation conditions corresponding to the complete trajectory data.

[0186] The data quality verification module, located in the support service layer, performs quality verification on the entire trajectory data, filters out valid trajectory data, and generates quality feedback based on the verification results for adjusting attack and defense tasks or updating attack and defense strategies. The data quality verification module also manages the attack and defense strategy versions involved in the quality verification process, recording the strategy generation, verification status, stage coverage, and reward performance for each strategy version. Based on the quality verification results, it determines whether a strategy version should continue to be used, retrained, resampled, or rolled back.

[0187] The data quality verification module receives the full-process trajectory data generated by the Cangjie data module and performs integrity verification, consistency verification, stage coverage verification, reward decay consistency verification, trajectory diversity verification, label quality verification, abnormal trajectory verification, duplicate sample filtering, and data format normalization on the full-process trajectory data. The integrity verification, abnormal trajectory filtering, duplicate sample filtering, invalid field handling, and data format normalization together constitute the data cleaning process. Through this data cleaning process, the data quality verification module obtains valid trajectory data that meets preset data format and quality requirements. The data quality verification module also performs data release access control on the cleaned valid trajectory data. Based on the quality verification results, it determines whether the corresponding trajectory batch is allowed to be released, generates a release record including the quality verification results, abnormal trajectory filtering results, strategy version information, and release status, and sends the allowed valid trajectory data and the release record to the data asset management module for versioned organization and release.

[0188] Phase coverage verification is used to determine the actual attack chain phases experienced by the attack and defense trajectory; reward decay consistency verification is used to compare the empirical reward decay determined based on the full-process trajectory data with the theoretical reward decay corresponding to the attack behavior profile; trajectory diversity verification is used to analyze the differences between different attack and defense trajectories in terms of action sequence, reward change and phase timing; abnormal trajectory verification is used to identify attack and defense trajectories with abnormal rewards, abnormal phase dwell time or degraded strategy output.

[0189] The data quality verification module filters the entire process trajectory data based on the quality verification results, identifying those that meet the quality verification criteria as valid trajectory data and removing abnormal trajectories that fail the quality verification. For trajectory batches that fail the quality verification, the data quality verification module generates quality feedback based on the type of verification failure.

[0190] When the phase coverage does not meet the quality verification conditions, the quality feedback is used to instruct the attack and defense task preparation module or the director orchestration module to adjust the task configuration of the corresponding attack chain phase; when the reward decay consistency does not meet the quality verification conditions, the quality feedback is used to instruct the attack and defense task preparation module to adjust the attack behavior profile, phase dwell time distribution, or reward parameters, or to instruct the attack execution module and the defense response module to adjust the attack and defense strategies; when the trajectory diversity does not meet the quality verification conditions, the quality feedback is used to instruct the attack execution module and the defense response module to select different strategy versions for attack and defense game; when there are abnormal trajectories, the quality feedback is used to instruct the Cangjie data module to retain the corresponding error records, and the data quality verification module to filter the abnormal trajectories.

[0191] The data quality verification module sends quality feedback to the attack / defense task preparation module, the director / arrangement module, the attack execution module, the defense response module, or the Cangjie data module. This allows the corresponding module to adjust the attack / defense tasks, phased tasks, attack strategies, defense strategies, or data recording methods based on the quality feedback. After the adjustment is completed, the attack execution module and the defense response module re-execute the attack / defense game, the Cangjie data module supplements and generates full-process trajectory data, and the data quality verification module performs quality verification on the supplemented full-process trajectory data again.

[0192] The data asset management module, located in the data asset layer, is used to version-organize valid trajectory data to obtain an APT attack and defense dataset. The data asset management module receives valid trajectory data filtered by the data quality verification module and organizes it according to attack behavior profiles, attack strategy generations, defense strategy generations, or data generation batches.

[0193] The data asset management module configures dataset version identifiers for the effective trajectory data after organization, and associates the effective trajectory data with attack and defense task configurations, phase dwell time distribution configurations, strategy version information, random seeds, field masking rules, quality verification records, and reproduction information to form an APT attack and defense dataset that can be reproduced and traced.

[0194] In one optional embodiment, the data asset management module masks fields in the valid trajectory data according to field masking rules, and generates a strongly supervised data view, a semi-supervised data view, and an observation data view based on the same batch of valid trajectory data. The strongly supervised data view retains the complete fields of the seven-tuple; the semi-supervised data view masks the actual attack phase of some trajectories; and the observation data view retains the visible fields specified by the environmental observation and data publishing tasks.

[0195] The APT attack and defense dataset includes attack traffic, log data, alarm events, defense response records, attack phases, attack actions, defense actions, attack rewards, defense rewards, phase sequence, attack strategy versions, defense strategy versions, attack and defense results, replay information, and strategy evolution information.

[0196] The data asset management module can also organize APT attack and defense datasets according to their intended use, forming a security large-scale model fine-tuning corpus and a defense agent upgrade sample library. The security large-scale model fine-tuning corpus includes attack chain stages, environmental observations, attack actions, defense actions, attack and defense results, and debriefing information, used for attack chain understanding, alarm analysis, attack and defense debriefing, strategy generation, and security operation auxiliary decision-making tasks. The defense agent upgrade sample library includes attack strategy changes, attack samples, detection results, deception results, blocking results, and defense strategy update records, used for training and updating the defense strategy model.

[0197] In actual deployment, the APT attack and defense data generation system based on intelligent game theory and confrontation operates in the following order: attack and defense task configuration, attack chain script arrangement, attack action execution, defense detection and response, full-process data recording, data quality verification, quality feedback and data asset organization.

[0198] First, the attack and defense task preparation module configures attack and defense tasks based on target environment information, attack behavior profiles, and stage dwell time distributions to obtain the attack and defense game environment. The director scheduling module schedules multi-stage attack chain scripts based on the attack and defense tasks to obtain staged tasks. The attack execution module and the defense response module execute multiple rounds of attack and defense games based on the staged tasks in the attack and defense game environment, advance the attack chain stages according to the stage dwell time distribution, and train the attack strategy model and defense strategy model based on the attack and defense results to obtain the attack and defense game process.

[0199] During the multi-round attack and defense game, the Cangjie data module continuously records and labels the attack and defense game process according to the seven-tuple, and obtains the trajectory data of the whole process; the data quality verification module performs quality verification on the trajectory data of the whole process, filters out the valid trajectory data, and provides quality feedback to the corresponding module; the data asset management module organizes the valid trajectory data in a versioned manner to obtain the APT attack and defense dataset.

[0200] Although the present invention has been disclosed above with reference to embodiments, it is not intended to limit the present invention. Appropriate modifications or equivalent substitutions made by those skilled in the art to the technical solutions of the present invention should be covered within the protection scope of the present invention, which is defined by the claims.

Claims

1. A method for generating APT attack and defense data based on intelligent game theory, characterized in that, Includes the following steps: Based on target environment information, attack behavior profile, phase dwell time distribution and attack and defense actions, configure attack and defense tasks, and arrange multi-stage attack chain scripts to obtain the attack and defense game environment and phased tasks. Based on the aforementioned attack and defense game environment and phased tasks, the attacking agent executes attack actions, and the defending agent detects, deceives, and blocks responses. The attack and defense strategy is updated according to the attack and defense results, resulting in a multi-round attack and defense game process. The multi-round attack and defense game process is recorded and labeled according to a seven-tuple consisting of attack phase, environmental observation, attack action, defense action, attack reward, defense reward and phase sequence to obtain the full process trajectory data; The entire process trajectory data is subjected to quality verification, and the attack and defense tasks, attack and defense strategies or abnormal trajectories are adjusted or filtered according to the verification results to obtain trajectory data that passes the quality verification. The trajectory data that has passed quality verification is versioned, organized, and packaged to obtain the APT attack and defense dataset.

2. The method as described in claim 1, characterized in that, Based on target environment information, attack behavior profiles, phase dwell time distribution, and attack and defense actions, attack and defense tasks are configured, and multi-phase attack chain scripts are arranged to obtain the attack and defense game environment and phased tasks, including: Based on the target environment information and attack and defense actions, configure the network topology, key assets, vulnerability information, and attack and defense action space to obtain environmental constraint information; Based on the attack behavior profile and the stage dwell time distribution, the attack chain stages, the dwell rules of each attack chain stage, and the stage transition rules are determined to obtain the attack chain configuration information. Based on the environmental constraints and attack chain configuration information, attack and defense tasks are configured, and multi-stage attack chain scripts are arranged according to the attack chain stages to obtain the attack and defense game environment and staged tasks.

3. The method as described in claim 2, characterized in that, Based on the attack behavior profile and stage dwell time distribution, the attack chain stages, dwell rules for each attack chain stage, and stage transition rules are determined, resulting in attack chain configuration information, including: Determine the attack chain phases and the behavioral configurations for each attack chain phase based on the attack behavior profile; The dwell rules for each attack chain stage are determined based on the dwell time distribution of the phase-type stage corresponding to each attack chain stage. Based on the offensive and defensive actions, the behavior configuration, and the dwell rules, the phase transition rules are determined, and the attack chain configuration information is obtained.

4. The method as described in claim 1, characterized in that, Based on the aforementioned attack-defense game environment and phased tasks, an attacking agent executes attack actions, while a defending agent detects, deceives, and blocks responses. The attack-defense strategy is updated based on the attack-defense results, resulting in a multi-round attack-defense game process, including: The attacking agent generates attack actions based on the phased tasks, the current attack phase, and defense feedback. The defensive agent determines the attack phase confidence distribution based on environmental observations and action history, and determines detection actions, deception actions, and blocking actions based on the attack phase confidence distribution to obtain defensive actions; In the attack and defense game environment, the attack action and the defense action are executed. The phase transition, residence time, attack reward and defense reward are determined according to the phase residence time distribution to obtain the attack and defense result. The attack strategy model and defense strategy model are trained based on the attack and defense results to obtain updated attack and defense strategies. The attack and defense game is then repeatedly executed based on the updated attack and defense strategies to obtain a multi-round attack and defense game process.

5. The method as described in claim 4, characterized in that, The attack strategy model and defense strategy model are trained based on the attack and defense results to obtain updated attack and defense strategies. The attack and defense game is then repeatedly executed based on these updated strategies to obtain a multi-round attack and defense game process, including: The multiple attack and defense results are aggregated to obtain a batch of attack and defense results; Based on the batches of attack and defense results, the cumulative reward, stage coverage and strategy output distribution are determined, and when the strategy update conditions are met, the attack strategy model and defense strategy model are trained by multi-agent reinforcement learning to obtain the updated attack strategy and defense strategy. Generate policy version identifiers and policy generations for the updated attack and defense strategies, and store the corresponding policy versions and historical policy versions in the policy pool; The sampling mode is determined from the training sampling mode, the evolution sampling mode and the steady-state sampling mode, and the latest policy version, the historical policy version or a combination of different policy versions are selected from the policy pool according to the sampling mode to obtain candidate attack policies and candidate defense policies. When the strategy output distribution meets the degradation condition or the stage coverage decreases, the candidate attack strategy or candidate defense strategy is replaced with a historical strategy version that meets the stability condition, and the strategy exploration rate is increased. The attack and defense game is repeatedly executed based on the candidate attack strategy and candidate defense strategy to obtain a multi-round attack and defense game process.

6. The method as described in claim 1, characterized in that, The multi-round attack-defense game process is recorded and labeled according to a seven-tuple consisting of attack phase, environmental observation, attack action, defense action, attack reward, defense reward, and phase sequence, to obtain the full-process trajectory data, including: Before selecting attack and defense actions at each time step, record the attack phase and environmental observations. After simultaneously selecting attack and defense actions at each time step, the attack and defense actions are recorded. After the phase transition is completed in the attack and defense game environment, the attack reward, defense reward and phase sequence are recorded, and the attack phase, environment observation, attack action, defense action, attack reward, defense reward and phase sequence are combined into a time step seven-tuple. By associating the seven-tuples of each time step in chronological order and then labeling the trajectory of each associated seven-tuple, the trajectory data of the entire process can be obtained.

7. The method as described in claim 6, characterized in that, By associating the seven-tuples of each time step in chronological order and then labeling the trajectory of each associated seven-tuple, the entire process trajectory data is obtained, including: Obtain attack behavior profile parameters, attack and defense strategy generation, strategy version identifier, random seed and field masking rules, and generate trajectory metadata; The seven-tuples of each time step are associated with the trajectory identifier and time sequence, and the associated seven-tuples of each time step are stored in correspondence with the trajectory metadata to obtain the labeled trajectory data. The publication visibility of each field in the labeled trajectory data is marked according to the field masking rules to obtain the full-process trajectory data.

8. The method as described in claim 1, characterized in that, The entire trajectory data is subjected to quality verification, and the attack and defense tasks, attack and defense strategies, or abnormal trajectories are adjusted based on the verification results to obtain trajectory data that passes the quality verification, including: Based on the full-process trajectory data, determine the stage coverage rate, the consistency index of reward decay, the trajectory diversity index, and the abnormal trajectory status; The stage coverage, reward decay consistency index, trajectory diversity index, and abnormal trajectory status are compared with the corresponding quality verification conditions to obtain the quality verification results. Based on the quality verification results, adjust the attack and defense tasks and strategies or filter abnormal trajectories, and supplement the entire process trajectory data based on the adjusted attack and defense tasks or strategies to obtain trajectory data that has passed the quality verification.

9. The method as described in claim 1, characterized in that, The trajectory data that has passed quality verification is versioned, organized, and packaged for release to obtain the APT attack and defense dataset, including: The attack and defense strategy hierarchy is obtained, and the trajectory data that has passed the quality verification is organized according to the attack behavior profile and the attack and defense strategy hierarchy to obtain the trajectory data packet to be released. The field masking rules are determined based on the published task, and the fields in the trajectory data packet to be published are masked based on the field masking rules to obtain a strongly supervised data view, a semi-supervised data view, and an observation data view. The strongly supervised data view, semi-supervised data view, and observation data view are associated with the attack and defense task configuration, stage dwell time configuration, quality verification records, and reproduction information to generate a dataset version and release and encapsulate it to obtain the APT attack and defense dataset.

10. An APT attack and defense data generation system based on intelligent game theory, characterized in that, It includes modules for attack and defense task preparation, director and choreography, attack execution, defense response, Cangjie data, data quality verification, and data asset management. The attack and defense task preparation module is used to configure attack and defense tasks based on target environment information, attack behavior profile and stage dwell time distribution to obtain the attack and defense game environment. The director scheduling module is used to schedule multi-stage attack chain scripts according to the attack and defense tasks to obtain staged tasks. The attack execution module and the defense response module are used to execute multiple rounds of attack and defense games based on the phased tasks in the attack and defense game environment, advance the attack chain stages according to the phase dwell time distribution, and train the attack strategy model and defense strategy model according to the attack and defense results to obtain the attack and defense game process. The Cangjie data module is used to record and label the attack and defense game process according to a seven-tuple including stage information, environmental observation, attack action, defense action, attack reward, defense reward and stage sequence, to obtain the trajectory data of the whole process; The data quality verification module is used to verify the quality of the entire process trajectory data, filter out valid trajectory data, and generate quality feedback based on the verification results for adjusting attack and defense tasks or updating attack and defense strategies. The data asset management module is used to version the valid trajectory data to obtain an APT attack and defense dataset.