Intelligent agent adversarial evolution simulation and game verification system in virtual parallel space

By constructing an agent-based adversarial evolutionary simulation and game verification system in a virtual parallel space, the problems of insufficient environmental change mapping and poor policy robustness in existing multi-agent simulation technologies are solved, and stable policy iteration and reliable output are achieved.

CN121920231APending Publication Date: 2026-04-24HANGZHOU TUBU ER TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANGZHOU TUBU ER TECHNOLOGY CO LTD
Filing Date
2026-01-21
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively reflect dynamic environmental changes in the physical world during multi-agent simulations. Poor policy robustness and unclear coupling between the environmental model and the agent decision-making model result in insufficient policy adaptability to environmental changes, making it difficult to achieve parallel mapping and continuous verification of the real system.

Method used

A simulation and game-theoretic verification system for adversarial evolution of intelligent agents in a virtual parallel space is constructed. By collecting static environmental data and dynamic situational data from the physical world, a virtual parallel simulation space is built, and state perception input, action output, and resource constraints are configured. An improved NFSP algorithm is used to update the policy parameters through adversarial evolution, and quantitative calculation and policy selection are performed through a game-theoretic verification index system.

Benefits of technology

It achieves a close mapping between the simulation environment and the real scene, the strategy iteration and update are stable, the game verification results are quantifiable and reusable, the credibility and adaptability of the strategy output are improved, and the problems of insufficient policy adaptability and fragmented verification in the existing technology are solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121920231A_ABST
    Figure CN121920231A_ABST
Patent Text Reader

Abstract

The invention discloses an agent adversarial evolution simulation and game verification system in a virtual parallel space, and the system comprises a virtual parallel simulation space construction module which is used for constructing the virtual parallel simulation space; the confrontation agent modeling module is used for constructing and configuring a first agent and a second agent; the confrontation interaction simulation module is used for executing autonomous decisions and interaction actions and generating a confrontation interaction sequence; the adversarial evolutionary updating module is used for executing adversarial evolutionary updating by adopting an improved NFSP algorithm and generating updated strategy parameters; the strategy iteration simulation module is used for generating an updated confrontation interaction sequence; the confrontation result analysis module is used for generating confrontation result data; the game verification module is used for performing quantitative calculation based on a preset game verification index system to generate a game verification result; the strategy screening output module is used for outputting a confrontation strategy result, and closed-loop cooperation of confrontation simulation, strategy evolution and game verification is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of agent simulation and game analysis, and in particular to an agent adversarial evolution simulation and game verification system in a virtual parallel space. Background Technology

[0002] With the development of artificial intelligence technology, agent-based simulation and decision-making systems have been widely used in complex system analysis, strategy deduction, adversarial assessment, and behavior verification. Especially in application scenarios with multiple participants, multiple policy spaces, and strong adversarial relationships, constructing agent models and having them interact and learn in a simulation environment has become an important technical means to verify the effectiveness and stability of policies.

[0003] In existing technologies, multi-agent simulations typically rely on pre-defined rules or a single learning model to perform policy evaluation in fixed or semi-fixed simulation environments. Some studies introduce game theory models or reinforcement learning methods, enabling agents to achieve policy optimization through repeated trials in adversarial environments. However, these methods often focus on policy training within a single simulation space, frequently lacking the ability to map complex real-world environments and failing to effectively reflect the dynamic changes and multi-dimensional constraints of the physical world. Regarding simulation environment construction, existing technologies typically treat the environment as a static or weakly dynamic system, with limited changes in environmental parameters during training, failing to fully model the dynamic situations in real-world scenarios. This results in insufficient adaptability and poor policy robustness when faced with environmental changes. Furthermore, the coupling between the environmental model and the agent's decision-making model is not clear enough, making it difficult to achieve parallel mapping and continuous verification of the real system. Summary of the Invention

[0004] One objective of this invention is to propose an agent-based adversarial evolution simulation and game-theoretic verification system in a virtual parallel space. This invention fully utilizes virtual parallel simulation space modeling, multi-agent adversarial interaction, an improved NFSP algorithm driven by adversarial trajectory data, and a game-theoretic verification index system. It constructs a closed-loop process from physical world data acquisition and fusion, adversarial simulation to generate adversarial interaction sequences, adversarial evolution to update strategy parameters, and game-theoretic verification to filter and output adversarial strategy results. It has the advantages of simulation closely resembling real-world scenarios, stable iterative updates of adversarial strategies, quantifiable and reusable game-theoretic verification results, high reliability of strategy output, and strong scalability and adaptability.

[0005] A simulation and game-theoretic verification system for adversarial evolution of intelligent agents in a virtual parallel space according to an embodiment of the present invention includes:

[0006] The virtual parallel simulation space construction module is used to collect static environmental data and dynamic situational data of the physical world, fuse and model the static environmental data and dynamic situational data, and construct a virtual parallel simulation space corresponding to the state of the physical world.

[0007] The adversarial agent modeling module is used to construct two adversarial agent models in a virtual parallel simulation space, which are respectively set as the first agent and the second agent. The module configures the state perception input set, action output set, resource constraints and initial policy parameters for the first agent and the second agent.

[0008] The adversarial interaction simulation module is used to drive the first and second intelligent agents to perform autonomous decision-making and interaction actions according to the time step in a virtual parallel simulation space, and generate adversarial interaction sequences.

[0009] The adversarial evolution update module is used to generate adversarial trajectory data based on adversarial interaction sequences, and to perform adversarial evolution update on the policy parameters of the first agent and the second agent using an improved NFSP algorithm to generate updated policy parameters.

[0010] The strategy iteration simulation module is used to load the updated strategy parameters into the first and second agents, execute adversarial interactions in a virtual parallel simulation space, and generate an updated adversarial interaction sequence.

[0011] The adversarial result parsing module is used to parse the updated adversarial interaction sequence and generate adversarial result data;

[0012] The game verification module is used to quantify the adversarial result data based on a preset game verification index system and generate game verification results.

[0013] The strategy filtering and output module is used to filter the set of strategy parameters that meet the preset verification conditions based on the game verification results, and output the set of strategy parameters as the adversarial strategy results verified by the game.

[0014] Optionally, modules can be integrated using the following methods:

[0015] Collect static environmental data and dynamic situational data from the physical world, and fuse them into a virtual parallel simulation space through modeling.

[0016] In a virtual parallel simulation space, two adversarial intelligent agent models are constructed, designated as the first intelligent agent and the second intelligent agent, respectively, and configured with state-aware input sets, action output sets, resource constraints, and initial policy parameters.

[0017] Based on a virtual parallel simulation space, the first and second intelligent agents are driven to make autonomous decisions and interactive actions according to the time step, forming an adversarial interaction sequence.

[0018] Based on the state sequence, action sequence and adversarial result sequence in the adversarial interaction sequence, adversarial trajectory data is generated. An improved NFSP algorithm is used to perform adversarial evolutionary updates on the policy parameters of the first agent and the second agent to generate updated policy parameters.

[0019] The updated strategy parameters are loaded into the first and second agents, and the first and second agents are driven to perform adversarial interactions in a virtual parallel simulation space to generate an updated adversarial interaction sequence.

[0020] The result information of the interaction between the first agent and the second agent in the updated adversarial interaction sequence is analyzed to generate adversarial result data;

[0021] Based on a pre-set game verification index system, the adversarial result data is quantitatively calculated to generate game verification results.

[0022] The set of strategy parameters that meet the preset verification conditions is selected based on the game verification results, and the set of strategy parameters is output as the adversarial strategy result of the game verification.

[0023] Optionally, the construction process involves collecting static environmental data and dynamic situational data from the physical world. The static environmental data refers to environmental information in the physical world, and the dynamic situational data refers to situational information in the physical world that changes over time. The static environmental data and the dynamic situational data are then time-correlated and fused to generate simulation environment data of physical world state changes at a unified time step. Based on the simulation environment data, a virtual operating environment corresponding to the physical world state is established in the computer system, forming a virtual parallel simulation space.

[0024] Optionally, the construction and configuration of the first and second intelligent agents specifically include:

[0025] Based on the environmental state of the virtual parallel simulation space, two adversarial intelligent agent models are constructed and designated as the first intelligent agent model and the second intelligent agent model, respectively. The construction process is as follows: in the virtual parallel simulation space, two independent decision model instances are established based on the same simulation environment state, and the same environment state acquisition method and interaction interface are configured for the two decision model instances, forming the first intelligent agent and the second intelligent agent that are adversarial in the same virtual parallel simulation space.

[0026] A state-aware input set is configured for the first intelligent agent and the second intelligent agent respectively. The state-aware input set refers to the environmental state information of the virtual parallel simulation space at the corresponding time step.

[0027] The first intelligent agent and the second intelligent agent are configured with action output sets respectively. The action output sets refer to the range of interactive actions that the first intelligent agent and the second intelligent agent can perform in the virtual parallel simulation space.

[0028] Resource constraints are configured for the first intelligent agent and the second intelligent agent respectively. The resource constraints refer to limiting the resource usage range of the first intelligent agent and the second intelligent agent when performing interactive actions in the action output set.

[0029] Initial policy parameters are configured for the first and second intelligent agents respectively. The initial policy parameters refer to the interactive actions in the action output set generated based on the state-aware input set.

[0030] Optionally, the formation of the adversarial interaction sequence specifically includes:

[0031] Arrange the spatial states of the virtual parallel simulation space into a state sequence in chronological order;

[0032] The spatial state is input into the first and second intelligent agents respectively to execute autonomous decision-making, generate interactive actions, and form an action sequence in chronological order. The generation of the interactive actions is to write the spatial state of the virtual parallel simulation space into the state perception input set of the first and second intelligent agents. The first and second intelligent agents respectively call the initial policy parameters to perform decision calculation on the spatial state and obtain the action selection result corresponding to the action output set.

[0033] At each time step, the interactive action is input into the virtual parallel simulation space to drive the spatial state update of the virtual parallel simulation space. The update process is to write the interactive action into the state update interface of the virtual parallel simulation space, call the environment state transition rules to calculate and process the current spatial state, generate the spatial state corresponding to the next time step, and write the spatial state corresponding to the next time step into the spatial state storage of the virtual parallel simulation space to complete the spatial state update.

[0034] Based on the updated spatial state and the corresponding interactive actions at each time step, the updated spatial state is calculated according to the adversarial result determination rules to obtain the adversarial result at each time step. The adversarial results obtained at each time step are recorded and collected in chronological order to form an adversarial result sequence.

[0035] The state sequence, action sequence, and adversarial result sequence are concatenated to form an adversarial interaction sequence.

[0036] Optionally, the generation of the updated strategy parameters specifically includes:

[0037] Adversarial trajectory data is constructed based on the state sequence, action sequence, and adversarial result sequence in the adversarial interaction sequence. An improved NFSP algorithm is used to perform adversarial evolutionary updates on the policy parameters of the first agent and the second agent. The improved NFSP algorithm includes an adversarial trajectory data sampling unit and a dual policy parameter update unit. The adversarial trajectory data sampling unit refers to performing sampling processing on the adversarial trajectory data to generate batches of sampled data. The dual policy parameter update unit refers to introducing a mirror adversarial update mechanism to perform adversarial evolutionary updates on the policy parameters of the first agent and the second agent and output the updated policy parameters.

[0038] In the adversarial trajectory data sampling unit, the adversarial trajectory data is sampled to generate a batch of sampled data. The sampling process involves indexing and marking the adversarial trajectory data, dividing the trajectory elements into sampling ranges according to the time step, selecting trajectory elements from the sampling range according to the sampling rules, and combining the selected trajectory elements in chronological order to form a batch of sampled data.

[0039] In the dual-policy parameter update unit, adversarial evolutionary updates are performed on the policy parameters of the first agent and the second agent based on the sampled data batch. The adversarial evolutionary update introduces a mirror adversarial update mechanism, which generates the update direction of the optimal response policy parameter and the update direction of the average policy parameter for the sampled data batch. The update direction is mapped to the mirror space to complete the parameter update. The updated parameters are then mapped from the mirror space back to the policy parameter space, and the updated policy parameters are output.

[0040] Optionally, the update process involves writing the updated policy parameters into the policy parameter storage of the first and second agents respectively in the virtual parallel simulation space. The first and second agents use the updated policy parameters as the current decision-making basis and drive the virtual parallel simulation space to run according to the time step. At each time step, the spatial state of the virtual parallel simulation space is input into the first and second agents. The first and second agents respectively execute autonomous decision-making output interaction actions and input the interaction actions into the virtual parallel simulation space to drive the spatial state update of the virtual parallel simulation space. The updated spatial state and the interaction actions are recorded in chronological order to form an updated adversarial interaction sequence.

[0041] Optionally, the generation of the adversarial result data specifically includes:

[0042] The updated adversarial interaction sequence is aligned to generate an aligned adversarial interaction sequence.

[0043] At each time step, the result information of the interaction between the first agent and the second agent is parsed based on the aligned adversarial interaction sequence.

[0044] The state difference in the spatial state update information is extracted from the interaction result information at each time step. The state difference is the difference between the state before the spatial state update and the state after the spatial state update at the same time step. The state difference is used as the basis for judging the adversarial result. Combined with the interaction actions of the first agent and the interaction actions of the second agent, adversarial result records for the corresponding time steps are generated. The adversarial result records are collected in chronological order to obtain adversarial result data.

[0045] Optionally, the generation of the game verification result specifically includes:

[0046] Based on each adversarial result record in the adversarial result data, extract the adversarial result value corresponding to the first agent and the adversarial result value corresponding to the second agent respectively.

[0047] The adversarial result values ​​corresponding to the first agent and the second agent are compared and statistically analyzed to generate the number of wins and the win rate of the first agent. The generation process is to read the adversarial result values ​​corresponding to the first agent and the second agent one by one in chronological order. When the adversarial result value corresponding to the first agent is greater than the adversarial result value corresponding to the second agent, the adversarial result record is counted as a win and the number of wins is accumulated. After the comparison of all adversarial result records is completed, the number of wins is divided by the total number of adversarial result records to obtain the win rate of the first agent.

[0048] The cumulative adversarial result value of the first agent is obtained by summing the adversarial result values ​​of the second agent, and the cumulative adversarial result value of the second agent is obtained by summing the adversarial result values ​​of the first agent and the second agent. The average value of the adversarial result difference is obtained by statistically calculating the difference between the adversarial result values ​​of the first agent and the second agent.

[0049] The game verification results are formed based on the number of wins, win rate, cumulative adversarial result value of the first agent, cumulative adversarial result value of the second agent, and the average of the adversarial result difference.

[0050] Optionally, the output of the adversarial strategy result specifically includes:

[0051] Based on the strategy parameters corresponding to the game verification results, a correspondence between strategy parameters and game verification results is formed. The process of forming this correspondence involves generating a unique identifier for each set of strategy parameters, writing the unique identifier into the corresponding game verification result, and using the unique identifier as an index during storage to associate and register the strategy parameters with the game verification results, thus forming a one-to-one correspondence between strategy parameters and game verification results.

[0052] For each set of strategy parameters in the corresponding relationship, a comprehensive calculation is performed on the mean of the number of wins, win rate, cumulative adversarial result value of the first agent, cumulative adversarial result value of the second agent, and the adversarial result difference in the game verification results, to generate a strategy parameter evaluation value corresponding to each set of strategy parameters;

[0053] The strategy parameter evaluation values ​​are matched with preset verification conditions to filter strategy parameters that meet the preset verification conditions, and these filters are then compiled into a strategy parameter set. This set of strategy parameters is used as the game-verified adversarial strategy result. The filtering process involves reading the corresponding strategy parameter evaluation value for each set of strategy parameters, comparing the evaluation value with the verification threshold in the preset verification conditions, and selecting the passed strategy parameter into the strategy parameter set when the evaluation value is greater than or equal to the verification threshold. When the evaluation value is less than the verification threshold, the failed strategy parameters are removed. After traversing all strategy parameters, the strategy parameter set is output as the game-verified adversarial strategy result.

[0054] The beneficial effects of this invention are:

[0055] This invention collects static environmental data and dynamic situational data from the physical world and performs time correlation and fusion modeling to construct a virtual parallel simulation space in a computer system that corresponds to the state of the physical world. This enables adversarial simulation to have a mappable and reproducible environmental foundation, thereby avoiding the problems of the simulation environment being disconnected from the real scene and the difficulty in depicting environmental changes leading to insufficient strategy adaptability in the prior art. It also improves the simulation environment's ability to withstand complex situational changes and its consistency.

[0056] This invention constructs two adversarial agent models in a virtual parallel simulation space. By structurally defining the perception boundaries, behavioral boundaries, and resource boundaries of the agents through a set of state-aware inputs, a set of action outputs, resource constraints, and initial policy parameters, the first and second agents can perform autonomous decisions and interactive actions under a unified constraint system, forming an adversarial interaction sequence. This overcomes the problems of unclear agent behavioral boundaries and difficulty in standardizing, aligning, and reusing interaction data in existing multi-agent simulations, and improves the controllability and analyzability of the adversarial interaction process.

[0057] This invention further constructs adversarial trajectory data based on adversarial interaction sequences, and uses an improved NFSP algorithm containing adversarial trajectory data sampling units and dual policy parameter update units to perform adversarial evolutionary updates on policy parameters. This enables policy updates to continuously iterate under the drive of adversarial trajectory data, and improves the data organization efficiency of policy updates by sampling data batches. By using dual policy parameter updates, it takes into account both immediate adversarial capabilities and long-term policy stability, thereby improving the shortcomings of existing technologies such as unbalanced distribution of adversarial training data, unstable policy updates, and difficulty in balancing short-term gains and long-term behavioral consistency.

[0058] This invention drives the first and second agents to generate updated adversarial interaction sequences in a virtual parallel simulation space after the strategy parameters are updated. The updated adversarial interaction sequences are then aligned and parsed to form adversarial result data. Based on a preset game verification index system, the adversarial result data is quantified to obtain the game verification result. The strategy parameter set is then filtered and output based on the strategy parameter evaluation value and verification threshold, forming a closed-loop process of strategy training, adversarial verification, and strategy filtering. This overcomes the problems of fragmented training and verification, single verification index, and difficulty in supporting strategy filtering output in the prior art, and improves the quantifiability of game verification results and the credibility, stability, and reusability of strategy output. Attached Figure Description

[0059] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0060] Figure 1 This is the overall flowchart of the virtual parallel space intelligent agent adversarial evolution simulation and game verification system proposed in this invention.

[0061] Figure 2 This is a schematic diagram showing the construction and configuration of the first and second agents in the virtual parallel space agent adversarial evolution simulation and game verification system proposed in this invention.

[0062] Figure 3 This is a schematic diagram illustrating the construction of updated strategy parameters for the agent adversarial evolution simulation and game verification system in a virtual parallel space proposed in this invention. Detailed Implementation

[0063] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0064] refer to Figures 1-3 A simulation and game-theoretic verification system for adversarial evolution of intelligent agents in a virtual parallel space, including:

[0065] The virtual parallel simulation space construction module is used to collect static environmental data and dynamic situational data of the physical world, fuse and model the static environmental data and dynamic situational data, and construct a virtual parallel simulation space corresponding to the state of the physical world.

[0066] The adversarial agent modeling module is used to construct two adversarial agent models in a virtual parallel simulation space, which are respectively designated as the first agent and the second agent. The module configures the state-aware input set, action output set, resource constraints, and initial policy parameters for the first agent and the second agent.

[0067] The adversarial interaction simulation module is used to drive the first and second intelligent agents to perform autonomous decision-making and interaction actions according to the time step in a virtual parallel simulation space, and generate adversarial interaction sequences.

[0068] The adversarial evolution update module is used to generate adversarial trajectory data based on adversarial interaction sequences, and to perform adversarial evolution update on the policy parameters of the first agent and the second agent using an improved NFSP algorithm to generate updated policy parameters.

[0069] The strategy iteration simulation module is used to load the updated strategy parameters into the first and second agents, execute adversarial interactions in a virtual parallel simulation space, and generate an updated adversarial interaction sequence.

[0070] The adversarial result parsing module is used to parse the updated adversarial interaction sequence and generate adversarial result data;

[0071] The game verification module is used to quantify the adversarial result data based on a preset game verification index system and generate game verification results.

[0072] The strategy filtering and output module is used to filter the set of strategy parameters that meet the preset verification conditions based on the game verification results, and output the set of strategy parameters as the adversarial strategy results verified by the game.

[0073] In this embodiment, the modules are connected through the following method:

[0074] Collect static environmental data and dynamic situational data from the physical world, and fuse them into a virtual parallel simulation space through modeling.

[0075] In a virtual parallel simulation space, two adversarial intelligent agent models are constructed, designated as the first intelligent agent and the second intelligent agent, respectively, and configured with state-aware input sets, action output sets, resource constraints, and initial policy parameters.

[0076] Based on a virtual parallel simulation space, the first and second intelligent agents are driven to make autonomous decisions and interactive actions according to the time step, forming an adversarial interaction sequence.

[0077] Based on the state sequence, action sequence and adversarial result sequence in the adversarial interaction sequence, adversarial trajectory data is generated. An improved NFSP algorithm is used to perform adversarial evolutionary updates on the policy parameters of the first agent and the second agent to generate updated policy parameters.

[0078] The updated strategy parameters are loaded into the first and second agents, and the first and second agents are driven to perform adversarial interactions in a virtual parallel simulation space to generate an updated adversarial interaction sequence.

[0079] The result information of the interaction between the first agent and the second agent in the updated adversarial interaction sequence is analyzed to generate adversarial result data;

[0080] Based on a pre-set game verification index system, the adversarial result data is quantitatively calculated to generate game verification results.

[0081] The set of strategy parameters that meet the preset verification conditions is selected based on the game verification results, and the set of strategy parameters is output as the adversarial strategy result of the game verification.

[0082] In this embodiment, the construction process involves collecting static environmental data and dynamic situational data from the physical world. The static environmental data refers to environmental information in the physical world, and the dynamic situational data refers to situational information in the physical world that changes over time. The static environmental data and the dynamic situational data are time-correlated and fused to generate simulation environment data of physical world state changes at a unified time step. Based on the simulation environment data, a virtual operating environment corresponding to the physical world state is established in the computer system, forming a virtual parallel simulation space.

[0083] In this embodiment, the construction and configuration of the first and second intelligent agents specifically include:

[0084] Based on the environmental state of the virtual parallel simulation space, two adversarial intelligent agent models are constructed and designated as the first intelligent agent model and the second intelligent agent model, respectively. The construction process is as follows: in the virtual parallel simulation space, two independent decision model instances are established based on the same simulation environment state, and the same environment state acquisition method and interaction interface are configured for the two decision model instances, forming the first intelligent agent and the second intelligent agent that are adversarial in the same virtual parallel simulation space.

[0085] A state-aware input set is configured for the first intelligent agent and the second intelligent agent respectively. The state-aware input set refers to the environmental state information of the virtual parallel simulation space at the corresponding time step.

[0086] The configuration process is as follows: based on the environmental state information generated by the virtual parallel simulation space at each time step, input interfaces for receiving environmental state information are established for the first intelligent agent and the second intelligent agent respectively, and the environmental state information is mapped to the input interfaces according to a predetermined data structure to form the state-aware input set;

[0087] The first intelligent agent and the second intelligent agent are configured with action output sets respectively. The action output sets refer to the range of interactive actions that the first intelligent agent and the second intelligent agent can perform in the virtual parallel simulation space.

[0088] The configuration process is as follows: based on the interaction behavior types supported by the virtual parallel simulation space, action interfaces for outputting interactive actions are established for the first agent and the second agent respectively, and the interaction behavior types are mapped to the action interfaces to form the action output set;

[0089] Resource constraints are configured for the first intelligent agent and the second intelligent agent respectively. The resource constraints refer to limiting the resource usage range of the first intelligent agent and the second intelligent agent when performing interactive actions in the action output set.

[0090] The configuration process is as follows: based on the resource types and resource restriction rules available for interactive behavior in the virtual parallel simulation space, resource occupancy restrictions corresponding to interactive actions are set for the first agent and the second agent respectively, and the resource occupancy restrictions are associated with the interactive actions in the action output set to form the resource constraints.

[0091] Initial policy parameters are configured for the first and second intelligent agents respectively. The initial policy parameters refer to the interactive actions in the action output set generated based on the state-aware input set.

[0092] The configuration process specifically involves: establishing decision parameter storage for the first and second intelligent agents to perform autonomous decision-making, and writing the parameter values ​​that drive the autonomous decision-making into the decision parameter storage to form the initial strategy parameters.

[0093] In this embodiment, the formation of the adversarial interaction sequence specifically includes:

[0094] Arrange the spatial states of the virtual parallel simulation space into a state sequence in chronological order;

[0095] The spatial state is input into the first and second intelligent agents respectively to execute autonomous decision-making, generate interactive actions, and form an action sequence in chronological order. The generation of the interactive actions is to write the spatial state of the virtual parallel simulation space into the state perception input set of the first and second intelligent agents. The first and second intelligent agents respectively call the initial policy parameters to perform decision calculation on the spatial state and obtain the action selection result corresponding to the action output set.

[0096] At each time step, the interactive action is input into the virtual parallel simulation space to drive the spatial state update of the virtual parallel simulation space. The update process is to write the interactive action into the state update interface of the virtual parallel simulation space, call the environment state transition rules to calculate and process the current spatial state, generate the spatial state corresponding to the next time step, and write the spatial state corresponding to the next time step into the spatial state storage of the virtual parallel simulation space to complete the spatial state update.

[0097] Based on the updated spatial state and the corresponding interactive actions at each time step, the updated spatial state is calculated according to the adversarial result determination rules to obtain the adversarial result at each time step. The adversarial results obtained at each time step are recorded and collected in chronological order to form an adversarial result sequence.

[0098] The adversarial result determination rule refers to using the state difference between the state before and after the spatial state update at the same time step as the determination basis, and combining the first agent interaction action and the second agent interaction action corresponding to the time step, to map the state difference to the adversarial result corresponding to that time step.

[0099] The state sequence, action sequence, and adversarial result sequence are concatenated to form an adversarial interaction sequence.

[0100] In this embodiment, the generation of the updated strategy parameters specifically includes:

[0101] Adversarial trajectory data is constructed based on the state sequence, action sequence, and adversarial result sequence in the adversarial interaction sequence. An improved NFSP algorithm is used to perform adversarial evolutionary updates on the policy parameters of the first agent and the second agent. The improved NFSP algorithm includes an adversarial trajectory data sampling unit and a dual policy parameter update unit. The adversarial trajectory data sampling unit refers to performing sampling processing on the adversarial trajectory data to generate batches of sampled data. The dual policy parameter update unit refers to introducing a mirror adversarial update mechanism to perform adversarial evolutionary updates on the policy parameters of the first agent and the second agent and output the updated policy parameters.

[0102] The construction process is as follows: In the adversarial interaction sequence, the state sequence, action sequence, and adversarial result sequence are aligned according to the time step. The spatial state, interaction action, and adversarial result under the same time step are combined and associated with the spatial state under the next time step to form trajectory elements for the corresponding time step. The trajectory elements formed by each time step are collected in chronological order to obtain adversarial trajectory data.

[0103] In the adversarial trajectory data sampling unit, the adversarial trajectory data is sampled to generate a batch of sampled data. The sampling process involves indexing and marking the adversarial trajectory data, dividing the trajectory elements into sampling ranges according to the time step, selecting trajectory elements from the sampling range according to the sampling rules, and combining the selected trajectory elements in chronological order to form a batch of sampled data.

[0104] In the dual-policy parameter update unit, adversarial evolutionary updates are performed on the policy parameters of the first agent and the second agent based on the sampled data batch. The adversarial evolutionary update introduces a mirror adversarial update mechanism, which generates the update direction of the optimal response policy parameter and the update direction of the average policy parameter for the sampled data batch. The update direction is mapped to the mirror space to complete the parameter update. The updated parameters are then mapped from the mirror space back to the policy parameter space, and the updated policy parameters are output.

[0105] In this embodiment, the update process involves writing the updated policy parameters into the policy parameter storage of the first and second agents in the virtual parallel simulation space. The first and second agents then use the updated policy parameters as the current decision-making basis and drive the virtual parallel simulation space to run according to the time step. At each time step, the spatial state of the virtual parallel simulation space is input into the first and second agents. The first and second agents then execute autonomous decision-making output interaction actions, and these interaction actions are input into the virtual parallel simulation space to drive the spatial state update. The updated spatial state and the interaction actions are recorded in chronological order to form an updated adversarial interaction sequence.

[0106] In this embodiment, the generation of the adversarial result data specifically includes:

[0107] The updated adversarial interaction sequence is aligned to generate an aligned adversarial interaction sequence.

[0108] The alignment process specifically involves: marking the spatial state, interaction action, and interaction result information in the updated adversarial interaction sequence with time indexes according to the time step; completing records with missing time steps; merging multiple records with the same time step; so that the spatial state, interaction action, and interaction result information correspond one-to-one at each time step, thereby obtaining the aligned adversarial interaction sequence.

[0109] At each time step, the result information of the interaction between the first agent and the second agent is parsed based on the aligned adversarial interaction sequence.

[0110] The specific parsing process is as follows: at each time step, read the record of the corresponding time step in the aligned adversarial interaction sequence, extract the interaction action of the first agent, the interaction action of the second agent, and the spatial state update information of the virtual parallel simulation space from the interaction result information, and associate and aggregate the interaction action of the first agent, the interaction action of the second agent, and the spatial state update information according to the time step to obtain the interaction result information of the first agent and the second agent.

[0111] The state difference in the spatial state update information is extracted from the interaction result information at each time step. The state difference is the difference between the state before the spatial state update and the state after the spatial state update at the same time step. The state difference is used as the basis for judging the adversarial result. Combined with the interaction actions of the first agent and the interaction actions of the second agent, adversarial result records for the corresponding time steps are generated. The adversarial result records are collected in chronological order to obtain adversarial result data.

[0112] In this embodiment, the generation of the game verification result specifically includes:

[0113] Based on each adversarial result record in the adversarial result data, extract the adversarial result value corresponding to the first agent and the adversarial result value corresponding to the second agent respectively.

[0114] The extraction process involves reading each adversarial result record in the adversarial result data, locating the data field corresponding to the first agent and the data field corresponding to the second agent according to the field identifier in the record, reading the value of the data field corresponding to the first agent as the adversarial result value corresponding to the first agent, and reading the value of the data field corresponding to the second agent as the adversarial result value corresponding to the second agent.

[0115] The adversarial result values ​​corresponding to the first agent and the second agent are compared and statistically analyzed to generate the number of wins and the win rate of the first agent. The generation process is to read the adversarial result values ​​corresponding to the first agent and the second agent one by one in chronological order. When the adversarial result value corresponding to the first agent is greater than the adversarial result value corresponding to the second agent, the adversarial result record is counted as a win and the number of wins is accumulated. After the comparison of all adversarial result records is completed, the number of wins is divided by the total number of adversarial result records to obtain the win rate of the first agent.

[0116] The cumulative adversarial result value of the first agent is obtained by summing the adversarial result values ​​of the second agent, and the cumulative adversarial result value of the second agent is obtained by summing the adversarial result values ​​of the first agent and the second agent. The average value of the adversarial result difference is obtained by statistically calculating the difference between the adversarial result values ​​of the first agent and the second agent.

[0117] The game verification results are formed based on the number of wins, win rate, cumulative adversarial result value of the first agent, cumulative adversarial result value of the second agent, and the average of the adversarial result difference.

[0118] In this embodiment, the output of the adversarial strategy result specifically includes:

[0119] Based on the strategy parameters corresponding to the game verification results, a correspondence between strategy parameters and game verification results is formed. The process of forming this correspondence involves generating a unique identifier for each set of strategy parameters, writing the unique identifier into the corresponding game verification result, and using the unique identifier as an index during storage to associate and register the strategy parameters with the game verification results, thus forming a one-to-one correspondence between strategy parameters and game verification results.

[0120] For each set of strategy parameters in the corresponding relationship, a comprehensive calculation is performed on the mean of the number of wins, win rate, cumulative adversarial result value of the first agent, cumulative adversarial result value of the second agent, and adversarial result difference in the game verification results to generate a strategy parameter evaluation value corresponding to each set of strategy parameters.

[0121] The calculation process involves multiplying the number of wins and the win rate by the corresponding weight coefficients, subtracting the cumulative adversarial result value of the second agent from the cumulative adversarial result value of the first agent and then multiplying it by the corresponding weight coefficient, multiplying the mean of the adversarial result difference by the corresponding weight coefficient, and adding the above weighted terms to obtain the policy parameter evaluation value of the policy parameters.

[0122] The strategy parameter evaluation values ​​are matched with preset verification conditions to filter strategy parameters that meet the preset verification conditions, and these filters are then compiled into a strategy parameter set. This set of strategy parameters is used as the game-verified adversarial strategy result. The filtering process involves reading the corresponding strategy parameter evaluation value for each set of strategy parameters, comparing the evaluation value with the verification threshold in the preset verification conditions, and selecting the passed strategy parameter into the strategy parameter set when the evaluation value is greater than or equal to the verification threshold. When the evaluation value is less than the verification threshold, the failed strategy parameters are removed. After traversing all strategy parameters, the strategy parameter set is output as the game-verified adversarial strategy result.

[0123] Example 1:

[0124] In this embodiment, an unmanned patrol and adversarial simulation of a closed port area is used as the application scenario. The closed port area has infrastructure such as a fixed road network, fence boundaries, storage yard areas, berth areas, fixed camera poles, and communication base stations. Simultaneously, during operation within the area, there are dynamic situations such as changes in vehicle flow, temporary obstacles, target object positions, and wireless communication conditions. Existing technologies in this type of adversarial simulation typically employ a single simulation map or static rule-driven simulation method. The trained strategies are easily sensitive to environmental changes, and training and verification are disconnected. The strategy output lacks quantitative filtering criteria, making it difficult to reuse the simulation results in subsequent strategy iterations or situational reviews of different days. The problems to be solved in this embodiment are mainly reflected in: insufficient mapping between the simulation environment and the real situation leading to poor strategy generalization; difficulty in standardizing and aligning adversarial interaction data leading to unstable evaluation; and lack of closed-loop game verification leading to insufficient credibility of strategy output.

[0125] When applying the method of this invention in this scenario, static environmental data and dynamic situational data of the port park are first collected, and a virtual parallel simulation space is formed in the computer system. Static environmental data includes the park's road topology, fence boundaries, locations of fixed obstacles in the storage yard, berth area traffic rules, and fixed equipment location information; dynamic situational data includes real-time location changes of inspection objects, changes in moving obstacles, changes in communication quality, and changes in vehicle flow lines within the park. After temporal correlation and fusion processing of the static environmental data and dynamic situational data, simulation environment data is generated at a unified time step, and a virtual operating environment corresponding to the physical world state is established accordingly, forming a virtual parallel simulation space. This allows subsequent adversarial simulations to no longer rely on a single static map, but can update the spatial state as the situation changes.

[0126] Subsequently, two adversarial agent models are constructed in a virtual parallel simulation space, designated as Agent 1 and Agent 2, respectively. In this embodiment, Agent 1 corresponds to the "inspection and protection side," whose objective is to inspect key areas within the park and block abnormal approach behaviors. Agent 2 corresponds to the "intrusion and interference side," whose objective is to approach and traverse key areas undetected or unblocked. A state-aware input set is configured for Agent 1 and Agent 2, enabling them to receive the spatial state of the virtual parallel simulation space at each time step; an action output set is configured, limiting their interactive actions to the range of executable behaviors within the park; resource constraints are configured, restricting resource consumption during action execution; and initial policy parameters are configured for each agent, enabling them to generate interactive actions from the action output set based on the state-aware input set. Through these configurations, the perception boundaries, behavioral boundaries, and resource boundaries of Agent 1 and Agent 2 are ensured to be consistent and controllable, avoiding the incomparability problem caused by the unclear behavioral range of agents in existing simulations.

[0127] During operation in the virtual parallel simulation space, the first and second agents are driven to perform autonomous decisions and interactive actions according to time steps, forming an adversarial interaction sequence. At each time step, the virtual parallel simulation space outputs a spatial state, which is written into the state-aware input sets of the first and second agents respectively. The first and second agents respectively call the initial policy parameters to perform decision calculations on the spatial state, obtaining the action selection results corresponding to the action output set as the interactive actions. The interactive actions are input into the virtual parallel simulation space, driving the spatial state update of the virtual parallel simulation space, and recording the state sequence, action sequence, and adversarial result sequence, ultimately forming the adversarial interaction sequence. This adversarial interaction sequence serves as the direct basis for the subsequent construction of adversarial trajectory data, ensuring that the data link is traceable and reproducible.

[0128] In the strategy update phase, adversarial trajectory data is constructed based on the state sequence, action sequence, and adversarial result sequence in the adversarial interaction sequence. The adversarial trajectory data is aligned by time steps, combining spatial states, interaction actions, and adversarial results at the same time step, and associating them with the spatial states at the next time step to form trajectory elements. These are then aggregated in chronological order to obtain the adversarial trajectory data. In the improved NFSP algorithm, the adversarial trajectory data sampling unit performs sampling processing on the adversarial trajectory data to generate sampled data batches. During the sampling process, trajectory elements are indexed and marked, and the sampling range is divided according to the time step. Then, trajectory elements are selected from the sampling range according to the sampling rules and combined in chronological order to form sampled data batches. Subsequently, in the dual-strategy parameter update unit, adversarial evolutionary updates are performed based on the sampled data batches. The strategy parameters include optimal response strategy parameters and average strategy parameters. Update directions for the optimal response strategy parameters and average strategy parameters are generated for the sampled data batches. A mirror adversarial update mechanism is introduced to map the update directions to the mirror space to complete the parameter update. The updated parameters are then mapped back from the mirror space to the strategy parameter space, and the updated strategy parameters are output. The updated policy parameters are loaded into the first and second agents, and adversarial interactions are performed again in the virtual parallel simulation space to generate an updated adversarial interaction sequence, thus forming a closed loop between policy updates and adversarial inference.

[0129] In the adversarial result analysis phase, the updated adversarial interaction sequence is analyzed to generate adversarial result data. The updated adversarial interaction sequence is first aligned, and spatial state, interaction actions, and interaction result information are time-indexed according to time steps. Missing time step records are completed, and multiple records at the same time step are merged, ensuring a one-to-one correspondence between spatial state, interaction actions, and interaction result information at each time step, resulting in an aligned adversarial interaction sequence. Then, based on the aligned adversarial interaction sequence, the result information of the interaction between the first and second agents is analyzed, extracting the first agent's interaction actions, the second agent's interaction actions, and spatial state update information, and aggregating them according to time steps. The state difference extracted from the spatial state update information at each time step is used as the basis for determining the adversarial result, and combined with the first and second agent's interaction actions to generate adversarial result records for the corresponding time steps, thus obtaining the adversarial result data. This adversarial result data is input into a game validation index system for quantitative calculation, generating game validation results, which are further used for strategy selection and output.

[0130] This embodiment uses two consecutive days of operational status data from the same port area for verification. The first day is used to form a virtual parallel simulation space mapping the initial strategy parameters to the environment. The second day is used for adversarial simulation and strategy selection under conditions of significant dynamic changes. A fixed number of adversarial interaction sequences are generated for each set of strategy parameters to form adversarial result data and calculate the game verification results. The game verification results include at least the number of wins, win rate, cumulative adversarial result value of the first agent, cumulative adversarial result value of the second agent, and the mean of the adversarial result difference. Based on the weight coefficients, a comprehensive calculation is performed to obtain the strategy parameter evaluation value, which is then compared with the verification threshold to obtain the strategy parameter set. For comparative verification, this embodiment sets the control schemes as "improved NFSP algorithm implementation without introducing a mirror adversarial update mechanism" and "static simulation environment implementation without using a virtual parallel simulation space". Both maintain the same action output set and resource constraints for the first and second agents to ensure comparability. Statistical results show that the present invention can maintain a higher win rate and a more stable mean difference in adversarial results under dynamic situational changes. At the same time, it can output a more stable set of strategy parameters with a higher pass rate during the strategy selection stage, demonstrating the advantages of higher reliability of strategy output and more effective verification loop.

[0131] Table 1. Statistics on the Verification and Strategy Selection of the Port Industrial Park's Confrontational Game Theory

[0132] Indicator Item Static simulation environment implementation Improved NFSP does not introduce a mirror-based adversarial update mechanism Method of the present invention Total number of adversarial outcome records (items) 1200 1200 1200 Number of times the first agent wins (times) 642 706 804 First agent's win rate 0.535 0.588 0.670 The first agent's cumulative adversarial result value 18240 19680 22410 The cumulative adversarial result value of the second agent 17610 18690 20220 Mean of the difference in adversarial outcomes 0.525 0.825 1.825 Number of strategy parameter groups (groups) 40 40 40 The number of policy parameters (groups) used to validate the threshold. 11 16 24 Strategy parameter set pass rate 0.275 0.400 0.600 Mean of strategy parameter evaluation value 0.61 0.74 0.92 Standard deviation of strategy parameter evaluation values 0.21 0.18 0.12 Number of missing completion attempts in the updated adversarial interaction sequence parsing (times) 37 22 8

[0133] As can be seen from the number of wins and win rate in Table 1, under the same total number of adversarial result records, the first agent implemented in this invention wins 804 times, with a win rate of 0.670, which is significantly higher than 0.535 achieved in the static simulation environment and also higher than the control scheme without the mirror adversarial update mechanism (0.588). This indicates that in the dynamically changing port park scenario, the spatial state formed based on the virtual parallel simulation space can more stably support the autonomous decision-making of the first agent. Furthermore, after the mirror adversarial update mechanism participates in the dual-strategy parameter update, the strategy parameters exhibit stronger adversarial effectiveness during the adversarial update process, enabling the first agent to achieve a higher win rate under the same interaction constraints.

[0134] The cumulative adversarial result value and the average difference between adversarial results show that the cumulative adversarial result value of the first agent implemented in this invention is higher than that of the two control schemes. Furthermore, the difference between the cumulative adversarial result values ​​of the first agent and the second agent is larger at the mean level, with the average difference in adversarial results reaching 1.825. While the control schemes can also improve adversarial performance to some extent, the improvement is limited under dynamic conditions. These results indicate that this invention constructs adversarial trajectory data through adversarial interaction sequences and performs adversarial evolutionary updates driven by batch sampling data. This makes the iterative direction of the strategy parameters more focused on the cumulative advantage of the adversarial gain difference, thereby improving the strategy's ability to generate sustained gains during the adversarial process.

[0135] The strategy selection results show that the number of strategy parameters passing the verification threshold in this invention reaches 24, with a pass rate of 0.600, significantly higher than the 0.275 achieved in the static simulation environment and the 0.400 of another control scheme. More importantly, the mean and standard deviation of the strategy parameter evaluation values ​​are higher, indicating that the strategy parameter evaluation values ​​calculated through the game-theoretic verification index system have better stability in this invention, and the strategy set output is more concentrated and controllable. This contrasts with the problems of fragmented training and verification and difficulty in stable strategy selection in existing technologies, demonstrating that this invention has stronger quantifiable verification capabilities and higher policy output credibility in the closed-loop chain of "adversarial evolutionary update—game-theoretic verification—strategy selection output".

[0136] The updated adversarial interaction sequence parsing missing data completion count shows that the number of missing data completions achieved by this invention is significantly lower than the control scheme. This indicates that the adversarial interaction sequences formed in the virtual parallel simulation space have better data consistency in the time step alignment and parsing stages, reducing the dependence on missing record completion and multi-record merging in the parsing stage. This makes the generation process of adversarial result data more stable, thus providing a more reliable data foundation for the quantitative calculation of game verification results. This further demonstrates the effectiveness of this invention in solving the existing technical defect of "difficulty in standardizing and aligning adversarial interaction data, leading to unstable evaluation."

[0137] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A simulation and game-theoretic verification system for adversarial evolution of intelligent agents in a virtual parallel space, characterized in that, include: The virtual parallel simulation space construction module is used to collect static environmental data and dynamic situational data of the physical world, fuse and model the static environmental data and dynamic situational data, and construct a virtual parallel simulation space corresponding to the state of the physical world. The adversarial agent modeling module is used to construct two adversarial agent models in a virtual parallel simulation space, which are respectively designated as the first agent and the second agent. The module configures the state-aware input set, action output set, resource constraints, and initial policy parameters for the first agent and the second agent. The adversarial interaction simulation module is used to drive the first and second intelligent agents to perform autonomous decision-making and interaction actions according to the time step in a virtual parallel simulation space, and generate adversarial interaction sequences. The adversarial evolution update module is used to generate adversarial trajectory data based on adversarial interaction sequences, and to perform adversarial evolution update on the policy parameters of the first agent and the second agent using an improved NFSP algorithm to generate updated policy parameters. The strategy iteration simulation module is used to load the updated strategy parameters into the first and second agents, execute adversarial interactions in a virtual parallel simulation space, and generate an updated adversarial interaction sequence. The adversarial result parsing module is used to parse the updated adversarial interaction sequence and generate adversarial result data; The game verification module is used to quantify the adversarial result data based on a preset game verification index system and generate game verification results. The strategy filtering and output module is used to filter the set of strategy parameters that meet the preset verification conditions based on the game verification results, and output the set of strategy parameters as the adversarial strategy results verified by the game.

2. The simulation and game-theoretic verification system for adversarial evolution of intelligent agents in a virtual parallel space according to claim 1, characterized in that, The modules are connected in the following way: Collect static environmental data and dynamic situational data from the physical world, and fuse them into a virtual parallel simulation space through modeling. In a virtual parallel simulation space, two adversarial intelligent agent models are constructed, designated as the first intelligent agent and the second intelligent agent, respectively, and configured with state-aware input sets, action output sets, resource constraints, and initial policy parameters. Based on a virtual parallel simulation space, the first and second intelligent agents are driven to make autonomous decisions and interactive actions according to the time step, forming an adversarial interaction sequence. Based on the state sequence, action sequence and adversarial result sequence in the adversarial interaction sequence, adversarial trajectory data is generated. An improved NFSP algorithm is used to perform adversarial evolutionary updates on the policy parameters of the first agent and the second agent to generate updated policy parameters. The updated strategy parameters are loaded into the first and second agents, and the first and second agents are driven to perform adversarial interactions in a virtual parallel simulation space to generate an updated adversarial interaction sequence. The result information of the interaction between the first agent and the second agent in the updated adversarial interaction sequence is analyzed to generate adversarial result data; Based on a pre-set game verification index system, the adversarial result data is quantitatively calculated to generate game verification results. The set of strategy parameters that meet the preset verification conditions is selected based on the game verification results, and the set of strategy parameters is output as the adversarial strategy result of the game verification.

3. The simulation and game-theoretic verification system for adversarial evolution of intelligent agents in a virtual parallel space according to claim 2, characterized in that, The construction process involves collecting static environmental data and dynamic situational data from the physical world. The static environmental data refers to environmental information in the physical world, and the dynamic situational data refers to situational information in the physical world that changes over time. The static environmental data and the dynamic situational data are time-correlated and fused to generate simulation environment data of physical world state changes at a unified time step. Based on the simulation environment data, a virtual operating environment corresponding to the physical world state is established in the computer system, forming a virtual parallel simulation space.

4. The simulation and game-theoretic verification system for adversarial evolution of intelligent agents in a virtual parallel space according to claim 2, characterized in that, The construction and configuration of the first and second intelligent agents specifically include: Based on the environmental state of the virtual parallel simulation space, two adversarial intelligent agent models are constructed and designated as the first intelligent agent model and the second intelligent agent model, respectively. The construction process is as follows: in the virtual parallel simulation space, two independent decision model instances are established based on the same simulation environment state, and the same environment state acquisition method and interaction interface are configured for the two decision model instances, forming the first intelligent agent and the second intelligent agent that are adversarial in the same virtual parallel simulation space. A state-aware input set is configured for the first intelligent agent and the second intelligent agent respectively. The state-aware input set refers to the environmental state information of the virtual parallel simulation space at the corresponding time step. The first intelligent agent and the second intelligent agent are configured with action output sets respectively. The action output sets refer to the range of interactive actions that the first intelligent agent and the second intelligent agent can perform in the virtual parallel simulation space. Resource constraints are configured for the first intelligent agent and the second intelligent agent respectively. The resource constraints refer to limiting the resource usage range of the first intelligent agent and the second intelligent agent when performing interactive actions in the action output set. Initial policy parameters are configured for the first and second intelligent agents respectively. The initial policy parameters refer to the interactive actions in the action output set generated based on the state-aware input set.

5. The simulation and game-theoretic verification system for adversarial evolution of intelligent agents in a virtual parallel space according to claim 2, characterized in that, The formation of the adversarial interaction sequence specifically includes: Arrange the spatial states of the virtual parallel simulation space into a state sequence in chronological order; The spatial state is input into the first and second intelligent agents respectively to execute autonomous decision-making, generate interactive actions, and form an action sequence in chronological order. The generation of the interactive actions is to write the spatial state of the virtual parallel simulation space into the state perception input set of the first and second intelligent agents. The first and second intelligent agents respectively call the initial policy parameters to perform decision calculation on the spatial state and obtain the action selection result corresponding to the action output set. At each time step, the interactive action is input into the virtual parallel simulation space to drive the spatial state update of the virtual parallel simulation space. The update process is to write the interactive action into the state update interface of the virtual parallel simulation space, call the environment state transition rules to calculate and process the current spatial state, generate the spatial state corresponding to the next time step, and write the spatial state corresponding to the next time step into the spatial state storage of the virtual parallel simulation space to complete the spatial state update. Based on the updated spatial state and the corresponding interactive actions at each time step, the updated spatial state is calculated according to the adversarial result determination rules to obtain the adversarial result at each time step. The adversarial results obtained at each time step are recorded and collected in chronological order to form an adversarial result sequence. The state sequence, action sequence, and adversarial result sequence are concatenated to form an adversarial interaction sequence.

6. The simulation and game-theoretic verification system for adversarial evolution of intelligent agents in a virtual parallel space according to claim 2, characterized in that, The generation of the updated strategy parameters specifically includes: Adversarial trajectory data is constructed based on the state sequence, action sequence, and adversarial result sequence in the adversarial interaction sequence. An improved NFSP algorithm is used to perform adversarial evolutionary updates on the policy parameters of the first agent and the second agent. The improved NFSP algorithm includes an adversarial trajectory data sampling unit and a dual policy parameter update unit. The adversarial trajectory data sampling unit refers to performing sampling processing on the adversarial trajectory data to generate batches of sampled data. The dual policy parameter update unit refers to introducing a mirror adversarial update mechanism to perform adversarial evolutionary updates on the policy parameters of the first agent and the second agent and output the updated policy parameters. In the adversarial trajectory data sampling unit, the adversarial trajectory data is sampled to generate a batch of sampled data. The sampling process involves indexing and marking the adversarial trajectory data, dividing the trajectory elements into sampling ranges according to the time step, selecting trajectory elements from the sampling range according to the sampling rules, and combining the selected trajectory elements in chronological order to form a batch of sampled data. In the dual-policy parameter update unit, adversarial evolutionary updates are performed on the policy parameters of the first agent and the second agent based on the sampled data batch. The adversarial evolutionary update introduces a mirror adversarial update mechanism, which generates the update direction of the optimal response policy parameter and the update direction of the average policy parameter for the sampled data batch. The update direction is mapped to the mirror space to complete the parameter update. The updated parameters are then mapped from the mirror space back to the policy parameter space, and the updated policy parameters are output.

7. The simulation and game-theoretic verification system for adversarial evolution of intelligent agents in a virtual parallel space according to claim 2, characterized in that, The update process involves writing the updated policy parameters into the policy parameter storage of the first and second agents in the virtual parallel simulation space. The first and second agents then use the updated policy parameters as the current decision-making basis and drive the virtual parallel simulation space to run according to the time step. At each time step, the spatial state of the virtual parallel simulation space is input into the first and second agents. The first and second agents then execute autonomous decision-making output interaction actions, and these interaction actions are input into the virtual parallel simulation space to drive the spatial state update. The updated spatial state and the interaction actions are recorded in chronological order to form an updated adversarial interaction sequence.

8. The simulation and game-theoretic verification system for adversarial evolution of intelligent agents in a virtual parallel space according to claim 2, characterized in that, The generation of the adversarial outcome data specifically includes: The updated adversarial interaction sequence is aligned to generate an aligned adversarial interaction sequence. At each time step, the result information of the interaction between the first agent and the second agent is parsed based on the aligned adversarial interaction sequence. The state difference in the spatial state update information is extracted from the interaction result information at each time step. The state difference is the difference between the state before the spatial state update and the state after the spatial state update at the same time step. The state difference is used as the basis for judging the adversarial result. Combined with the interaction actions of the first agent and the interaction actions of the second agent, adversarial result records for the corresponding time steps are generated. The adversarial result records are collected in chronological order to obtain adversarial result data.

9. The simulation and game-theoretic verification system for adversarial evolution of intelligent agents in a virtual parallel space according to claim 2, characterized in that, The generation of the game verification result specifically includes: Based on each adversarial result record in the adversarial result data, extract the adversarial result value corresponding to the first agent and the adversarial result value corresponding to the second agent respectively. The adversarial result values ​​corresponding to the first agent and the second agent are compared and statistically analyzed to generate the number of wins and the win rate of the first agent. The generation process is to read the adversarial result values ​​corresponding to the first agent and the second agent one by one in chronological order. When the adversarial result value corresponding to the first agent is greater than the adversarial result value corresponding to the second agent, the adversarial result record is counted as a win and the number of wins is accumulated. After the comparison of all adversarial result records is completed, the number of wins is divided by the total number of adversarial result records to obtain the win rate of the first agent. The cumulative adversarial result value of the first agent is obtained by summing the adversarial result values ​​of the second agent, and the cumulative adversarial result value of the second agent is obtained by summing the adversarial result values ​​of the first agent and the second agent. The average value of the adversarial result difference is obtained by statistically calculating the difference between the adversarial result values ​​of the first agent and the second agent. The game verification results are formed based on the number of wins, win rate, cumulative adversarial result value of the first agent, cumulative adversarial result value of the second agent, and the average of the adversarial result difference.

10. The simulation and game-theoretic verification system for adversarial evolution of intelligent agents in a virtual parallel space according to claim 2, characterized in that, The output of the adversarial strategy results specifically includes: Based on the strategy parameters corresponding to the game verification results, a correspondence between strategy parameters and game verification results is formed. The process of forming this correspondence involves generating a unique identifier for each set of strategy parameters, writing the unique identifier into the corresponding game verification result, and using the unique identifier as an index during storage to associate and register the strategy parameters with the game verification results, thus forming a one-to-one correspondence between strategy parameters and game verification results. For each set of strategy parameters in the corresponding relationship, a comprehensive calculation is performed on the mean of the number of wins, win rate, cumulative adversarial result value of the first agent, cumulative adversarial result value of the second agent, and adversarial result difference in the game verification results to generate a strategy parameter evaluation value corresponding to each set of strategy parameters. The strategy parameter evaluation values ​​are matched with preset verification conditions to filter strategy parameters that meet the preset verification conditions, and these filters are then compiled into a strategy parameter set. This set of strategy parameters is used as the game-verified adversarial strategy result. The filtering process involves reading the corresponding strategy parameter evaluation value for each set of strategy parameters, comparing the evaluation value with the verification threshold in the preset verification conditions, and selecting the passed strategy parameter into the strategy parameter set when the evaluation value is greater than or equal to the verification threshold. When the evaluation value is less than the verification threshold, the failed strategy parameters are removed. After traversing all strategy parameters, the strategy parameter set is output as the game-verified adversarial strategy result.