Intelligent agent training method and device based on FTS platform, terminal equipment and storage medium

Through the FTS platform-based agent training method, the physical component information of the agent is obtained, the status and action space are determined, the combat strategy is designed and the reward function is trained, which solves the problem of insufficient number of deep reinforcement learning platforms for multiple agents, and improves the collaboration and adaptability of the intelligent system in complex environments.

CN120373413APending Publication Date: 2025-07-25ARMOR ACADEMY OF CHINESE PEOPLES LIBERATION ARMY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510443997.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In the existing combat simulation systems, there are insufficient number of deep reinforcement learning platforms for multi-agents, making it difficult to achieve efficient collaboration and task allocation of heterogeneous agents in complex environments. Especially in urban combat scenarios, the simulation degree is insufficient, making it difficult to meet the needs of large-scale heterogeneous groups intelligent games.

Method used

Based on the FTS platform, by obtaining the physical component information of the agent, determining the state space and action space of the first and second agents, designing a battle strategy, and improving the adaptability of the intelligent system in complex environments through reward function training strategies.

Benefits of technology

The knowledge transfer and generalization capabilities of intelligent systems in complex environments have been improved, and the collaboration efficiency and adaptability of heterogeneous agents in urban combat scenarios have been enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373413A_ABST
    Figure CN120373413A_ABST
Patent Text Reader

Abstract

The invention discloses an agent training method and device based on an FTS platform, terminal equipment and a storage medium, and the method comprises the steps: obtaining the physical component information of agents, and determining a first agent and a second agent according to a scene combat demand and the physical component information; determining action space information of the first intelligent agent according to the state space information of the first intelligent agent, and determining action space information of the second intelligent agent according to the state space information of the second intelligent agent; determining a battle strategy according to the action space information of the first intelligent agent and the action space information of the second intelligent agent; the method comprises the following steps: executing a battle strategy to obtain a reward function, training the battle strategy according to the reward function, and continuously improving the long-term adaptive capacity of an intelligent system to a complex environment through interaction among a physical agent, a virtual agent and the environment. And the knowledge migration and generalization ability of the intelligent agent in a complex environment is more emphasized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of artificial intelligence technology, and particularly relates to an agent training method, device, terminal device and storage medium based on the FTS platform. Background Art

[0002] "Learning war from war" is the most fundamental method for studying the laws of war. However, with the rapid transformation of military technology, especially the transformation of the war form brought about by new weapons and equipment, traditional war experience has become difficult to guide new wars. "Moving from the laboratory to the future battlefield" combat simulation has become the mainstream. Computer Generated Force (CGF) refers to a virtual combat force object created by a computer that can autonomously control or direct all or part of its actions and behaviors. It can be understood as a type of Agent object oriented to combat scenarios and is a core component in combat simulation modeling.

[0003] Urban combat in the land battlefield is still the only means to achieve direct seizure and control, and often plays a decisive role in the war situation. Aiming at the characteristics of insufficient simulation degree of combat forces and combat environment in the multi-agent simulation process, an intelligent command and control entity that can support the combat command training at both the single-equipment and formation levels is constructed based on the application scenario of urban seizure and control.

[0004] Multi-agent controlled vehicles move and strike in the town, with the goal of eliminating all enemy units in the town terrain. The simulation platform is the core component of the combat experiment system. The simulation experiment environment is crucial for algorithm research. By means of unified experimental settings and standardized evaluation indicators, fair comparison between algorithms is realized, which becomes a key tool for moving from theoretical research to large-scale practical applications. However, as a relatively new research field, the number of existing open-source platforms for multi-agent deep reinforcement learning is significantly less than that of deep reinforcement learning.

[0005] In the complex environment of the real world, agents often need to possess different capabilities and undertake different roles, which requires them to play their roles in their respective areas of expertise while collaborating efficiently with other types of agents. In this context, the research on heterogeneous agent collaboration becomes particularly important. The challenges of heterogeneous agent collaboration not only involve how to make agents with different capabilities and resources work together, but also include how to efficiently allocate tasks and resources among different agents, and how to adjust the collaboration strategy to cope with environmental changes. An effective heterogeneous agent collaboration strategy can significantly improve the efficiency and adaptability of the entire system, enabling multi-agent deep reinforcement learning to be applied in a wider range of fields.

[0006] Currently, the decision-making behavior modeling in existing combat simulation systems mainly includes the following methods from two dimensions: human-like / rational and behavior acquisition methods:

[0007] 1. Based on sufficient military domain knowledge, directly describe the external decision-making logic of related tasks, such as finite state machines, script rules, etc. The acquisition of its behavior largely belongs to the category of knowledge engineering, does not involve complex internal reasoning, and emphasizes acting like a human.

[0008] 2. Combine military domain knowledge with a general cognitive architecture to simulate the cognitive process of the brain, conduct internal reasoning and decision-making, and generate decision outputs, such as SOAR, ACT-R, etc., emphasizing thinking like a human.

[0009] 3. Based on theoretical methods such as planning and decision-making in artificial intelligence, approximately implement the human decision-making and reasoning process, such as Bayesian networks, goal-oriented planning, Monte Carlo tree search, etc. It usually assumes that human behavior is fully rational and optimized, emphasizing acting reasonably.

[0010] In the case of considering real-world situations such as continuous three-dimensional space, complex dynamic constraints, limited observations, and complex disturbances, it is very challenging to achieve large-scale heterogeneous swarm intelligence games. How to meet the intelligent games of large-scale heterogeneous swarms is an urgent problem to be solved currently. Summary of the Invention

[0011] This application aims to provide an intelligent agent training method, device, terminal device, and storage medium based on the FTS platform to solve the deficiencies in the prior art. The technical problems to be solved by this application are achieved through the following technical solutions.

[0012] In a first aspect, an embodiment of this application provides an intelligent agent training method based on the FTS platform. The method includes:

[0013] Obtain the physical component information of the intelligent agent, where the physical component information at least includes an actuator, a sensor, a communication device, and a weapon;

[0014] Determine a first intelligent agent and a second intelligent agent according to the scenario combat requirements and the physical component information;

[0015] Determine the action space information of the first intelligent agent according to the state space information of the first intelligent agent, and determine the action space information of the second intelligent agent according to the state space information of the second intelligent agent;

[0016] Determine a battle strategy according to the action space information of the first intelligent agent and the action space information of the second intelligent agent;

[0017] Execute the battle strategy to obtain a reward function, and train the battle strategy according to the reward function.

[0018] Optionally, the state space information of the first agent includes at least longitude, latitude, heading, health, detection range, the number of detected enemy units, the azimuth angle of the target area, the relative distance to the target area, the relative azimuth angle of other agents, and the relative distance to other agents; the state space information of the second agent includes at least longitude, latitude, heading, and type.

[0019] Optionally, determining the battle strategy according to the action space information of the first agent and the action space information of the second agent includes:

[0020] According to the action space information of the first agent and the action space information of the second agent, determine the movement strategy of the first agent and the detection strategy of the second agent, where the movement strategy includes at least a movement range, a movement direction, and a movement distance.

[0021] Optionally, determining the movement strategy of the first agent and the detection strategy of the second agent according to the action space information of the first agent and the action space information of the second agent includes:

[0022] When the first agent moves to a preset area in a preset manner, determine that the second agent uses an operations research optimization algorithm to select an optimal detection area from multiple sub-areas, where the preset area and the multiple sub-areas are all battle sub-areas.

[0023] Optionally, the reward function includes at least a stage reward function and a final reward function, and the final reward is the number of detected enemy units.

[0024] Optionally, the stage reward is determined according to the reward for the distance to the target area and the distance penalty to other agents, and the final reward is determined according to the number of detected enemy units and the time penalty.

[0025] In a second aspect, an agent training device based on the FTS platform provided by an embodiment of the present application includes:

[0026] An acquisition module, configured to acquire physical component information of an agent, where the physical component information includes at least a mover, a sensor, a communication device, and a weapon;

[0027] A determination module, configured to determine a first agent and a second agent according to the scenario combat requirements and the physical component information;

[0028] A setting module, configured to determine the action space information of the first agent according to the state space information of the first agent, and determine the action space information of the second agent according to the state space information of the second agent;

[0029] A battle module, configured to determine a battle strategy according to the action space information of the first agent and the action space information of the second agent;

[0030] A training module, configured to execute the battle strategy to obtain a reward function, and train the battle strategy according to the reward function.

[0031] Optionally, the state space information of the first agent at least includes longitude, latitude, heading, blood volume, detection range, the number of detected enemy units, the azimuth angle of the target area, the relative distance of the target area, the relative azimuth angle of other agents, and the relative distance from other Agents; the state space information of the second agent at least includes longitude, latitude, heading, and type.

[0032] Optionally, the battle module is configured to:

[0033] Determine a movement strategy of the first agent and a detection strategy of the second agent according to the action space information of the first agent and the action space information of the second agent, where the movement strategy at least includes a movement range, a movement direction, and a movement distance.

[0034] Optionally, the battle module is configured to:

[0035] When the first agent moves to a preset area in a preset manner, determine that the second agent uses an operations research optimization algorithm to select an optimal detection area from multiple sub-areas, where the preset area and the multiple sub-areas are all battle sub-areas.

[0036] Optionally, the reward function at least includes a stage reward function and a final reward function, and the final reward is the number of detected enemy units.

[0037] Optionally, the stage reward is determined according to the reward for the distance to the target area and the distance penalty from other agents, and the final reward is determined according to the number of detected enemy units and the time penalty.

[0038] In a third aspect, an embodiment of the present application provides a terminal device, including: at least one processor and a memory;

[0039] The memory stores a computer program; the at least one processor executes the computer program stored in the memory to implement the intelligent agent training method based on the FTS platform provided in the first aspect.

[0040] Fourthly, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed, it implements the intelligent agent training method based on the FTS platform provided in the first aspect.

[0041] The embodiments of the present application include the following advantages:

[0042] The intelligent agent training method, device, terminal device and storage medium based on the FTS platform provided by the embodiments of the present application obtain the physical component information of the intelligent agent, where the physical component information at least includes a mover, a sensor, a communication device and a weapon; determine the first intelligent agent and the second intelligent agent according to the scenario combat requirements and the physical component information; determine the action space information of the first intelligent agent according to the state space information of the first intelligent agent, and determine the action space information of the second intelligent agent according to the state space information of the second intelligent agent; determine the battle strategy according to the action space information of the first intelligent agent and the action space information of the second intelligent agent; execute the battle strategy to obtain a reward function, and train the battle strategy according to the reward function. In the embodiments of the present application, according to the scenario combat requirements, the characteristics of controlling troops and weapons and equipment, the intelligent agent neural network structure, initial parameters, action space, state space, reward function, etc. are designed. Through the interaction between physical intelligent agents, virtual intelligent agents and the environment, the long-term adaptability of the intelligent system to complex environments is continuously improved. Compared with traditional artificial intelligence algorithms, it emphasizes more on the knowledge transfer and generalization ability of intelligent agents in complex environments. Description of the Drawings

[0043] In order to more clearly illustrate the embodiments of the present application or the existing technical solutions, the following will briefly introduce the drawings required for use in the description of the embodiments or the existing technical solutions. Obviously, the drawings in the following description are only some embodiments recorded in the present application. For those of ordinary skill in the art, other drawings can be obtained according to these drawings without creative efforts.

[0044] Figure 1 It is a flowchart of an intelligent agent training method based on the FTS platform in an embodiment of the present application;

[0045] Figure 2 It is a flowchart of another intelligent agent training method based on the FTS platform in an embodiment of the present application

[0046] Figure 3 It is the relationship of mover model classes in an embodiment of the present application;

[0047] Figure 4 It is the relationship of sensor model classes in an embodiment of the present application;

[0048] Figure 5 This is the class relationship of the communication device model in an embodiment of the present application;

[0049] Figure 6 This is the class relationship of the weapon model in an embodiment of the present application;

[0050] Figure 7 This is the schematic diagram of the confrontation strategy of the agent in an embodiment of the present application;

[0051] Figure 8 This is the schematic diagram of the confrontation strategy of another agent in an embodiment of the present application;

[0052] Figure 9 This is the schematic diagram of the update of the confrontation strategy of the agent in an embodiment of the present application;

[0053] Figure 10 This is the schematic diagram of the training effect in an embodiment of the present application;

[0054] Figure 11 This is the schematic diagram of the training effect in an embodiment of the present application;

[0055] Figure 12 This is the schematic diagram of the training effect in an embodiment of the present application;

[0056] Figure 13 This is the structural block diagram of an embodiment of the agent training device based on the FTS platform of the present application;

[0057] Figure 14 This is the structural schematic diagram of a terminal device of the present application. Detailed implementation manners

[0058] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with specific embodiments and the corresponding drawings. Apparently, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.

[0059] An embodiment of the present application provides an agent training method based on the FTS platform for training multiple agents. The execution subject of this embodiment is an agent training device based on the FTS platform, which is set on a terminal device. For example, the terminal device at least includes a computer terminal, etc.

[0060] The platform modeling adopts an object-oriented inheritance method. The platform type has a general definition. The sub-platform type can inherit attributes from the parent platform type, and unique characteristics are assigned in the sub-platform type definition. All FTS platforms are instantiations of the platform type.

[0061] The modeling method of FTS (Intelligent Union Fight Training System) is objective and scientific, reflecting the intelligent behavior of the war subject, and includes five basic components: the modeling of movers, sensors, communication devices, processors, and weapons.

[0062] Motion modeling is attached to the platform to maintain the platform's dynamic state and can model the six-degree-of-freedom motion state of the aircraft.

[0063] Sensor modeling creates trajectory metrics and forwards trajectory information through the link, and can model antennas, transmitters, and receivers.

[0064] Communication device modeling uses external links to forward or receive messages, and can model the physical links of the communication system using transmitters, receivers, and antennas, and can simulate communication protocols such as the network layer, transport layer, and physical layer, as well as the data types transmitted by wired or wireless devices.

[0065] Processor modeling uses a behavior modeling method based on the RIPR framework to define behaviors or computational algorithms similar to the human brain or computer, and most processors can be defined through scenario scripts.

[0066] Weapon modeling can model the combat effectiveness, fuzes, warheads, etc. of weapons, and interfere with, suppress, or destroy their targets, such as missiles, bombs, artillery, lasers, jammers, etc.

[0067] The flexible model interaction architecture of FTS includes sensor detection, communication, command and control, engagement, etc. The model interaction architecture is shown in Figure 2 As shown, the platform components include physical components and intellectual / computational components. The physical components include movers, communication devices, sensors, and weapons. The intellectual / computational component is the processor, where the processor interacts with the physical components through internal links, and the communication device interacts with other platforms through external links. The platform supports four types of information, namely perception, task, trajectory, and measurement information. The platform attributes include target characteristics and attributes defined by additional data.

[0068] FTS uses four architectural elements (attributes, elements, components, and links) to describe each platform model in the simulation:

[0069] (1) Attributes. Include general data, such as platform name, type, and affiliation. This sub-element can be extended to include mission-unique information such as radar, optical, and infrared signature data to determine whether the aircraft is vulnerable to detection by enemy sensors.

[0070] (2) Information. It includes the data resident on the platform and detailed information on how the recipient perceives this data. For an aircraft, this will include the types of data to be displayed to the pilot (altitude, speed, heading, radar indication, etc.) and the large amount of raw data driving these displays.

[0071] (3) Components. Composed of various models that directly control the behavior of the platform. These models describe how the platform moves in space and time, perceives its surrounding environment, processes the information it collects, communicates with other platforms, and uses its kinetic and non-kinetic weapon arsenals against hostile platforms, as well as performs various other tasks. Finally, the link element coordinates the data exchange between the various subsystems on the platform and the communication with other platforms. Communicates with other platforms, uses its kinetic and non-kinetic weapon arsenals against adversary platforms, and performs various other tasks.

[0072] (4) Link. Coordinates the data exchange between the various subsystems on the platform and the communication with other platforms. Communicates with other platforms, uses its kinetic and non-kinetic weapon arsenals against adversary platforms, and performs various other tasks. Finally, the link element coordinates the data exchange between the various subsystems on the platform and the communication with other platforms.

[0073] Based on the model components, different platform models can be pieced together. The entity interacts with the outside through the command and control link constructed by the communication component.

[0074] Refer to Figure 1 , which shows the step flowchart of an embodiment of the agent training method based on the FTS platform of the present application. The method may specifically include the following steps:

[0075] S101. Obtain the physical component information of the agent, where the physical component information at least includes a mover, a sensor, a communication device, and a weapon;

[0076] Specifically, the terminal device is the processor as in Figure 2 . This processor can obtain the physical component information of the physical components, that is, it can obtain the physical component information of the mover, the sensor, the communication device, the weapon, etc., and based on this physical component information, model the corresponding physical devices, that is, generate virtual components corresponding to the physical components in the manner shown in Figure 3 . The virtual component at least includes a sensor model, a communication device model, a weapon model, a processor model, an electronic protection model, and an electronic attack model.

[0077] Among them, Figure 3 is the mover model class relationship. The mover model class mainly includes a sensor model, a communication device model, a weapon model, a processor model, an electronic protection model, an electronic attack model, etc. The sensor model class relationship is shown in Figure 4As shown, for the class relationship of the communication device model, see Figure 5 As shown, for the class relationship of the weapon model, see Figure 6 As shown.

[0078] S102. Determine the first agent and the second agent according to the scenario combat requirements and physical component information;

[0079] Specifically, the processor on the terminal device generates corresponding virtual components based on the physical component information, and determines the first agent and the second agent according to the scenario combat requirements and the virtual components. The first agent and the second agent are used for combat. In the embodiments of the present application, the first agent is the friendly side, and the second agent is the enemy side.

[0080] S103. Determine the action space information of the first agent according to the state space information of the first agent, and determine the action space information of the second agent according to the state space information of the second agent;

[0081] Specifically, the processor on the terminal device obtains the state space information of the first agent, such as longitude, latitude, heading, blood volume, detection range, the number of enemy units detected, etc., and also obtains the state space information of the second agent, such as longitude, latitude, heading, and type; during the combat process, the friendly side, that is, the first agent, not only needs to obtain its own state space information, but also needs to obtain the state space information of other friendly agents and the state space information of the enemy.

[0082] The terminal device determines the action space information of the first agent according to the state space information of the first agent, that is, according to the state space information of the first agent, such as longitude, latitude, heading, blood volume, detection range, the number of enemy units detected, etc., determines which area the first agent moves to, as well as the moving direction and moving distance, and determines the corresponding action space information of the second agent corresponding to the action space information of the first agent.

[0083] S104. Determine the combat strategy according to the action space information of the first agent and the action space information of the second agent;

[0084] The terminal device determines the combat strategy according to the action space information of the first agent and the action space information of the second agent, that is, when the friendly agent performs a certain operation, the enemy agent gives corresponding countermeasures.

[0085] S105. Execute the combat strategy to obtain a reward function, and train the combat strategy according to the reward function.

[0086] The terminal device controls the virtual component to execute the battle strategy. The purpose of this battle strategy is to detect the largest number of enemy units. For each step of the execution, a stage reward is calculated. After the execution is completed, the final reward is calculated, that is, the final reward is the number of detected enemy units. In order to detect the largest number of enemy units, the above battle strategy is continuously adjusted, and finally the optimal agent is obtained.

[0087] Another embodiment of this application further supplements and explains the agent training method based on the FTS platform provided in the above embodiment.

[0088] The 3v3 multi-agent task confrontation space discretizes a large area into 9 sub-areas according to the street environment. The red side units include the 6 types of units mentioned in the indicators. Except for the UAV agent, the units of the red and blue sides in the rest of the scenarios are symmetric, only the initial positions are different. The red side is initialized at the lower side of the map, and the blue side is initialized at the upper side of the map. The task of the red side UAV agent is to detect all the blue side units, and the task objectives of the other 5 types of multi-agents are to eliminate all the blue side units; the positioning of the cooperative unmanned aerial vehicle (UAV) and the unmanned ground vehicle (UGV) is a complex technology, involving the data fusion of multiple sensors and multi-robot cooperation. UAVs and UGVs are usually equipped with different types of sensors, such as GPS, inertial measurement unit (IMU), vision sensors, etc.

[0089] Optionally, the state space information of the first agent at least includes longitude, latitude, heading, blood volume, detection range, the number of detected enemy units, the azimuth of the target area, the relative distance of the target area, the relative azimuth of other agents, and the relative distance from other Agents; as shown in Table 1, the state space information of the second agent at least includes longitude, latitude, heading, and type, as shown in Table 2.

[0090] Table 1 Characteristics of friendly units

[0091]

[0092] Table 2 Characteristics of enemy units

[0093] Name Type Meaning lon PlainFeature Longitude lat PlainFeature Latitude direction PlainFeature Course type PlainFeature Type

[0094] Optionally, according to the action space information of the first agent and the action space information of the second agent, determine the battle strategy, including:

[0095] According to the action space information of the first agent and the action space information of the second agent, determine the movement strategy of the first agent and the detection strategy of the second agent, where the movement strategy at least includes the movement range, movement direction, and movement distance.

[0096] Exemplarily, the terminal device determines the moving direction, which is discretized into 8 directions with due north as the reference. The angle between every two directions is 360 / 8 = 45°. Among them, 0° is the due north direction, and it rotates clockwise, as Figures 7 - 9 shown.

[0097] Optionally, according to the action space information of the first agent and the action space information of the second agent, determine the moving strategy of the first agent and the detection strategy of the second agent, including:

[0098] When the first agent moves to the preset area in a preset manner, determine that the second agent uses the operation research optimization algorithm to select the optimal detection area from multiple sub-areas. Among them, the preset area and multiple sub-areas are all combat sub-areas. As shown in Table 3, for various task scenarios, such as: UGV: unmanned ground vehicle; UGV_small: small unmanned vehicle; UGV_mid: medium unmanned vehicle; UGV_big: large unmanned vehicle; Tank: tank; Attackvehicle: fast attack vehicle.

[0099] Table 3

[0100]

[0101]

[0102] Optionally, the reward function includes at least a stage reward function and a final reward function, and the final reward is the number of enemy units discovered.

[0103] Optionally, the stage reward is determined according to the reward for the distance to the target area and the distance penalty from other agents, and the final reward is determined according to the number of enemy units discovered and the time penalty.

[0104] Specifically, since the types, functions, and characteristics of the friendly units are the same, only the quantities are different, a multi-agent training architecture with shared network parameters is adopted, and the friendly unit information is shared, including not only the enemy unit information but also the information of the friendly units other than itself. Since the task objectives of the three agents are the same, the reward functions are the same;

[0105] The reward function consists of two parts, the final reward and the stage reward. First, the final reward is the number of enemy units discovered. Since the task is to detect enemy units as much as possible, the more discovered, the greater the reward. Second, the stage reward is the distance penalty from other agents. The idea of this design is to prevent the distances between multiple agents from being too close as much as possible, so as to increase the reconnaissance coverage rate of the drones. α is the learning rate in the formula.

[0106] Reward = R final + Rstep

[0107] R final = α × number of enemy units detected × time penalty

[0108] R step = α × (reward for distance to target area + distance penalty to other agents)

[0109] The training effect of the agent is as Figures 10 - 12 shown

[0110] The agent training method based on the FTS platform provided by the embodiments of the present application obtains the physical component information of the agent, where the physical component information at least includes a mover, a sensor, a communication device, and a weapon; determines a first agent and a second agent according to the scenario combat requirements and the physical component information; determines the action space information of the first agent according to the state space information of the first agent, and determines the action space information of the second agent according to the state space information of the second agent; determines a battle strategy according to the action space information of the first agent and the action space information of the second agent; executes the battle strategy to obtain a reward function, and trains the battle strategy according to the reward function. In the embodiments of the present application, according to the scenario combat requirements, the control of troops, and the characteristics of weapons and equipment, the neural network structure and initial parameters, action space, state space, reward function, etc. of the agent are designed. Through the interaction between physical agents, virtual agents, and the environment, the long-term adaptability of the intelligent system to complex environments is continuously improved. Compared with traditional artificial intelligence algorithms, it emphasizes more on the knowledge transfer and generalization ability of agents in complex environments.

[0111] Another embodiment of the present application provides an agent training device based on the FTS platform, which is used to execute the agent training method based on the FTS platform provided by the above embodiment.

[0112] Referring to Figure 13 , a structural block diagram of an embodiment of an agent training device based on the FTS platform of the present application is shown. The device may specifically include the following modules: an acquisition module 1301, a determination module 1302, a setting module 1303, a battle module 1304, and a training module 1305, where:

[0113] The acquisition module 1301 is used to acquire the physical component information of the agent, where the physical component information at least includes a mover, a sensor, a communication device, and a weapon;

[0114] The determination module 1302 is used to determine a first agent and a second agent according to the scenario combat requirements and the physical component information;

[0115] The setting module 1303 is configured to determine the action space information of the first agent according to the state space information of the first agent, and determine the action space information of the second agent according to the state space information of the second agent;

[0116] The battle module 1304 is configured to determine a battle strategy according to the action space information of the first agent and the action space information of the second agent;

[0117] The training module 1305 is configured to execute the battle strategy to obtain a reward function, and train the battle strategy according to the reward function.

[0118] Another embodiment of the present application further supplements the intelligent agent training device based on the FTS platform provided in the above embodiment.

[0119] Optionally, the state space information of the first agent at least includes longitude, latitude, heading, blood volume, detection range, the number of detected enemy units, the azimuth angle of the target area, the relative distance of the target area, the relative azimuth angle of other agents, and the relative distance from other Agents; the state space information of the second agent at least includes longitude, latitude, heading, and type.

[0120] Optionally, the battle module is configured to:

[0121] Determine the movement strategy of the first agent and the detection strategy of the second agent according to the action space information of the first agent and the action space information of the second agent, where the movement strategy at least includes a movement range, a movement direction, and a movement distance.

[0122] Optionally, the battle module is configured to:

[0123] When the first agent moves to a preset area in a preset manner, determine that the second agent uses an operations research optimization algorithm to select an optimal detection area from multiple sub-areas, where the preset area and the multiple sub-areas are all battle sub-areas.

[0124] Optionally, the reward function at least includes a stage reward function and a final reward function, and the final reward is the number of discovered enemy units.

[0125] Optionally, the stage reward is determined according to the reward for the distance to the target area and the distance penalty from other agents, and the final reward is determined according to the number of discovered enemy units and the time penalty.

[0126] For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and for the related parts, please refer to the partial description of the method embodiment.

[0127] The intelligent agent training device based on the FTS platform provided by the embodiments of the present application obtains the physical component information of the intelligent agent, where the physical component information at least includes a mover, a sensor, a communication device, and a weapon; determines a first intelligent agent and a second intelligent agent according to the scenario combat requirements and the physical component information; determines the action space information of the first intelligent agent according to the state space information of the first intelligent agent, and determines the action space information of the second intelligent agent according to the state space information of the second intelligent agent; determines a confrontation strategy according to the action space information of the first intelligent agent and the action space information of the second intelligent agent; executes the confrontation strategy to obtain a reward function, and trains the confrontation strategy according to the reward function. In the embodiments of the present application, according to the scenario combat requirements, the characteristics of controlling troops and weaponry, the intelligent agent neural network structure, initial parameters, action space, state space, reward function, etc. are designed. Through the interaction between physical intelligent agents, virtual intelligent agents, and the environment, the long-term adaptability of the intelligent system to complex environments is continuously improved. Compared with traditional artificial intelligence algorithms, it emphasizes more on the knowledge transfer and generalization ability of intelligent agents in complex environments.

[0128] Another embodiment of the present application provides a terminal device for executing the intelligent agent training method based on the FTS platform provided in the above embodiment.

[0129] Figure 14 is a schematic structural diagram of a terminal device of the present application, as Figure 14 shown, the terminal device includes: at least one processor 1401 and a memory 1402;

[0130] The memory stores a computer program; at least one processor executes the computer program stored in the memory to implement the intelligent agent training method based on the FTS platform provided in the above embodiment.

[0131] The terminal device provided in this embodiment obtains the physical component information of the intelligent agent, where the physical component information at least includes a mover, a sensor, a communication device, and a weapon; determines the first intelligent agent and the second intelligent agent according to the scenario combat requirements and the physical component information; determines the action space information of the first intelligent agent according to the state space information of the first intelligent agent, and determines the action space information of the second intelligent agent according to the state space information of the second intelligent agent; determines the battle strategy according to the action space information of the first intelligent agent and the action space information of the second intelligent agent; executes the battle strategy to obtain a reward function, and trains the battle strategy according to the reward function. In the embodiment of the present application, according to the scenario combat requirements, the characteristics of controlling troops and weapons and equipment, the neural network structure and initial parameters, action space, state space, reward function, etc. of the intelligent agent are designed. Through the interaction between physical intelligent agents, virtual intelligent agents, and the environment, the long-term adaptation ability of the intelligent system to complex environments is continuously improved. Compared with traditional artificial intelligence algorithms, it emphasizes more on the knowledge transfer and generalization ability of intelligent agents in complex environments.

[0132] Another embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed, it implements the intelligent agent training method based on the FTS platform provided in any of the above embodiments.

[0133] According to the computer-readable storage medium of this embodiment, it obtains the physical component information of the intelligent agent, where the physical component information at least includes a mover, a sensor, a communication device, and a weapon; determines the first intelligent agent and the second intelligent agent according to the scenario combat requirements and the physical component information; determines the action space information of the first intelligent agent according to the state space information of the first intelligent agent, and determines the action space information of the second intelligent agent according to the state space information of the second intelligent agent; determines the battle strategy according to the action space information of the first intelligent agent and the action space information of the second intelligent agent; executes the battle strategy to obtain a reward function, and trains the battle strategy according to the reward function. In the embodiment of the present application, according to the scenario combat requirements, the characteristics of controlling troops and weapons and equipment, the neural network structure and initial parameters, action space, state space, reward function, etc. of the intelligent agent are designed. Through the interaction between physical intelligent agents, virtual intelligent agents, and the environment, the long-term adaptation ability of the intelligent system to complex environments is continuously improved. Compared with traditional artificial intelligence algorithms, it emphasizes more on the knowledge transfer and generalization ability of intelligent agents in complex environments.

[0134] It should be noted that the above detailed description is exemplary and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.

[0135] It should be noted that the terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly dictates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they specify the presence of the features, steps, operations, devices, components, and / or combinations thereof.

[0136] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that these terms can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order different from those illustrated or described herein.

[0137] In addition, the terms "comprise" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0138] For ease of description, spatial relative terms such as "above", "on top of", "on the upper surface", "over", etc. may be used herein to describe the spatial positional relationship of one device or feature to other devices or features as shown in the figures. It should be understood that the spatial relative terms are intended to encompass different orientations in use or operation in addition to the orientation depicted in the figures. For example, if the device in the figures is inverted, the device described as "above" or "on top of" other devices or structures will then be positioned "below" or "beneath" other devices or structures. Thus, the exemplary term "above" can include both the orientation of "above" and "below". The device can also be positioned in other different ways, such as rotated 90 degrees or in other orientations, and the corresponding explanations for the spatial relative descriptions used herein will be made.

[0139] In the detailed description above, reference has been made to the drawings which form a part hereof. In the drawings, like reference numerals typically identify like components, unless the context indicates otherwise. The illustrated embodiments described in the detailed specification, drawings, and claims are not meant to be limiting. Other embodiments may be used and other changes may be made without departing from the spirit or scope of the subject matter presented herein.

[0140] The above are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. An intelligent agent training method based on the FTS platform, characterized in that, The method includes: Obtaining the physical component information of the agent, where the physical component information at least includes a mover, a sensor, a communication device, and a weapon; Determining a first agent and a second agent according to the scenario combat requirements and the physical component information; Determining the action space information of the first agent according to the state space information of the first agent, and determining the action space information of the second agent according to the state space information of the second agent; Determining a battle strategy according to the action space information of the first agent and the action space information of the second agent; Executing the battle strategy to obtain a reward function, and training the battle strategy according to the reward function.

2. The intelligent agent training method based on the FTS platform according to claim 1, wherein, The state space information of the first agent at least includes longitude, latitude, heading, blood volume, detection range, the number of detected enemy units, the azimuth angle of the target area, the relative distance of the target area, the relative azimuth angle of other agents, and the relative distance from other agents; the state space information of the second agent at least includes longitude, latitude, heading, and type.

3. The intelligent agent training method based on the FTS platform according to claim 2, characterized in that The determining the battle strategy according to the action space information of the first agent and the action space information of the second agent includes: Determining the movement strategy of the first agent and the detection strategy of the second agent according to the action space information of the first agent and the action space information of the second agent, where the movement strategy at least includes a movement range, a movement direction, and a movement distance.

4. The intelligent agent training method based on the FTS platform according to claim 3, characterized in that, The determining the movement strategy of the first agent and the detection strategy of the second agent according to the action space information of the first agent and the action space information of the second agent includes: When the first agent moves to a preset area in a preset manner, determining that the second agent uses an operations research optimization algorithm to select an optimal detection area from multiple sub-areas, where the preset area and the multiple sub-areas are all battle sub-areas.

5. The intelligent agent training method based on the FTS platform according to claim 1, wherein, The reward function at least includes a stage reward function and a final reward function, and the final reward is the number of detected enemy units.

6. The intelligent agent training method based on the FTS platform according to claim 5, wherein The stage reward is determined according to the reward for the distance to the target area and the distance penalty from other agents, and the final reward is determined according to the number of detected enemy units and the time penalty.

7. An intelligent agent training device based on the FTS platform, characterized in that, The device includes: An acquisition module, configured to acquire the physical component information of the agent, where the physical component information at least includes a mover, a sensor, a communication device, and a weapon; A determination module, configured to determine a first agent and a second agent according to the scenario combat requirements and the physical component information; A setting module, configured to determine the action space information of the first agent according to the state space information of the first agent, and determine the action space information of the second agent according to the state space information of the second agent; A battle module, configured to determine a battle strategy according to the action space information of the first agent and the action space information of the second agent; A training module, configured to execute the battle strategy to obtain a reward function, and train the battle strategy according to the reward function.

8. The intelligent agent training device based on the FTS platform according to claim 7, wherein the state space information of the first intelligent agent at least includes longitude, latitude, heading, blood volume, detection range, the number of detected enemy units, the azimuth angle of the target area, the relative distance of the target area, the relative azimuth angle of other intelligent agents, and the relative distance from other Agents; the state space information of the second intelligent agent at least includes longitude, latitude, heading, and type.

9. A terminal device, characterized in that, Comprising: at least one processor and a memory; the memory stores a computer program; the at least one processor executes the computer program stored in the memory to implement the intelligent agent training method based on the FTS platform according to any one of claims 1-6.

10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and when the computer program is executed, it implements the intelligent agent training method based on the FTS platform according to any one of claims 1-6.

Citation Information

Cited By

  • Intelligent agent training management method, system and equipment

    CN121526274A