Intelligent agent control method and device, terminal equipment and medium
By obtaining and processing battlefield situation information and using neural network models for agent control, the problem of comprehensive control of agents in distributed land warfare systems is solved, and efficient combat command and simulation training is achieved.
Patent Information
- Application Number
- CN202510216702.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-17
AI Technical Summary
It is difficult for the prior art to effectively control agents in distributed land warfare systems, especially in different scenarios, how to perform comprehensive control is an urgent issue.
By obtaining battlefield situation information, feature extraction and fusion are performed, pre-trained neural network-based intelligent model is used for training, battlefield decision information is generated, and control instructions are generated based on this.
It realizes comprehensive control of the intelligent body, can support single-equipped and formation-level combat command training, provides flexible battlefield situation display, and improves the intelligence level of combat command simulation training.
Smart Images

Figure CN120163048A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of artificial intelligence technology, and particularly relates to a control method, device, terminal device and medium for an intelligent agent. Background Art
[0002] Based on the characteristics of the future land battlefield, the challenges faced by land operations are mainly the "sponge-like" combat environment and the "sand-like" combat objects. The traditional centralized combat system is tantamount to "fighting in the open against the hidden" and "using the large to fight the small", and will inevitably suffer heavy losses and it is difficult to obtain the expected military operation benefits. The wide-area decentralized deployment of smaller combat forces has become an important direction for the development of land operations. The "layered immersion" combat concept in the multi-domain battlefield space can enable the combat force formation to transform from "centralized" to "distributed", and is expected to promote the generation of new combat capabilities that meet future combat requirements.
[0003] In the above context, the need for technical research on an intelligent command and control system for a distributed land combat system has been drawn out. Based on typical application scenarios such as urban combat, airport seizure and mountain combat, how to comprehensively control intelligent agents in combination with different scenarios is an urgent problem to be solved at present. Summary of the Invention
[0004] This application aims to provide a control method, device, terminal device and storage medium for an intelligent agent to solve the deficiencies in the prior art. The technical problems to be solved by this application are achieved through the following technical solutions.
[0005] In a first aspect, an embodiment of this application provides a control method for an intelligent agent, and the method includes:
[0006] Obtain battlefield situation information, where the battlefield situation information at least includes our situation information and the opponent's situation information;
[0007] Extract features from the battlefield situation information to obtain a feature vector corresponding to the battlefield situation information;
[0008] Perform fusion processing on the feature vector to obtain a fused feature vector;
[0009] Train a pre-trained intelligent agent model according to the fused feature vector to obtain battlefield decision information; the pre-trained intelligent agent model is based on a neural network;
[0010] Generate a control instruction corresponding to the battlefield decision information according to the battlefield decision information.
[0011] Optionally, the pre-trained intelligent agent model at least includes a single-unit intelligent agent model and a formation intelligent agent model. The formation intelligent agent model is used to output a command strategy for controlling the formation according to the battlefield situation information.
[0012] Optionally, training the pre-trained agent model based on the fused feature vector to obtain battlefield decision-making information, including:
[0013] Inputting the fused feature vector into the intelligent decision-making algorithm model to evaluate and make decisions on the overall battlefield situation;
[0014] The intelligent decision-making algorithm model includes a situation feature extraction unit, a situation evaluation unit, and an inference and decision-making unit. The situation feature extraction unit, the situation evaluation unit, and the inference and decision-making unit are deep neural network models; the situation evaluation unit is used to evaluate the overall situation through a value model according to the overall situation of the current situation; the situation feature extraction unit is used to transform the fused feature vector to obtain an input vector that meets the requirements of the neural network model;
[0015] The inference and decision-making unit is used to train the neural network model according to the input vector to generate the battlefield decision-making information, and the battlefield decision-making information includes at least movement tasks and strike tasks.
[0016] Optionally,
[0017] Generating control instructions corresponding to the battlefield decision-making information according to the battlefield decision-making information, including:
[0018] Decoding the battlefield decision-making information into a target selection instruction, a sensor selection instruction, and an action selection instruction corresponding to the battlefield decision-making information;
[0019] At each decision-making time, determining the next battlefield decision-making information according to the current battlefield decision-making information and the next battlefield situation information.
[0020] Optionally, the method further includes:
[0021] During the neural network training process, setting a reward function in combination with the simulation environment;
[0022] Specifically including:
[0023] Constructing the reward function in a way that takes the global reward as the main and the local reward as the supplement. The local reward includes at least the exploration area reward, the relative distance penalty, and the firing reward; the global reward includes at least the friendly force loss penalty, the enemy force loss reward, and the time penalty.
[0024] Optionally, the friendly force situation information includes at least general features such as speed, heading, and blood volume, vector features such as its own position, weapons, relative positions of friendly forces, and relative positions of enemy forces, unique features such as the friendly force's affiliation and the enemy force's affiliation, and a terrain grid map as a spatial feature of a preset size.
[0025] In a second aspect, an embodiment of the present application provides a control device for an intelligent agent, the device comprising:
[0026] An acquisition module, configured to acquire battlefield situation information, where the battlefield situation information at least includes our situation information and the opponent's situation information;
[0027] An extraction module, configured to perform feature extraction on the battlefield situation information to obtain a feature vector corresponding to the battlefield situation information;
[0028] A fusion module, configured to perform a fusion process on the feature vectors to obtain a fused feature vector;
[0029] A training module, configured to train a pre-trained intelligent agent model according to the fused feature vector to obtain battlefield decision information; the pre-trained intelligent agent model is obtained based on a neural network;
[0030] A generation module, configured to generate a control instruction corresponding to the battlefield decision information according to the battlefield decision information.
[0031] Optionally, the pre-trained intelligent agent model at least includes a single-unit intelligent agent model and a formation intelligent agent model, and the formation intelligent agent model is configured to output a command strategy for controlling the formation according to the battlefield situation information.
[0032] Optionally, the training module is configured to:
[0033] Input the fused feature vector into an intelligent decision-making algorithm model to evaluate and make decisions on the overall battlefield situation;
[0034] The intelligent decision-making algorithm model includes a situation feature extraction unit, a situation evaluation unit, and an inference and decision-making unit, and the situation feature extraction unit, the situation evaluation unit, and the inference and decision-making unit are deep neural network models; the situation evaluation unit is configured to evaluate the overall situation through a value model according to the overall situation of the current situation; the situation feature extraction unit is configured to convert the fused feature vector to obtain an input vector that meets the requirements of the neural network model;
[0035] The inference and decision-making unit is configured to train the neural network model according to the input vector to generate the battlefield decision information, and the battlefield decision information at least includes movement tasks and strike tasks.
[0036] Optionally, the generation module is configured to:
[0037] Decode the battlefield decision information into a target selection instruction, a sensor selection instruction, and an action selection instruction corresponding to the battlefield decision information;
[0038] At each decision-making moment, determine the next battlefield decision-making information based on the current battlefield decision-making information and the next battlefield situation information.
[0039] Optionally, the training module is configured to:
[0040] During the neural network training process, combine the simulation environment and set the reward function;
[0041] Specifically, it includes:
[0042] Construct the reward function in a way that takes the global reward as the main part and the local reward as the auxiliary part. The local reward includes at least the exploration area reward, the relative distance penalty, and the firing reward; the global reward includes at least the friendly damage penalty, the enemy damage reward, and the time penalty.
[0043] Optionally, the friendly situation information includes at least the speed, heading, and blood volume as general features, the own position, weapon, relative position of the friendly side, and relative position of the enemy side as vector features, the friendly side and the enemy side to which the own side belongs as unique features, and the terrain grid map as a spatial feature of a preset size.
[0044] In a third aspect, an embodiment of the present application provides a terminal device, including: at least one processor and a memory;
[0045] The memory stores a computer program; the at least one processor executes the computer program stored in the memory to implement the control method of the intelligent agent provided in the first aspect.
[0046] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed, it implements the control method of the intelligent agent provided in the first aspect.
[0047] The embodiments of the present application have the following advantages:
[0048] The control method, device, terminal device and medium of an intelligent agent provided by an embodiment of the present application obtain battlefield situation information, where the battlefield situation information at least includes our side's situation information and the opponent's situation information; extract features from the battlefield situation information to obtain a feature vector corresponding to the battlefield situation information; perform fusion processing on the feature vector to obtain a fused feature vector; train a pre-trained intelligent agent model according to the fused feature vector to obtain battlefield decision information; the pre-trained intelligent agent model is obtained based on a neural network; according to the battlefield decision information, generate a control instruction corresponding to the battlefield decision information. The embodiment of the present application constructs an intelligent command and control entity capable of supporting single-unit and formation-level combat command training through an artificial intelligence algorithm. By combining a rule-based decision-making method with deep reinforcement learning technology, an artificial intelligence training platform that supports intelligent training, simulation deduction, and integrates key technologies such as deep learning and reinforcement learning is built. Further, a virtual intelligent agent for command decision-making is generated, providing intelligent opponents at different levels for combat command simulation training, and can provide flexible battlefield situation displays according to different scenarios, and intuitively display the evaluation results of the analysis and evaluation module. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0050] Figure 1 It is a flowchart of a control method for an intelligent agent in an embodiment of the present application;
[0051] Figure 2 It is a schematic diagram of the composition of an intelligent agent model in an embodiment of the present application;
[0052] Figure 3 It is a schematic diagram of the network structure framework of an intelligent agent in an embodiment of the present application;
[0053] Figure 4 It is a schematic diagram of the network structure of a spatial feature encoder in an embodiment of the present application;
[0054] Figure 5 It is a schematic diagram of the network structure of an entity feature encoder in an embodiment of the present application;
[0055] Figure 6 It is a schematic diagram of the network structure of a general feature encoder in an embodiment of the present application;
[0056] Figure 7Schematic diagram of the Dense feature aggregator network structure in an embodiment of the present application;
[0057] Figure 8 Schematic diagram of the LSTM feature aggregator network structure in an embodiment of the present application;
[0058] Figure 9 Schematic diagram of the GRU feature aggregator network structure in an embodiment of the present application;
[0059] Figure 10 Schematic diagram of the discrete action decoder network structure in an embodiment of the present application;
[0060] Figure 11 Schematic diagram of the ordered unit selection decoder network structure in an embodiment of the present application;
[0061] Figure 12 Schematic diagram of the unordered unit selection decoder network structure in an embodiment of the present application;
[0062] Figure 13 Schematic diagram of the single unit selection decoder network structure in an embodiment of the present application;
[0063] Figure 14 Schematic diagram of the situation input module structure in an embodiment of the present application;
[0064] Figure 15 Schematic diagram of the decision result generation module structure in an embodiment of the present application;
[0065] Figure 16 Schematic diagram of the control instruction output module structure in an embodiment of the present application;
[0066] Figure 17 Schematic diagram of the model neural network structure in an embodiment of the present application;
[0067] Figure 18 Schematic diagram of the situation extraction module structure in an embodiment of the present application;
[0068] Figure 19 Schematic diagram of the decision result generation module structure in an embodiment of the present application;
[0069] Figure 20 Schematic diagram of the reward function design in an embodiment of the present application;
[0070] Figure 21 Schematic diagram of the neural network structure in an embodiment of the present application;
[0071] Figure 22 Schematic diagram of the overall architecture of the large-scale distributed training technology in an embodiment of the present application;
[0072] Figure 23Schematic diagram of the operation process of a large-scale distributed training architecture in an embodiment of the present application;
[0073] Figure 24 Schematic diagram of task success rate prediction in an embodiment of the present application;
[0074] Figure 25 Schematic diagram of horizontal comparison of combat winning rates of different agents in an embodiment of the present application;
[0075] Figure 26 Demonstration interface of the real-time state of confrontation in an embodiment of the present application;
[0076] Figure 27 Screenshot of real-time winning rate estimation in an embodiment of the present application;
[0077] Figure 28 Block diagram of the structure of an embodiment of a control device for an agent of the present application;
[0078] Figure 29 Schematic diagram of the structure of a terminal device of the present application. Specific embodiments
[0079] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with specific embodiments and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the scope of protection of the present application.
[0080] An embodiment of the present application provides a control method for an agent, which is used to control the agent. The execution subject of this embodiment is a control device for the agent, which is set on a terminal device. For example, the terminal device at least includes a computer terminal, etc.
[0081] The agent model takes the deep reinforcement learning algorithm as the core, constructs a deep neural network, and adopts large-scale distributed training technology to generate a single-agent model and a formation agent model with a certain level of intelligence to support the human-machine / machine-machine confrontation demonstration in the confrontation deduction environment. The agent model mainly includes four parts: agent construction, agent call interface, agent training, and agent deployment. As Figure 2 shown, it includes: agent construction, agent call interface, agent training, and agent deployment.
[0082] The embodiments of the present application have the following advantages: The scale of the simulation environment is large, supporting the strategy training of a large-scale group; There are various types of environments, which can simulate various complex factors in the real scenario, such as the interaction of agents based on the physics engine, etc.; It not only supports the reinforcement learning training of a single group, but also supports multiple teams to perform policy iteration simultaneously; It can quickly build a new simulation environment. The training framework and training algorithms are easy to learn and maintain; The framework is reasonably designed, with less computing memory occupancy when training large-scale group strategies; The visualization effect is good, and it can intuitively display the learning effect of the group strategy; There is a complete and unified training result archiving and plotting module.
[0083] The agent network structure provides a programming framework for users to quickly build single and formation agents, mainly composed of three modules: a feature encoder, a feature aggregator, and a feature decoder. Users can customize the neural network structure of the agent based on this framework. The framework composition is as Figure 3 shown.
[0084] I. Feature Encoder
[0085] The feature encoder provides encoders for processing spatial features, entity features, and general features, and is used for feature encoding or feature extraction of various categories. The feature encoder includes a spatial feature encoder, an entity feature encoder, and a general feature encoder.
[0086] (1) Spatial Feature Encoder
[0087] The spatial feature encoder is mainly used to extract high-dimensional matrix feature information such as images, such as extracting the information contained in the current battlefield environment photo. Its network structure consists of an input layer, an intermediate layer, and an output layer. The input layer uses a fully connected neural network, the intermediate layer uses a convolutional neural network or a residual network, and the output layer uses a fully connected neural network. Among them, the number of hidden layers and the number of neurons in the output layer of the spatial feature encoder are both designed based on the network representation ability. The network structure of the spatial feature encoder is as Figure 4 shown.
[0088] Users can customize the spatial feature encoder, including customizing the number of neurons in the output layer, the number of layers of the convolutional neural network and the residual network.
[0089] (2) Entity Feature Encoder
[0090] The entity feature encoder is mainly used to extract entity feature information, such as unit features such as the model and status of the equipment platform. It first processes the processed features through a fully connected neural network once, then uses a transformer for further feature extraction, and finally performs pooling sampling. Among them, the number of hidden layers and the number of neurons in the output layer of the entity feature encoder are both designed based on the network representation ability. The network structure of the entity feature encoder is as Figure 5as shown
[0091] Users can customize the entity feature encoder, including the number of custom fully connected layers, whether to use transformers, the type of pooling method used, etc.
[0092] (3) General feature encoder
[0093] The general feature encoder is mainly used to extract general feature information, such as statistical information in the environment like the number of equipment survivors. It mainly uses a multi-layer fully connected neural network to extract feature information. Among them, the number of hidden layers and the number of neurons in the output layer of the general feature encoder are both designed based on the network representation ability. The network structure of the general feature encoder is as Figure 6 shown
[0094] Users can customize the general feature encoder, including the number of hidden layers and the number of neurons in the output layer.
[0095] II. Feature aggregator
[0096] An aggregator that provides feature aggregation for spatial features, entity features, and general features. The feature aggregator is used to aggregate the features processed by different feature encoders and splice them into complete information. It includes Dense feature aggregator, LSTM feature aggregator, and GRU feature aggregator.
[0097] (1) Dense feature aggregator
[0098] The Dense feature aggregator is a simple fully connected feature aggregator. Users can simply splice and aggregate the features processed by different encoders through the Dense feature aggregator. It mainly uses a multi-layer fully connected neural network for output. Among them, the number of hidden layers and the number of neurons in the output layer of the Dense feature aggregator are both designed based on the network representation ability. The network structure of the Dense feature aggregator is as Figure 7 shown
[0099] Users can customize the Dense feature aggregator, including the size of the fully connected neural network.
[0100] (2) LSTM feature aggregator
[0101] The LSTM feature aggregator is a feature aggregator that uses the recurrent neural network LSTM. For partially observable Markov decision process (POMDP) problems, the recurrent neural network LSTM can help memorize past information and predict the future. Its principle is to first concatenate multiple features and process them through a multi-layer fully connected neural network, then process them through multiple modules (LSTM Cells) used to memorize long-term dependencies, and finally use a fully connected network for output. Among them, the number of hidden layers and the number of neurons in the output layer of the LSTM feature aggregator are both designed based on the network representation ability. The network structure of the LSTM feature aggregator is as shown in Figure 8 shown.
[0102] Users can customize the LSTM feature aggregator, including the number of fully connected neural networks, LSTM Cells, and neurons in the output layer.
[0103] (3) GRU feature aggregator
[0104] The GRU feature aggregator is a feature aggregator that uses the recurrent neural network GRU. Similar to the LSTM feature aggregator, for partially observable Markov decision process (POMDP) problems, the recurrent neural network GRU can help memorize past information and predict the future. Its principle is to first concatenate multiple features and process them through a multi-layer fully connected neural network, then process them through multiple modules (GRU Cells) used to memorize long-term dependencies, and finally use a fully connected network for output. Among them, the number of hidden layers and the number of neurons in the output layer of the GRU feature aggregator are both designed based on the network representation ability. The network structure of the GRU feature aggregator is as shown in Figure 9 shown.
[0105] Users can customize the GRU feature aggregator, including the number of fully connected neural networks, GRU Cells, and neurons in the output layer.
[0106] III. Feature decoder
[0107] A decoder is provided to decode features and output decision actions such as discrete actions, ordered unit selections, unordered unit selections, and single unit selections. The feature decoder is the exit of the entire neural network, used to decode the features aggregated by the aggregator and output actions, that is, to make decisions. It includes a discrete action decoder, an ordered unit selection decoder, an unordered unit selection decoder, and a single unit selection decoder.
[0108] (1) Discrete action decoder
[0109] The discrete action decoder is a classification action decoder used for making decisions on discrete actions, such as walking forward and walking backward. Its principle is to first input the aggregated features into a multi-layer fully connected layer, then into a fully connected layer with the number of output neurons being the action dimension x to obtain the policy, and finally use the policy sampling to obtain the behavioral actions. Among them, the number of hidden layers of the discrete action decoder is designed based on the network representation ability, and the number of neurons in the output layer is designed based on the action space. The network structure of the discrete action decoder is as shown in Figure 10 shown.
[0110] The user can customize the discrete action decoder, including the fully connected neural network and the number of discrete actions.
[0111] (2) Ordered unit selection decoder
[0112] The ordered unit selection decoder is used for making an ordered selection of multiple units, that is, making a decision on ordered multi-body selection. Among them, the number of hidden layers of the ordered unit selection decoder is designed based on the network representation ability, and the number of neurons in the output layer is designed based on the action space. The network structure of the ordered unit selection decoder is as shown in Figure 11 shown.
[0113] The user can customize the ordered unit selection decoder, including the maximum number of unit selections.
[0114] (3) Unordered unit selection decoder
[0115] The unordered unit selection decoder is used for making an unordered selection of multiple units, that is, making a decision on unordered multi-body selection. Among them, the number of hidden layers of the unordered unit selection decoder is designed based on the network representation ability, and the number of neurons in the output layer is designed based on the action space. The unordered unit selection decoder is as shown in Figure 12 shown.
[0116] The user can customize the unordered unit selection decoder, including the size of the attention element.
[0117] (4) Single unit selection decoder
[0118] The single unit selection decoder is used for making a selection of a single unit, that is, making a decision on single-body selection. Among them, the number of hidden layers of the single unit selection decoder is designed based on the network representation ability, and the number of neurons in the output layer is designed based on the action space. The network structure of the single unit selection decoder is as shown in Figure 13 shown.
[0119] The user can customize the single unit selection decoder, including the size of the attention element.
[0120] The agent construction framework provides a standardized game confrontation agent development architecture, supporting algorithm developers to develop single-unit agents for 3-on-3 (3V3) scenarios and formation agents for urban combat, airport seizure, and mountain combat. It includes development interfaces such as agent network structure, observed situation input, and decision-making action output, supporting the construction of single-agent and multi-agent models based on deep neural networks, real-time sensing of simulation situations, and issuing instructions to control the behavior of simulation entities. The specific description is as follows:
[0121] (1) Agent model construction
[0122] Based on the agent construction framework, the agent model construction module designs and constructs an agent model for realizing the confrontation simulation and deduction of unmanned equipment, mainly including three parts: single-unit agent model, formation agent model, and rule agent model. The specific description is as follows.
[0123] Refer to Figure 1 , which shows the step flowchart of an embodiment of the control method of an agent in this application. The method may specifically include the following steps:
[0124] S101. Obtain battlefield situation information, which at least includes our side's situation information and the opponent's situation information;
[0125] Specifically, during the construction of the single-unit agent model and the formation agent model, the terminal device obtains battlefield situation information in different scenarios, which at least includes our side's situation information and the opponent's situation information. Optionally, our side's situation information at least includes general features such as speed, heading, and blood volume, vector features such as its own position, weapon, relative position of our side, and relative position of the opponent, unique features such as our side's affiliation and the opponent's affiliation, and terrain grid map as a spatial feature of a preset size.
[0126] S102. Extract features from the battlefield situation information to obtain a feature vector corresponding to the battlefield situation information;
[0127] S103. Perform fusion processing on the feature vector to obtain a fused feature vector;
[0128] S104. Train the pre-trained agent model according to the fused feature vector to obtain battlefield decision information; the pre-trained agent model is based on a neural network;
[0129] S105. Generate a control instruction corresponding to the battlefield decision information according to the battlefield decision information.
[0130] Another embodiment of this application further supplements the control method of the agent provided in the above embodiment.
[0131] Optionally, the pre-trained agent model at least includes a single-agent model and a formation agent model, and the formation agent model is used to output a command strategy for controlling the formation according to the battlefield situation information.
[0132] Specifically, (2) single-agent model
[0133] The single-agent model will be constructed based on deep reinforcement learning technology. After construction, the single-agent deep reinforcement learning agent can form a standard data call interface with the simulation model of its corresponding equipment, and can obtain the sensor data, weapon system performance data, and damage data of the simulation model in real time. It can reasonably use the terrain for concealment and shooting, and complete autonomous maneuvering and fire strikes based on the payload of the unmanned platform itself (including various sensors and weapon systems).
[0134] The single-agent model includes 7 individual single-intelligent behavior models such as small unmanned vehicles, medium unmanned vehicles, heavy unmanned vehicles, fixed-wing drones, quadrotor drones, wheeled vehicles, and tanks, and supports autonomous decision-making in a 3-on-3 (3V3) scenario.
[0135] The single-agent based on deep reinforcement learning technology mainly includes a situation input module, a decision result generation module, and a control instruction output module. The specific process is as follows:
[0136] 1) Situation input:
[0137] The battlefield information changes rapidly. Modern warfare is an integrated joint operation under information conditions and needs to be carried out under an integrated combat situation. The key to forming such an integrated combat situation is to form integrated combat situation data in real time, and various different data fusion problems encountered in forming the combat situation need to be solved.
[0138] The situation input module is mainly used to form a combat situation during real-time combat, integrate battlefield situation information, and provide the overall battlefield situation for the decision-making intelligent algorithm, including the global situation, as well as the local situation of our own units and the local situation of enemy units. After being extracted, fused, and sorted by the situation fusion module, it forms consistent and shared overall battlefield situation information, which is used as the main input of the neural network. Accurate battlefield situation is a prerequisite for ensuring subsequent decision-making operations.
[0139] Taking the 3v3 unmanned vehicle scenario as an example, the main function of the situation input module is to extract data and convert it into features that the agent can learn. Among them, speed, heading, and blood volume in our own features are general features, own position, weapon, relative position of our own side, and relative position of the enemy side are vector features, the side to which we belong and the side to which the enemy belongs are one-hot features, and the terrain grid map is a spatial feature of size Matrix[256, 256]. The specific structure is as Figure 14 shown:
[0140] Through this feature, the Agent can not only recognize its own capabilities, but also recognize its own position and relative position to the enemy, which is convenient for the Agent to effectively plan paths and strike at the right time.
[0141] Optionally, train the pre-trained agent model according to the fused feature vector to obtain battlefield decision-making information, including:
[0142] Input the fused feature vector into the intelligent decision-making algorithm model to evaluate and make decisions on the overall battlefield situation;
[0143] The intelligent decision-making algorithm model includes a situation feature extraction unit, a situation evaluation unit, and an inference and decision-making unit. The situation feature extraction unit, the situation evaluation unit, and the inference and decision-making unit are deep neural network models; the situation evaluation unit is used to evaluate the overall situation according to the overall situation of the current situation through a value model; the situation feature extraction unit is used to transform the fused feature vector to obtain an input vector that meets the requirements of the neural network model;
[0144] The inference and decision-making unit is used to train the neural network model according to the input vector to generate battlefield decision-making information, and the battlefield decision-making information includes at least movement tasks and strike tasks.
[0145] After the decision result generation module completes the fusion of the overall combat situation, it is input into the intelligent decision-making algorithm model to further evaluate and make decisions on the overall battlefield situation. The situation feature extraction unit, the situation evaluation unit, and the inference and decision-making unit are all composed of a series of specific deep neural network models. In the situation evaluation unit, it is mainly to understand the overall situation of the current situation and evaluate the overall situation through a value model. In the intelligent decision-making algorithm model, it is divided into different unit structures, including a situation extraction unit and an inference and decision-making unit. Among them, the situation extraction unit receives the input of the situation fusion module, further processes this information to form an expression that can be understood and recognized by the machine, and provides it for the inference and decision-making unit to calculate and reason. The inference and decision-making unit mainly relies on reinforcement learning technology, based on the training data encountered during the learning process, observes the command action patterns that can form effective decisions or high-return decisions, so as to achieve the effect of effective command when making real decisions.
[0146] The decision instructions mainly include motion control and payload control, etc., enabling the agent to reasonably use the terrain for concealment and shooting, and complete autonomous maneuvering and fire strikes based on the payload of the unmanned platform itself (including various sensors and weapon systems). For example Figure 15As shown: The decision result is the output of the neural network, which is a vector Vector(1,4) with 1 row and 4 columns. Each head corresponds to an action. During the parsing process, the action head will be parsed first. If the move action is selected, the target unit will not execute the strike on the enemy target. Similarly, if the strike action is selected, the target unit will not execute the move action. According to the action selected by the neural network, the Agent will execute different instructions. For the movement task, the neural network is mainly used to select the polar angle, and the result is the discrete 8 parts of the angle in the range of [0, 360]. For the strike task, the neural network is mainly used to select the enemy unit ID.
[0147] Optionally, according to the battlefield decision information, control instructions corresponding to the battlefield decision information are generated, including:
[0148] Decode the battlefield decision information into a target selection instruction, a sensor selection instruction, and an action selection instruction corresponding to the battlefield decision information;
[0149] At each decision-making time, according to the current battlefield decision information and the next battlefield situation information, determine the next battlefield decision information.
[0150] Specifically, the control instruction output module includes a target selection unit, a sensor selection unit, and an action selection unit, which can decode into specific target selection instructions, sensor selection instructions, and action selection instructions according to the inference decision information, and output through the simulation environment docking interface to control the simulation model. At each decision-making time, the intelligent body will output the command and control for each unit, then collect the states of each unit and the overall situation, and then call the next decision-making command. As Figure 16 shown:
[0151] Optionally, the method further includes:
[0152] During the neural network training process, in combination with the simulation environment, set the reward function;
[0153] Specifically include:
[0154] Construct the reward function in the way of mainly using the global reward and supplemented by the local reward. The local reward includes at least the exploration area reward, the relative distance penalty, and the firing reward; the global reward includes at least the penalty for our own battle damage, the reward for the enemy's battle damage, and the time penalty.
[0155] To achieve an accurate assessment of the situation, the neural network parameters of the inference decision unit in the decision result generation module need to be combined with the simulation environment for the design of the reward function.
[0156] The design of the reward function for common deep reinforcement learning agent training tasks is mainly divided into two ideas. One is the global nature of the reward, that is, a global reward is given for each complete game; the other is to give several reward values in real time according to the situation information during a game. Compared with the former, the second idea will provide more guidance for the learning of the agent; the reward value given by the first idea is more accurate than the second, and usually can help the agent explore a better solution.
[0157] The embodiment of this application adopts the method of mainly using global rewards and supplemented by local rewards to shape the reward function. In this way, on the one hand, the agent can be guided to learn in the direction of the ultimate goal through global rewards, and on the other hand, the problem of sparse rewards can be solved through local rewards, and the learning speed can be accelerated. In addition, the exploration ability of the agent can be increased by flexibly adjusting the local reward function (such as adding perturbations and adding exploration reward values), so as to avoid falling into local optimal solutions.
[0158] The specific reward function design is as follows:
[0159] Local rewards: Exploration area reward + Relative distance penalty + Firing reward
[0160] Global rewards: Our side's battle damage penalty + Enemy's battle damage reward + Time penalty
[0161] For local rewards, the purpose of this reward function design is to encourage the Agent to explore areas while not getting too close to each other. Because in the case of limited time, although squad exploration can effectively ensure survival rate, it will greatly reduce the exploration efficiency. The purpose of designing the firing reward is that in the choice between survival and eliminating the enemy, eliminating the enemy should obtain a greater reward, so as to prevent the Agent from only choosing to explore and not choosing to strike.
[0162] For global rewards, the purpose of this reward function design is to try to exchange the smallest loss for the greatest killing, and at the same time design a time penalty to prevent the agent from being obsessed with exploration and ignoring attacks.
[0163] The neural network structure design of the single-agent is as Figure 17 shown.
[0164] The formation agent model will be constructed based on deep reinforcement learning technology. The constructed formation deep reinforcement learning agent is a formation-level command strategy model for the Fengjia combat unit, which can output the command strategy for controlling the formation according to the situation analysis and judgment collected from the battlefield. It has the ability to complete mission-level formation operations and can achieve mission-level strategy coordination of multiple unmanned platforms based on the global situation.
[0165] The formation intelligent agent model can give full play to the ability characteristics of the Fengjia combat unit, realize the command and operation of unmanned equipment in three typical battlefield environments: urban combat, airport seizure and mountain combat, support the input of three types of situation data: space, time, and statistics, and can simultaneously identify more than 30 combat units such as small unmanned vehicles, medium unmanned vehicles, heavy unmanned vehicles, unmanned aerial vehicles, wheeled vehicles, and tanks. The scale of neural network parameters is not less than 1 million, and it supports the scheduling and control of more than 15 simulation entities.
[0166] 1) Situation input:
[0167] The battlefield is selected as mountain combat. The red side is the attacking side. After capturing the enemy's high ground, it can prepare for the subsequent attack on the city. The blue side is the defending side. The situation extraction module of the formation intelligent agent is as Figure 18 shown:
[0168] The decision-making cycle of the formation-level intelligent agent is different from that of the single-equipment intelligent agent. Its main purpose is to select the appropriate task at the appropriate time after comprehensively considering the battlefield environment and force statistics. Therefore, the situation input is all the combat units of one's own side and the detected enemy units. The overall force situation is used as a feature to input to the Agent to train its situation awareness ability.
[0169] 2) Decision result generation:
[0170] Considering the scenario scale and training difficulty, the result output by the formation-level intelligent agent is at the task level. As the combat time progresses, the intelligent agent will only output which tasks should be executed at the current moment, which can be one task or multiple tasks. As for the execution of the tasks, it is handed over to the rule intelligent agent to implement. The specific design is as Figure 19 shown.
[0171] 3) Reward function design
[0172] This application embodiment adopts the global reward as the main one. The specific reward function evaluation is as Figure 20 shown: In each decision-making cycle, the intelligent agent selects the combat task according to the current situation and hands it over to the reward function evaluation matrix for evaluation. If all the preconditions, own conditions, and enemy conditions are met, a reward can be obtained; otherwise, a penalty will be obtained.
[0173] 4) Neural network architecture
[0174] In order to better control all the participating troops and avoid the problems of communication and coordination between different commanders, a global commander neural network architecture is designed, that is, a global commander controls all the participating units. In this architecture, the action of the intelligent agent is mainly a one-row six-column matrix Matrix(1, 6).
[0175] According to the overall scenario and assumption of the application embodiment, a single-agent neural network structure is designed. The scale of the neural network parameters is not less than 1 million. The specific structure is as follows Figure 21 As shown, in terms of state input, this solution uses a fully connected network to extract general features, and uses a feature transformation network to extract the unit features of the red and blue parties respectively. After merging the three feature information, the long short-term memory network is used to process historical information. In terms of action output, a fully connected network is used to generate predicates, an attention mechanism is used to select objects, and a fully connected network is used to select relevant parameters.
[0176] In addition, this application also provides an agent call interface module, which constructs standard communication interfaces between agents and between agents and the simulation environment, realizes the combat mission communication between the single-agent model and the formation-agent model, and supports the command and control of the overall cooperation between the formation and the single entity; at the same time, it supports the formation of a standard data call interface between the agent and the simulation model of the corresponding equipment in the simulation environment, docks and exchanges to generate game confrontation data for the agent to train and learn. The agent call interface can convert the observation information obtained from the environment into combat situation information that the agent can recognize, and at the same time convert the action information sent from the agent deployment module into decision instructions that the simulation deduction platform can receive.
[0177] The agent call interface provides a complete set of development interfaces. Users can complete data interaction by inputting data and parameters according to the requirements of the development document description and the format of the interface function. The agent call interface mainly includes a situation information processor interface, an agent training interface, an agent prediction interface, an action command conversion interface, and a reward value calculation interface. The specific descriptions are shown in Table 1:
[0178] Table 1
[0179]
[0180] The agent training module uses agent training technology with deep reinforcement learning algorithm as the core to conduct large-scale distributed training on the constructed deep neural network to achieve the expected intelligent effect, and shows the process of the ability evaluation index improving with the learning iteration in the form of charts. Our unit will follow the confidentiality requirements of the user unit and conduct on-site development for the training of some agents.
[0181] The agent training module mainly includes two parts: large-scale distributed training technology and visualization of ability evaluation indicators. The specific descriptions are as follows.
[0182] The large-scale distributed training technology supports agent training, effectively utilizes resources without sacrificing data efficiency and resource utilization rate, and realizes the rapid iteration of agents to achieve the expected intelligent effect. The architecture diagram of the large-scale distributed training technology is as follows Figure 22 As shown:
[0183] The large-scale distributed training technology includes a data generation engine module, a continuous learning engine module, a prediction and inference engine module, and an intelligent engine master control module, where:
[0184] Data generation engine module: mainly used for interacting with the high-concurrency simulation environment to generate high-quality sample data. Sample the sample data through multiple simulation environments and supply it to the continuous learning engine module for use;
[0185] Continuous learning engine module: consume massive data and use relevant algorithms to optimize the agent;
[0186] Prediction and inference engine module: quickly respond to and drive the simulation environment to generate data;
[0187] Intelligent engine master control module: mainly responsible for saving the state of the reinforcement learning task and the workflow, and is also responsible for the life cycle and resource management of the intelligent engine.
[0188] The traditional distributed reinforcement learning engine adopts a data generation-training two-layer architecture. Among them, the data generation engine module is responsible for continuously generating training samples, and the continuous learning engine module is responsible for continuously updating the decision model parameters according to the training samples. The sampling engine of the traditional reinforcement learning engine runs completely on the CPU, and the learning engine and the prediction engine run on the GPU. A large amount of data communication will occur between the CPU and the GPU. Due to the bottleneck of the network bandwidth, a large amount of data transmission and neural network parameter synchronization will reduce the working efficiency of the engine. In addition, the larger the parameters of the neural network model, the significantly lower the speed of forward inference using the CPU. Finally, the sampling engine needs to schedule CPU resources to perform simulation deduction and neural network inference simultaneously, and cannot fully improve the utilization rate of CPU multi-threaded resources.
[0189] Therefore, the large-scale distributed training technology provided by the embodiments of the present application adds an additional prediction and inference engine module, adopts a data generation-training-prediction three-layer architecture, uses GPU resources to accelerate the neural network inference speed, effectively solves several problems of the traditional distributed reinforcement learning engine, converts the large-scale computing power into large-scale data processing power, and realizes the efficient reinforcement learning training of the decision model under distributed computing power.
[0190] The data generation engine module running in parallel interacts with multiple simulation environments and performs high-speed network transmission through the virtual network management component to obtain simulation situation information. After completing the format conversion of the simulation environment data to the input data of the agent neural network (including deletion of useless data, data normalization, OneHot encoding, etc.), the neural network features are input to the prediction and inference engine module, which performs calculations and outputs the neural network decision output. After completing the mapping from the neural network decision output to the simulation environment instructions, the data generation engine module sends specific instructions to the simulation environment. The training data formed in this process, that is, the <State, Action, Reward> triple data stream, is input to the training data buffer pool for use by the subsequent continuous learning engine module.
[0191] After the training data is formed, the continuous learning engine module asynchronously reads the training data buffer pool. Here, "asynchronously" means that the continuously learning engine module deployed in a container reads data in real time, rather than waiting for the slowest training data to arrive and then reading it uniformly. The continuous learning engine module calls the reinforcement learning algorithm, calculates the gradient through the loss function (Loss), determines the update direction of the agent neural network, and performs gradient backpropagation to update the neural network parameters. During the update process of the neural network model parameters, the Ring-AllReduce method is adopted based on the virtual network management component. After the update is completed, the neural network model parameters are synchronized to the prediction and inference engine module.
[0192] As Figure 23 shown, the running relationships of each component are as follows:
[0193] The data generation engine interacts with multiple simulation environments in parallel. In each interaction, the simulation environment outputs the situation to the data generation engine, and the data generation engine returns instructions to the environment;
[0194] The generated data is stored in the sample pool in a streaming manner and will be read in batches when the learning engine has a need;
[0195] Each instance of the continuous learning engine independently obtains samples from the sample pool, performs forward calculations on the neural network, calculates the gradient according to the loss function defined by the reinforcement learning algorithm, determines the update direction of the network parameters, and then performs backward calculations to update the parameters;
[0196] The new version model generated by the learning engine will be loaded by the prediction and inference engine. When the data generation engine has an interaction requirement, it will send the observation of the environment to the prediction and inference engine, and return the action to the data generation engine after forward calculation;
[0197] The intelligent engine master control module is the process manager of the entire system. It will send different instruction flows to the above three modules according to the optimization processes of different reinforcement learning algorithms to complete the continuous operation of the reinforcement learning process.
[0198] Based on the above three-layer architecture, the sampling module runs a batch of simulation deductions on the CPU and conducts real-time interaction of situation data and decision-making data with the neural network inference running on the GPU. Compared with the model data and complete batch data under the traditional engine architecture, the data communication volume is greatly reduced, and the network bandwidth is effectively utilized. On the other hand, the continuous learning engine module and the prediction module run on the GPU at the same time, making full use of the high concurrency ability of the GPU to improve the speed of neural network batch inference, reduce the latency of model parameter update, and accelerate the efficiency of model parameter synchronization.
[0199] Visualization of the ability evaluation index shows the process of improvement of the intelligent agent ability evaluation indexes such as task success rate and winning rate during the learning iteration in the form of charts such as line charts and bar charts, as Figure 24 and Figure 25 shown.
[0200] The intelligent agent deployment module deploys the constructed single-unit and formation intelligent agent models on the Windows and mainstream Linux platforms, conducts human-machine confrontation demonstrations in the confrontation deduction environment, and outputs information such as battle damage, consumption, and victory and defeat determination relationships in real time. Examples are shown as Figure 26 and 27 shown.
[0201] The intelligent agent control method provided by the embodiment of the present application obtains battlefield situation information, where the battlefield situation information at least includes our side's situation information and the opponent's situation information; extracts features from the battlefield situation information to obtain a feature vector corresponding to the battlefield situation information; performs fusion processing on the feature vector to obtain a fused feature vector; trains a pre-trained intelligent agent model according to the fused feature vector to obtain battlefield decision information; the pre-trained intelligent agent model is obtained based on a neural network; according to the battlefield decision information, a control instruction corresponding to the battlefield decision information is generated. The embodiment of the present application constructs an intelligent command and control entity that can support single-unit and formation two-level combat command training through artificial intelligence algorithms. By combining the rule-based decision-making method with the deep reinforcement learning technology, an artificial intelligence training platform that can support intelligent training and simulation deduction and integrates key technologies such as deep learning and reinforcement learning is built. Further, a virtual intelligent agent for command decision-making is generated, providing intelligent opponents at different levels for combat command simulation training, and can provide flexible battlefield situation displays according to different scenarios, and intuitively display the evaluation results of the analysis and evaluation module.
[0202] Another embodiment of the present application provides a control device for an intelligent agent, which is used to execute the control method for the intelligent agent provided in the above embodiment.
[0203] Referring to Figure 28 , a structural block diagram of an embodiment of a control device for an intelligent agent according to the present application is shown. The device may specifically include the following modules: an acquisition module 2801, an extraction module 2802, a fusion module 2803, a training module 2804, and a generation module 2805, where:
[0204] The acquisition module 2801 is used to acquire battlefield situation information, and the battlefield situation information at least includes our side's situation information and the opponent's situation information;
[0205] The extraction module 2802 is used to extract features from the battlefield situation information to obtain a feature vector corresponding to the battlefield situation information;
[0206] The fusion module 2803 is used to perform fusion processing on the feature vectors to obtain a fused feature vector;
[0207] The training module 2804 is used to train a pre-trained intelligent agent model according to the fused feature vector to obtain battlefield decision-making information; the pre-trained intelligent agent model is obtained based on a neural network;
[0208] The generation module 2805 is used to generate a control instruction corresponding to the battlefield decision-making information according to the battlefield decision-making information.
[0209] Another embodiment of the present application further supplements the control device for the intelligent agent provided in the above embodiment.
[0210] Optionally, the pre-trained intelligent agent model at least includes a single-unit intelligent agent model and a formation intelligent agent model, and the formation intelligent agent model is used to output a command strategy for controlling the formation according to the battlefield situation information.
[0211] Optionally, the training module is used for:
[0212] Input the fused feature vector into an intelligent decision-making algorithm model to evaluate and make decisions on the overall battlefield situation;
[0213] The intelligent decision-making algorithm model includes a situation feature extraction unit, a situation evaluation unit, and an inference and decision-making unit. The situation feature extraction unit, the situation evaluation unit, and the inference and decision-making unit are deep neural network models; the situation evaluation unit is used to evaluate the overall situation through a value model according to the overall situation of the current situation; the situation feature extraction unit is used to transform the fused feature vector to obtain an input vector that meets the requirements of the neural network model;
[0214] The inference and decision-making unit is used to train the neural network model based on the input vector to generate battlefield decision-making information, and the battlefield decision-making information includes at least movement tasks and strike tasks.
[0215] Optionally, a generation module, configured to:
[0216] Decode the battlefield decision-making information into corresponding target selection instructions, sensor selection instructions, and action selection instructions;
[0217] At each decision-making moment, determine the next battlefield decision-making information based on the current battlefield decision-making information and the next battlefield situation information.
[0218] Optionally, a training module, configured to:
[0219] During the neural network training process, combine the simulation environment to set the reward function;
[0220] Specifically, it includes:
[0221] Construct the reward function in a way that takes the global reward as the main part and the local reward as the auxiliary part. The local reward includes at least exploration area reward, relative distance penalty, and firing reward; the global reward includes at least our side's battle damage penalty, the opponent's battle damage reward, and time penalty.
[0222] Optionally, our side's situation information includes at least speed, heading, and blood volume as general features, own position, weapons, relative position of our side, and relative position of the opponent as vector features, the side to which our side belongs and the side to which the opponent belongs as unique features, and the terrain grid map as a spatial feature with a preset size.
[0223] The control device of the intelligent agent provided by the embodiment of the present application, through
[0224] Obtain battlefield situation information, where the battlefield situation information includes at least our side's situation information and the opponent's situation information; extract features from the battlefield situation information to obtain a feature vector corresponding to the battlefield situation information; perform fusion processing on the feature vector to obtain a fused feature vector; train a pre-trained agent model based on the fused feature vector to obtain battlefield decision-making information; the pre-trained agent model is obtained based on a neural network; generate a control instruction corresponding to the battlefield decision-making information according to the battlefield decision-making information. In the embodiments of the present application, an intelligent command and control entity that can support the training of single-unit and formation-level combat command is constructed through an artificial intelligence algorithm. By combining a rule-based decision-making method with deep reinforcement learning technology, an artificial intelligence training platform that can support intelligent training, simulation deduction, and integrates key technologies such as deep learning and reinforcement learning is built. Further, a virtual intelligent agent for command decision-making is generated, providing intelligent opponents at different levels for combat command simulation training. Flexible battlefield situation displays can be provided according to different scenarios, and the evaluation results of the analysis and evaluation module are intuitively displayed.
[0225] Another embodiment of the present application provides a terminal device for executing the control method of the intelligent agent provided in the above embodiment.
[0226] Figure 29 is a schematic structural diagram of a terminal device of the present application, as Figure 29 shown, the terminal device includes: at least one processor 701 and a memory 702;
[0227] The memory stores a computer program; at least one processor executes the computer program stored in the memory to implement the control method of the intelligent agent provided in the above embodiment.
[0228] Another embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed, it implements the control method of the intelligent agent provided in any of the above embodiments.
[0229] It should be noted that the above detailed description is exemplary and is intended to provide further illustration of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.
[0230] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0231] It should be noted that the terms "first", "second", etc. in the description, claims and the above-mentioned drawings of this application are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here.
[0232] In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that comprises a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0233] For ease of description, spatial relative terms such as "above", "over", "on the upper surface", "above-mentioned", etc. may be used here to describe the spatial positional relationship of one device or feature to other devices or features as shown in the figures. It should be understood that spatial relative terms are intended to encompass different orientations in use or operation in addition to the orientation depicted in the figures. For example, if the device in the figures is inverted, the device described as "above" or "over" other devices or structures will then be positioned "below" or "under" other devices or structures. Thus, the exemplary term "above" can include both the orientations of "above" and "below". The device can also be positioned in other different ways, such as rotated 90 degrees or in other orientations, and corresponding interpretations of the spatial relative descriptions used here will be made.
[0234] In the detailed description above, reference has been made to the accompanying drawings, which form a part hereof. In the drawings, like reference numerals typically identify like components, unless the context indicates otherwise. The illustrated embodiments described in the detailed description, the drawings and the claims are not meant to be limiting. Other embodiments may be used and other changes may be made without departing from the spirit or scope of the subject matter presented here.
[0235] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. A control method for an intelligent agent, characterized in that: The method comprises: Acquiring battlefield situation information, wherein the battlefield situation information includes at least our situation information and the opponent's situation information; Extracting features from the battlefield situation information to obtain a feature vector corresponding to the battlefield situation information; Performing fusion processing on the feature vectors to obtain fused feature vectors; Training a pre-trained intelligent agent model according to the fused feature vector to obtain battlefield decision information; the pre-trained intelligent agent model is obtained based on a neural network; Based on the battlefield decision information, a control instruction corresponding to the battlefield decision information is generated.
2. The control method of an intelligent agent according to claim 1, characterized in that: The pre-trained intelligent agent model includes at least a single-agent agent model and a formation intelligent agent model, and the formation intelligent agent model is used to output a command strategy for controlling the formation according to the battlefield situation information.
3. The control method of an intelligent agent according to claim 2, characterized in that: The pre-trained agent model is trained according to the fused feature vector to obtain battlefield decision information, including: The fused feature vector is input into the intelligent decision-making algorithm model to evaluate and make decisions on the overall battlefield situation; The intelligent decision-making algorithm model includes a situation feature extraction unit, a situation assessment unit and a reasoning decision unit, wherein the situation feature extraction unit, the situation assessment unit and the reasoning decision unit are deep neural network models; the situation assessment unit is used to assess the overall situation through a value model according to the overall status of the current situation; the situation feature extraction unit is used to convert the fused feature vector to obtain an input vector that meets the requirements of the neural network model; The inference decision unit is used to train the neural network model according to the input vector to generate the battlefield decision information, and the battlefield decision information includes at least a mobile task and a strike task.
4. The control method of the intelligent agent according to claim 3, characterized in that: The step of generating a control instruction corresponding to the battlefield decision information according to the battlefield decision information comprises: Decoding the battlefield decision information into target selection instructions, sensor selection instructions, and action selection instructions corresponding to the battlefield decision information; At each decision-making, the next battlefield decision information is determined based on the current battlefield decision information and the next battlefield situation information.
5. The control method of an intelligent agent according to claim 2, characterized in that: The method further comprises: During the neural network training process, the reward function is set in combination with the simulation environment; Specifically include: The reward function is constructed in a manner of taking global reward as the main and local reward as the auxiliary, wherein the local reward at least includes exploration area reward, relative distance penalty and firing reward; the global reward at least includes our side's damage penalty, the opponent's damage reward and time penalty.
6. The control method of an intelligent agent according to claim 1, characterized in that: The situation information of our side at least includes speed, heading, and health as common features, own position, weapons, relative position of our side, and relative position of the opponent as vector features, our side and the opponent's side as unique features, and terrain grid map as a spatial feature of preset size.
7. A control device for an intelligent agent, characterized in that: The device comprises: An acquisition module, used to acquire battlefield situation information, wherein the battlefield situation information at least includes our situation information and the opponent's situation information; An extraction module, used for performing feature extraction on the battlefield situation information to obtain a feature vector corresponding to the battlefield situation information; A fusion module, used for fusing the feature vectors to obtain a fused feature vector; A training module, used for training a pre-trained intelligent agent model according to the fused feature vector to obtain battlefield decision information; the pre-trained intelligent agent model is obtained based on a neural network; A generation module is used to generate control instructions corresponding to the battlefield decision information based on the battlefield decision information.
8. The control device of the intelligent body according to claim 7, characterized in that: The pre-trained intelligent agent model includes at least a single-agent agent model and a formation intelligent agent model, and the formation intelligent agent model is used to output a command strategy for controlling the formation according to the battlefield situation information.
9. A terminal device, characterized in that: include: at least one processor and memory; The memory stores a computer program; The at least one processor executes the computer program stored in the memory to implement the control method of the intelligent agent according to any one of claims 1-6.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed, implements the control method of the intelligent agent according to any one of claims 1 to 6.
Citation Information
Cited By
Training method and device of aircraft control model
CN121254661A
A method and apparatus for training an aircraft control model
CN121254661B