Unmanned aerial vehicle air confrontation intelligent decision algorithm management, training and verification integrated system

By designing an integrated system for intelligent decision-making algorithm management, training and verification in the air confrontation of drones, and integrating algorithm library management, intelligent model design, training and simulation verification, the problems of intelligent decision-making algorithm management and training in the air confrontation scenario of drones are solved, and the algorithm development efficiency and application effect are improved.

CN120217828APending Publication Date: 2025-06-27AIR FORCE UNIV PLA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510199127.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the aerial confrontation scenario of drone, it is difficult for the existing technology to effectively manage, train and verify intelligent decision-making algorithms, resulting in an increase in decision-making complexity and difficulty, affecting the development efficiency and application effect of the algorithm.

Method used

Design an integrated system for the management, training and verification of intelligent decision-making algorithms in the air against drone, integrate intelligent decision-making algorithm library management, drone decision-making intelligent body model design, intelligent decision-making model training and simulation verification evaluation to realize the full process development and management of the algorithm.

Benefits of technology

Through this system, the development efficiency and application effect of intelligent decision-making algorithms of drones are improved, the development and use thresholds and costs are reduced, and the rapid formation and application of research results are promoted.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120217828A_ABST
    Figure CN120217828A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned aerial vehicle air confrontation intelligent decision-making algorithm management, training and verification integrated system, relates to the technical field of unmanned aerial vehicle intelligent decision-making, and particularly aims at solving the problem that in the prior art, in an existing unmanned aerial vehicle intelligent decision-making algorithm, an existing decision-making algorithm cannot be used. The management, training and verification integrated system for the unmanned aerial vehicle air confrontation intelligent decision-making algorithm aims at solving the key problems that reasonable management, control and verification of the training process of the unmanned aerial vehicle intelligent decision-making algorithm and guarantee of effectiveness and reliability of the algorithm in practical application are all difficult. Design development, learning training, verification evaluation and sharing management functions of the intelligent decision-making algorithm can be integrated on the same platform, so that the development efficiency and application effect of the unmanned aerial vehicle intelligent decision-making algorithm are effectively improved, the development and use threshold and cost are reduced, and rapid formation of unmanned aerial vehicle intelligent decision-making algorithm research results is promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of intelligent decision-making for unmanned aerial vehicles, and particularly relates to an integrated system for the management, training, and verification of an intelligent decision-making algorithm for unmanned aerial vehicle air combat. Background Art

[0002] The world's unmanned aerial vehicle technology has developed rapidly, and the application scope of unmanned equipment such as unmanned aerial vehicles has been continuously expanded, and the application level has been continuously improved. The application environment of unmanned aerial vehicles is also moving from low-intensity conflict and low-threat environments to high-intensity confrontation and medium-high-threat environments, which has had an important impact on all aspects of the military field. In information-based warfare, various types of information are intertwined and difficult to distinguish, greatly increasing the complexity and difficulty of intelligent decision-making. Therefore, intelligent decision-making technology has emerged. It uses artificial intelligence technologies such as the Internet of Things, cloud computing, deep learning, data mining, and mathematical modeling to enable the unmanned aerial vehicle system to perceive the rapid changes in the situation and perform analysis, judgment, and decision-making.

[0003] With the in-depth development of intelligent technology in the aviation field, air combat has become an important application direction for unmanned aerial vehicles, and its core technology lies in the intelligent decision-making method for autonomous combat of unmanned aerial vehicles. Due to the characteristics of strong confrontation and high dynamics in air combat decision-making, it has become a high-difficulty problem in the field of confrontation decision-making. Over the years, scholars have conducted a large number of explorations on air combat decision-making methods. From the early methods based on statistical decision-making and knowledge reasoning, to the optimal decision-making methods based on models, and then to the methods based on artificial intelligence, people's exploration of unmanned aerial vehicle air decision-making is becoming more and more inclined to new intelligent methods. The methods based on artificial intelligence endow unmanned aerial vehicles with the ability to learn and self-expand, enabling intelligent unmanned aerial vehicles to have more flexible and adaptable decision-making capabilities.

[0004] When designing an intelligent decision-making algorithm for unmanned aerial vehicle air combat, it is necessary to comprehensively consider the technical requirements of multiple aspects such as autonomous perception, decision-making, and execution of unmanned aerial vehicles. At the same time, how to effectively manage the intelligent decision-making algorithm of unmanned aerial vehicles (including storage, invocation, update, etc.) of unmanned aerial vehicles, how to reasonably control the training process of the intelligent decision-making algorithm of unmanned aerial vehicles, and how to verify the intelligent decision-making algorithm of unmanned aerial vehicles to ensure the effectiveness and reliability of the algorithm in actual applications are all key and difficult problems in the current development and application process of the intelligent decision-making algorithm of unmanned aerial vehicles.

[0005] Therefore, there is an urgent need to design an integrated system for the management, training, and verification of an intelligent decision-making algorithm for unmanned aerial vehicle air combat to solve the above problems. Summary of the Invention

[0006] Aiming at the defects existing in the above-mentioned prior art, the purpose of the present invention is to provide an integrated system for the management, training and verification of an intelligent decision-making algorithm for unmanned aerial vehicle (UAV) air combat, which can integrate the functions of the design and development, learning and training, verification and evaluation, and sharing and management of the intelligent decision-making algorithm on the same platform, thereby effectively improving the development efficiency and application effect of the UAV intelligent decision-making algorithm, while reducing the development and use threshold and cost, and promoting the rapid formation of research results on the UAV intelligent decision-making algorithm. It aims to improve the quality and efficiency of the development and application of the UAV intelligent decision-making algorithm through the full-process application of the intelligent decision-making algorithm from design, training, management, verification to evaluation.

[0007] In order to achieve the above purpose, the technical solutions adopted by the present invention are as follows: An integrated system for the management, training and verification of an intelligent decision-making algorithm for UAV air combat, characterized by comprising: An intelligent decision-making algorithm library management subsystem, which is used to support users to re-develop the algorithms in the air game confrontation intelligent algorithm library according to needs, and perform distributed training management, deployment and operation scheduling management on the algorithm library; A UAV decision-making intelligent agent model design subsystem, which provides a standardized development architecture for UAV air combat decision-making intelligent agents, supports algorithm developers to develop intelligent agents for different air mission scenarios, supports the construction of single-intelligent-agent and multi-intelligent-agent models based on deep neural networks, real-time perceives the simulation situation, and issues instructions to control the behavior of simulation entities; An intelligent decision-making model training subsystem, which is used to support the virtual self-play training of the intelligent agent model to improve the level of the intelligent agent; An intelligent decision-making model simulation verification and evaluation subsystem, which is used to provide simulation deduction, intelligent agent ability verification and human-machine confrontation ability, and perform evaluation according to the data of simulation deduction, intelligent agent ability verification and human-machine confrontation ability.

[0008] Preferably, the intelligent decision-making algorithm library management subsystem includes: A re-development module, which is used to support users to re-develop the algorithms in the air game confrontation intelligent algorithm library according to needs, and provides a programming call interface, a neural network call component interface, an intelligent agent training and prediction interface, and an algorithm re-development data flow interface; Among them, the programming call interface can open the algorithm editing template in the air game confrontation intelligent algorithm library and complete the call of reinforcement learning algorithms, transfer learning algorithms, incremental learning algorithms, few-shot learning algorithms, meta-learning algorithms, and machine learning algorithms; The neural network calls component interfaces, which are used to quickly build components of the neural network. The components of the neural network include: fully connected network MLP, convolutional neural network CNN, deep residual network ResNet, recurrent neural network RNN, long short-term memory network LSTM, gated recurrent unit GRU, attention mechanism Attention, feature transformation network Transformer, pointer network Pointer Net; The intelligent agent training and prediction interface facilitates users to set training parameters, start the intelligent agent training, and predict the situation, output the original action information, which is converted into instructions and then transmitted to the simulation environment; The algorithm secondary development data flow interface is used to complete data interaction by inputting data and parameters; the algorithm secondary development data flow interface includes: situation information processor interface, intelligent agent training interface, intelligent agent prediction interface, action command conversion interface, and reward value calculation interface; The algorithm training management module is used to manage and deploy the algorithm training tasks of the algorithms in the UAV air combat confrontation intelligent algorithm library; The algorithm deployment and scheduling module is used to schedule and manage the deployment and operation of the UAV air combat confrontation intelligent algorithm library, realize the docking with the third-party system, and deploy and run the algorithm library and related operation support environment on the private cloud and domestic typical public cloud platforms, supporting the unified management and call of the algorithm library by the third-party system.

[0009] Preferably, the algorithms in the air combat confrontation intelligent algorithm library all support containerized deployment and scheduling; The algorithm training management module manages the algorithm training tasks of the algorithms in the UAV air combat confrontation intelligent algorithm library, including: supporting operations such as starting, terminating, viewing, viewing logs, selecting, editing, and deleting during the training process, supporting the selection of the local service, GPUs server, and Kubernetes cluster for the training task, where the Kubernetes cluster supports single-machine and distributed training methods, and also supports the selection of multiple algorithms for training.

[0010] Preferably, the UAV decision-making intelligent agent model design subsystem includes: The intelligent agent network structure design module provides users with a quick programming framework for building intelligent agents for different air task scenarios. Users can customize the neural network structure of the intelligent agent based on this framework. The intelligent agent network structure design module includes a feature encoder, a feature aggregator, and a feature decoder; Observation Situation Input Design Module, by establishing various standard situation data structures, supports R & D personnel to carry out the input environment situation data structure of the intelligent game model according to the unified standard, and realizes the entity state reading of the simulation platform and the overall situation acquisition in the intelligent decision-making algorithm library management subsystem; the Observation Situation Input Design Module includes an agent policy network situation input sub-module, an agent command decision-making training structure sub-module, and an observation situation input interface sub-module; among them, the agent policy network situation input sub-module obtains the state of our unit, the state of the enemy unit, and the global situation of the deduction through the simulation environment docking interface, inputs them into the situation data preprocessing unit, and performs situation fusion, data completion, missing value repair, outlier processing, and normalization operations. The agent command decision-making training structure sub-module performs situation fusion on the global situation, the local situation of our unit, and the local situation of the enemy unit, and uses the processed feature information as the input of the neural network. The observation situation input interface sub-module realizes the situation data interaction with the simulation environment; Action Space Output Module, by establishing various standard decision data structures, supports R & D personnel to carry out action modeling and development according to this interface; The Action Space Output Module includes an agent output action generation sub-module and an action space output interface sub-module; the Action Space Output Module includes an agent output action generation sub-module and an action space output interface sub-module; among them, the agent output action generation sub-module includes a target selection unit, a sensor selection unit, and an action selection unit, which can decode into specific target selection instructions, sensor selection instructions, and action selection instructions according to the inference decision information, generate meta-actions using a fully connected network, select the red side unit and the enemy target respectively using the attention mechanism, generate the x and y coordinates of the target position using a fully connected network, output through the simulation environment interface and control the simulation entity. The action space output interface sub-module realizes the task instruction interaction with the simulation environment.

[0011] Preferably, the feature encoder is used for feature encoding and feature extraction of various categories. The feature encoder includes: a spatial feature encoder, an entity feature encoder, and a general feature encoder. Among them, the spatial feature encoding is used to extract high-dimensional matrix feature information such as images, the entity feature encoder is used to extract entity feature information, and the general feature encoder is used to extract general feature information; The feature aggregator is used to aggregate the features processed by different feature encoders and splice them into complete information; the feature aggregator includes: a Dense feature aggregator, an LSTM feature aggregator, and a GRU feature aggregator; The feature decoder is used to decode the features aggregated by the feature aggregator and output them. The feature decoder includes: a discrete action decoder, an ordered unit selection decoder, an unordered unit selection decoder, and a single unit selection decoder. The discrete action decoder is used for discrete action decision-making. The ordered unit selection decoder is used for ordered selection of multiple units. The unordered unit selection decoder is used for unordered selection of multiple units. The single unit selection decoder is used for single unit selection decision-making; In the observation situation input design module, the multiple standard situation data structures include spatial situation representation, temporal situation representation, and statistical situation representation; In the action space output module, the multiple standard decision data structures include multi-selection data structure, single-selection data structure, multi-classification data structure, binary classification data structure, and continuous value prediction data structure.

[0012] Preferably, the intelligent decision-making model training subsystem includes: The agent training architecture design module is used to support agent training and realize functions such as generation, collection, feature processing, agent sequential decision-making, frequency conversion decision-making, and task success rate prediction of sample data without sacrificing data efficiency and resource utilization rate; The typical agent training method module is used to combine with the agent training architecture design module to train the agent using different training methods. Among them, different training methods include synchronous symmetric self-play, asynchronous symmetric self-play, synchronous asymmetric self-play, and asynchronous asymmetric self-play.

[0013] Preferably, the agent training architecture design module includes: The data generation engine sub-module is used to interact with the high-concurrency simulation environment to generate high-quality sample data; The continuous learning engine sub-module is used to consume massive data and optimize the agent; The prediction inference engine sub-module is used to quickly respond and drive the simulation environment to generate data; The intelligent engine master control sub-module is used for reinforcement learning task state preservation and workflow, and is responsible for the life cycle and resource management of the intelligent engine.

[0014] Preferably, the intelligent decision-making model simulation verification and evaluation subsystem includes: The scenario editing module supports users to design air combat confrontation scenarios, construct near-range, medium-range, and far-range air mission scenarios, as well as surface target penetration strike mission scenarios, and two / three-dimensional synchronous editing of scenario files; The digital earth module provides real-time simulation and deduction functions based on the three-dimensional digital earth display environment, and supports the deployment of local map servers; The simulation model module is used to deploy typical anti-drone systems, manned aircraft, and unmanned aircraft simulation models, construct a simulation model system covering the mission environment, entity units, and action rules, and support typical simulation deductions in the air combat scenario; The simulation deduction module is used to provide simulation deduction and man-machine confrontation functions in typical air combat confrontation scenarios, and at the same time meet the rapid training requirements of intelligent game algorithms; Among them, the simulation deduction module includes: The simulation process driving sub-module is used to drive the model or simulation system to complete the simulation of the scenario plan and generate corresponding results; The simulation operation control sub-module completes the initialization data loading and setting of the mission simulation, the management and control of the simulation operation, and the simulation time advancement according to the equipment information, scenario data, and action plan in the input data; The simulation status monitoring sub-module is used to monitor the states of various entities and entity interactions during the simulation process, collect the states of simulation entities, model data, and entity interaction information, and is used to establish simulation logs and simulation process data; The intelligent decision-making algorithm evaluation module supports users to customize training effect evaluation indicators, automatically generates the training effect curve of the intelligent agent, and gradually converges as the intelligent agent training progresses.

[0015] Preferably, the intelligent decision-making algorithm evaluation module includes: The intelligent agent evaluation feedback calculation sub-module is responsible for constructing an evaluation index system for the intelligent agent around the mission objectives and content according to the evaluation requirements, forming the calculation method of the index and the evaluation criteria for the decision-making ability of the intelligent agent decision body, and supporting the iterative optimization of the intelligent agent; The intelligent agent evaluation feedback calculation sub-module is responsible for receiving the intelligent game simulation confrontation deduction data, and according to the evaluation index requirements, realizing the transformation and processing of the deduction data and the calculation of the deduction index results, and supporting the calculation of the evaluation indexes of the intelligent agent's task ability, task completion situation, and AI decision-making cost-effectiveness ratio.

[0016] The second object of the present invention is to provide a usage method for an integrated system for the management, training, and verification of an intelligent decision-making algorithm for unmanned aircraft air combat, and the method includes: S1. Simulation environment docking; the specific process of the simulation environment docking includes: S101. Perform containerization transformation and packaging on the simulation environment, and deploy the packaging information to the countermeasure knowledge graph platform and the intelligent game algorithm library system; S102. Simulation platform interface docking, providing interface docking between the countermeasure knowledge graph platform and the intelligent game algorithm library system and the typical true three-dimensional air combat simulation software; S103. Design of air mission scenarios. Based on the actual tasks of the air confrontation business scenarios, scenario design is carried out; S2. Agent design. The specific process of the agent design includes: S201. Neural network architecture design. The neural network architecture design process includes: designing the neural network structure and initial parameters of the agent according to the task requirements, control units, and equipment performance of the scenario; S202. Intelligent optimization algorithm design. The intelligent optimization algorithm design process includes: designing and defining the architecture form and optimization objectives of the agent's reinforcement learning algorithm according to the task requirements of the scenario, the characteristics of the control unit strength, and the requirements of the agent's working mode; S3. Agent training. The agent training process includes: S301. Design of evolutionary training method. The evolutionary training method design process includes: designing the agent's evolutionary training method according to the task constraints and the designed intelligent game algorithm; S302. Distributed training of agents. The distributed training process of agents includes: conducting large-scale distributed deep reinforcement learning training on the agents; S303. Migration and deployment of agents. The migration and deployment process of agents is: migrating the trained agents to a typical true three-dimensional air confrontation simulation environment for confrontation deduction; S4. Verification and evaluation of agents. The verification and evaluation process of agents includes: S401. Management of evaluation index system. The management process of the evaluation index system includes: constructing an evaluation index system and model for agents around the task objectives and contents according to the evaluation requirements, forming a calculation method for the indexes, and a judgment standard for the decision-making ability, intelligence level, and agent effectiveness of the agent model; S402. Agent evaluation. The agent evaluation process includes: conducting multiple rounds of confrontation deductions, collecting the confrontation deduction data generated during the deduction process and the confrontation training process, analyzing and evaluating the confrontation deduction results according to the management of the evaluation index system, analyzing the intelligence ability level and agent effectiveness of the agents, and feeding back the evaluation and optimization results to the users and algorithm R & D personnel.

[0017] The beneficial effects of the present invention are: The present invention discloses an integrated system for the management, training, and verification of an intelligent decision-making algorithm for UAV air confrontation. Compared with the prior art, the improvements of the present invention are: Through the design of an integrated system for the management, training, and verification of an intelligent decision-making algorithm for unmanned aerial vehicle (UAV) air combat, this invention provides a service platform for the full-process development experiment of UAV intelligent decision-making algorithms, as well as algorithm design and development, iterative update, sharing, and reuse. It can effectively improve the development efficiency and application effect of UAV intelligent decision-making algorithms, while reducing the development and usage thresholds and costs, and promoting the rapid formation of research results on UAV intelligent decision-making algorithms. The aim is to improve the quality and efficiency of the development and application of UAV intelligent decision-making algorithms through the full-process application of intelligent decision-making algorithms from design, training, management, verification, and evaluation. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 FIG. is a schematic diagram of the composition of the integrated system for the management, training, and verification of the intelligent decision-making algorithm for UAV air combat according to this invention; Figure 2 FIG. is a flowchart of the use of the integrated system for the management, training, and verification of the intelligent decision-making algorithm for UAV air combat according to this invention; Figure 3 FIG. is a schematic diagram of the graphical configuration of the deep neural network structure of the integrated system for the management, training, and verification of the intelligent decision-making algorithm for UAV air combat according to this invention; Figure 4 FIG. is a schematic diagram of the agent training architecture process of the integrated system for the management, training, and verification of the intelligent decision-making algorithm for UAV air combat according to this invention; Figure 5 FIG. is a schematic diagram of the visual presentation interface for the algorithm performance evaluation of the integrated system for the management, training, and verification of the intelligent decision-making algorithm for UAV air combat according to this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] In order to enable those of ordinary skill in the art to better understand the technical solutions of this invention, the following further describes the technical solutions of this invention in conjunction with the drawings and embodiments. The following embodiments are used for illustrating this invention, but cannot be used to limit the scope of this invention.

[0020] Embodiment 1: Referring to an integrated system for the management, training, and verification of an intelligent decision-making algorithm for UAV air combat shown in the attached Figures 1-5 drawings, it includes: an intelligent decision-making algorithm library management subsystem, a UAV decision-making intelligent agent model design subsystem, an intelligent decision-making model training subsystem, and an intelligent decision-making model simulation verification and evaluation subsystem; 1. Intelligent decision-making algorithm library management subsystem Specifically, the intelligent decision-making algorithm library management subsystem is used to support users in secondary development of the algorithms in the air combat confrontation intelligent algorithm library according to their needs, and to manage distributed training, deployment, and operation scheduling of the algorithm library. The intelligent decision-making algorithm library management subsystem includes an algorithm secondary development module, an algorithm training management module, and an algorithm deployment and scheduling module; among them, (1) The algorithm secondary development module is used to support users in secondary development of the algorithms in the air combat confrontation intelligent algorithm library according to their needs, and provides programming call interfaces, neural network call component interfaces, agent training and prediction interfaces, and algorithm secondary development data flow interfaces.

[0021] Furthermore, when providing secondary development interfaces to support users in secondary development of the algorithms in the air combat confrontation intelligent algorithm library according to their needs, the algorithm source code, supporting documentation, and demonstration cases are provided to support users' development work, and important classes and functions in the source code are annotated. Among them, Programming call interface Through this interface, the algorithm editing template in the algorithm library can be opened. After the algorithm editing is completed, calls to reinforcement learning algorithms, transfer learning algorithms, incremental learning algorithms, few-shot learning algorithms, meta-learning algorithms, and machine learning algorithms can be completed.

[0022] Neural network call component interface Components for quickly building neural networks, including basic neural network components such as fully connected network MLP, convolutional neural network CNN, deep residual network ResNet, recurrent neural network RNN, long short-term memory network LSTM, gated recurrent unit GRU, attention mechanism Attention, feature transformation network Transformer, pointer network Pointer Net, etc.

[0023] Agent training and prediction interface It is convenient for users to set training parameters through the agent training interface and quickly start agent training; predict the situation through the agent prediction interface, output the original action information, and transmit it to the simulation environment after being converted into instructions.

[0024] Algorithm secondary development data flow interface Provides a complete set of algorithm secondary development interfaces. Users can input data and parameters according to the requirements of the development documentation and the format of the interface functions to interact with this platform. The algorithm secondary development data flow interface includes a situation information processor interface, an agent training interface, an agent prediction interface, an action command conversion interface, and a reward value calculation interface.

[0025] (2) The algorithm training management module is used to manage the algorithm training tasks and deploy and schedule the algorithms in the UAV air combat confrontation intelligent algorithm library; all algorithms in the air combat confrontation intelligent algorithm library support containerized deployment and scheduling, and support the training tasks to select three resources: the local service, GPUs server, and Kubernetes cluster. The Kubernetes cluster supports single-machine and distributed training modes, and also supports selecting multiple algorithms for training.

[0026] Containerized deployment and scheduling In order to perform distributed reinforcement learning training on the air combat confrontation intelligent agent to improve the intelligent level, all intelligent algorithms in the air combat confrontation intelligent algorithm library need to build a lightweight algorithm library image package under the Linux environment based on containerization technology. The algorithm library after containerized deployment can be cluster-scheduled for the deduction training of the intelligent agent.

[0027] Containerization technology can quickly and reliably deploy the intelligent algorithm library on a large scale and has the following advantages: Continuous operation: After the end of a simulation deduction process, containerization technology can enable the system to quickly reload the scenario file and start a new game, reducing the time overhead caused by system restart.

[0028] Stable fault tolerance: Based on containerization technology, the system can stably tolerate faults and distinguish error scenarios. If it is a fatal error (such as memory out-of-bounds, OOM, etc.), the system can crash; if it is a user command error, the user will be informed of the error message; if an error unacceptable in the business scenario occurs, the user will be informed and let the user decide whether to start a new game; if it is other errors, internal recovery will be attempted as much as possible (such as destroying the error unit or restoring the initial state, retry mechanism).

[0029] Resource controllability: System resources such as memory, log files, and file descriptors will not continue to skyrocket.

[0030] Independent operation of multiple environments on a single machine: Assuming that a simulation environment occupies 1.2 CPU cores, more than 60 environments can run concurrently on an 80-core single machine. Each environment independently uses local resources (specified port numbers, specified directories) instead of shared global resources (such as fixed port numbers, fixed absolute paths, or other limited system resources).

[0031] Limited resource acceleration operation: Under the condition of only using a single-digit number of CPU cores (no more than 3, and the CPU utilization rate of one environment is preferably below 150%) and ensuring the normal business logic, a simulation acceleration of more than 10 times can be achieved. Generally, the fewer CPU resources, the more concurrent running environments, the higher the acceleration ratio, and the more training rounds. Under these premises, the intelligent agent can reach a higher intelligent level faster.

[0032] The algorithm training management module manages the algorithm training tasks for the algorithms in the UAV air combat confrontation intelligent algorithm library, including: providing users with management functions for training tasks, supporting operations such as starting, terminating, viewing, viewing logs, selecting, editing, and deleting during the training process, supporting the selection of three types of resources for training tasks: local services, GPUs servers, and Kubernetes clusters. The Kubernetes cluster supports single-machine and distributed training methods, and also supports the selection of multiple algorithms for training.

[0033] Viewing training tasks: It is possible to view all currently created tasks and brief information of the tasks from the training task list page. It includes three aspects of information. One is task information, including running status, project name, creator, project group, all algorithms, creation time, number of CPU cores, number of GPUs, etc. The second is the resource template, including the number of training templates, sampling templates, prediction templates, number of CPU cores, number of GPUs, etc. The third is the node status, and the fourth is the configuration file.

[0034] Selecting training tasks: Supports the selection of three types of resources: local services, GPUs servers, and Kubernetes clusters. The Kubernetes cluster supports single-machine and distributed training methods, and also supports the selection of multiple algorithms for training.

[0035] Starting and terminating training tasks: Training tasks can be started and terminated.

[0036] Editing and deleting training tasks: Before the training task is started, the configuration file of the task can be edited. After the training task is started, the content in the "description" column can be edited, and the current task can also be deleted.

[0037] Viewing logs: Provides log records for the intelligent agent training experiment process. In the training logs, users can view various log files generated by the system for each individual in each component during training, which is convenient for users to conduct targeted and detailed analysis in case of task errors or when fine-tuning is required. And for the need of unified management and improved management efficiency, the system also generates global logs to help users more quickly determine which component has problems in case of errors, enabling users to locate component logs faster. Users can also view and download them locally on the platform.

[0038] (3) The algorithm deployment and scheduling module is used to schedule and manage the deployment and operation of the UAV air combat confrontation intelligent algorithm library, realize docking with third-party systems, and deploy and run the algorithm library and related operation support environments on private clouds and typical domestic public cloud platforms, supporting the unified management and invocation of the algorithm library by third-party systems.

[0039] 2. UAV decision-making intelligent agent model design subsystem Specifically, for the UAV decision-making agent model design subsystem, this subsystem provides a standardized development architecture for UAV air combat decision-making agents, supporting algorithm developers to develop agents for different air mission scenarios. It includes key element development and design modules such as the agent network structure design module, the observation situation input design module, and the action space output module, supporting the construction of single-agent and multi-agent models based on deep neural networks, real-time perception of the simulation situation, and issuing instructions to control the behavior of simulation entities.

[0040] Through neural network architecture design, according to the characteristics of mission requirements, mission units, equipment performance, etc. in the scenario, design the neural network structure and initial parameters of the agent, supporting the direct import of typical intelligent decision-making models from the algorithm library and parameter configuration adjustment, and also supporting the redesign of the model.

[0041] Through intelligent optimization algorithm design, according to the characteristics of mission requirements and control unit strength in the scenario and the requirements of the agent working mode, design and limit the architecture form and optimization objectives of the agent reinforcement learning algorithm.

[0042] (1) The agent network structure design module provides a programming framework for users to quickly build agents for different air mission scenarios, mainly composed of three sub-modules: a feature encoder, a feature aggregator, and a feature decoder. Users can customize the neural network structure of the agent based on this framework.

[0043] And for the graphical configuration of the neural network structure in the agent network structure design module, it can provide users with a graphical way to build the agent model. This system encapsulates various basic neural network models in neural networks such as decoders, encoders, and aggregators, and through the man-machine interaction interface of the system, provides users with visual construction methods such as dragging, moving, sequential linking, and editing, as Figure 3 shown. Users can drag and move each neural network component on the interface, such as the entity feature encoder, the general feature encoder, the GRU feature aggregator, and the ordered unit selection decoder. Users can link each neural network component on the interface according to the input and output order. For example, link the outputs of the entity feature encoder and the general feature encoder to the input of the GRU feature aggregator, and then link the output of the GRU feature aggregator to the input of the ordered unit selection decoder. Users can edit the parameters of each neural network component on the interface, and finally form the network architecture of the decision-making agent model in a typical scenario.

[0044] Feature encoder The feature encoder provides encoders for processing spatial features, entity features, and general features, used for feature encoding and feature extraction of multiple categories. The feature encoder includes a spatial feature encoder, an entity feature encoder, and a general feature encoder. Among them, The spatial feature encoder is mainly used to extract feature information of high-dimensional matrices such as images, such as extracting the information contained in the current environmental photo. Its network structure consists of an input layer, an intermediate layer, and an output layer. The input layer uses a fully connected neural network, the intermediate layer uses a convolutional neural network or a residual network, and the output layer uses a fully connected neural network. Among them, the number of hidden layers and the number of neurons in the output layer of the spatial feature encoder are both designed based on the network representation ability.

[0045] The entity feature encoder is mainly used to extract entity feature information, such as unit features such as the model and status of an aircraft. It first processes the processed features through a fully connected neural network once, then uses a transformer for further feature extraction, and finally performs pooling sampling. Among them, the number of hidden layers and the number of neurons in the output layer of the entity feature encoder are both designed based on the network representation ability.

[0046] The general feature encoder is mainly used to extract general feature information, such as statistical information in the environment such as the number of surviving aircraft. It mainly uses a multi-layer fully connected neural network to extract feature information. Among them, the number of hidden layers and the number of neurons in the output layer of the general feature encoder are both designed based on the network representation ability.

[0047] Feature aggregator Provides an aggregator for aggregating spatial features, entity features, and general features. The feature aggregator is used to aggregate the features processed by different feature encoders and splice them into complete information; the feature aggregator in the present invention includes a Dense feature aggregator, an LSTM feature aggregator, and a GRU feature aggregator.

[0048] The Dense feature aggregator is a simple fully connected feature aggregator. Users can simply splice and aggregate the features processed by different encoders through the Dense feature aggregator. It mainly uses a multi-layer fully connected neural network for output. Among them, the number of hidden layers and the number of neurons in the output layer of the Dense feature aggregator are both designed based on the network representation ability.

[0049] The LSTM feature aggregator is a feature aggregator that uses the recurrent neural network LSTM. For partially observable Markov process (POMDP) problems, the recurrent neural network LSTM can help remember past information and predict the future. Its principle is to splice multiple features and first process them through a multi-layer fully connected neural network, then process them through multiple modules (LSTM Cells) used to remember long-term dependencies, and finally use a fully connected network for output. Among them, the number of hidden layers and the number of neurons in the output layer of the LSTM feature aggregator are both designed based on the network representation ability.

[0050] The GRU feature aggregator is a feature aggregator that uses the recurrent neural network GRU. Similar to the LSTM feature aggregator, for partially observable Markov decision process (POMDP) problems, the recurrent neural network GRU can help memorize past information and predict the future. Its principle is to first splice multiple features and process them through a multi-layer fully connected neural network, then process them through multiple modules (GRU Cells) used to memorize long-term dependencies, and finally use a fully connected network for output. Among them, the number of hidden layers and the number of neurons in the output layer of the GRU feature aggregator are both designed based on the network representation ability.

[0051] Feature decoder The feature decoder provides a decoder that decodes features and outputs decision-making actions such as discrete actions, ordered unit selection, unordered unit selection, and single unit selection. The feature decoder is the exit of the entire neural network, used to decode the features aggregated by the aggregator and output actions, that is, to make decisions. The feature decoder of the present invention includes a discrete action decoder, an ordered unit selection decoder, an unordered unit selection decoder, and a single unit selection decoder.

[0052] The discrete action decoder is a classification action decoder used for discrete action decisions, such as walking forward and walking backward. Its principle is to first input the aggregated features into a multi-layer fully connected layer, then input them into a fully connected layer with the number of output neurons being the action dimension x to obtain a policy, and finally use policy sampling to obtain behavioral actions. Among them, the number of hidden layers of the discrete action decoder is designed based on the network representation ability, and the number of neurons in the output layer is designed based on the action space.

[0053] The ordered unit selection decoder is used for making decisions on the sequential selection of multiple units, that is, making decisions on sequential multi-agent selection. Among them, the number of hidden layers of the ordered unit selection decoder is designed based on the network representation ability, and the number of neurons in the output layer is designed based on the action space.

[0054] The unordered unit selection decoder is used for making decisions on the unordered selection of multiple units, that is, making decisions on unordered multi-agent selection. Among them, the number of hidden layers of the unordered unit selection decoder is designed based on the network representation ability, and the number of neurons in the output layer is designed based on the action space.

[0055] The single unit selection decoder is used for making decisions on the selection of a single unit, that is, making decisions on single-agent selection. Among them, the number of hidden layers of the single unit selection decoder is designed based on the network representation ability, and the number of neurons in the output layer is designed based on the action space.

[0056] (2) Observation Situation Input Design Module. By establishing various standard situation data structures, such as spatial situation representation, temporal situation representation, statistical situation representation, etc., it supports R & D personnel to input environmental situation data structures of intelligent game models according to unified standards, and realizes the reading of entity states of the simulation platform and the acquisition of overall situation by algorithms in the intelligent decision-making algorithm library management subsystem. This module includes an agent policy network situation input sub-module, an agent command decision-making training structure sub-module, and an observation situation input interface sub-module.

[0057] The agent policy network situation input sub-module obtains the status of our unit, the status of the enemy unit, and the overall situation of the deduction through the simulation environment docking interface, and inputs them into the situation data preprocessing unit for situation fusion, data completion, missing value repair, outlier processing, and normalization. In this solution, a fully connected network is used to extract general features, and a feature transformation network (transformer) is used to extract the features of the red and blue units respectively. After merging the three feature information, the historical information is processed through a GRU unit.

[0058] The agent command decision-making training structure sub-module performs situation fusion on the overall situation, the local situation of our unit, and the local situation of the enemy unit, and uses the processed feature information as the input of the neural network. Then, prediction models and several related intermediate layer model structures are generated using the situation fusion information, including memory read-write units, attention units, communication units, and reasoning units. These intermediate layer model structures will be combined and used to control the behavior of the agent from aspects such as force selection, action selection, position selection, and target selection. On the other hand, the input of the prediction model will be further processed to generate a situation assessment network. The situation assessment network directly assesses the winning rate on the current field, thereby indirectly affecting the decision-making of the agent model.

[0059] The observation situation input interface sub-module realizes the situation data interaction with the simulation environment. The interface is called by the battle framework using the RpyC remote call method, adopts the Protobuf data protocol, can be encapsulated in Python language, and the interface can determine different environment interface parameters according to the specific encapsulation form of the adversarial game software to realize the customization of data content.

[0060] The observation situation input interface mainly includes multi-dimensional information such as the number, position, status, and attributes of simulation entities.

[0061] (3) The action space output module, by establishing a variety of standard decision data structures, such as multiple-choice data structure, single-choice data structure, multi-classification data structure, binary classification data structure, continuous value prediction data structure, etc., supports R&D personnel to perform action modeling and development according to the interface, supports different personnel to output the action space data structure of the intelligent agent model according to a unified standard, and realizes the control of the intelligent agent over the simulation entity. This module consists of two parts: the intelligent agent output action generation submodule and the action space output interface submodule.

[0062] The agent output action generation submodule includes a target selection unit, a sensor selection unit and an action selection unit. It can decode the inference decision information into specific target selection instructions, sensor selection instructions and action selection instructions, use a fully connected network to generate meta-actions, use an attention mechanism to select red units and enemy targets respectively, and use a fully connected network to generate the x and y coordinates of the target position, which are output through the simulation environment interface and control the simulation entity.

[0063] The action space output interface submodule realizes the interaction with the task instructions in the simulation environment. The interface is called by the combat framework using the RpyC remote call method, adopts the Protobuf data protocol, and can be encapsulated in Python language.

[0064] The interface mode uses the mature gRPC framework, which is a remote procedure call framework that can be used on multiple platforms. It provides functions such as load balancing, tracking, intelligent monitoring, identity authentication, and high-speed connection, which can effectively reduce interface development and maintenance costs.

[0065] This solution is designed to use the mature gRPC interface framework to develop a simulation environment integration SDK. For each simulation environment, it provides a unified standard integrated development programming call interface and development environment, opens up the interface with the agent training / deduction environment, realizes the agent's control over the simulation unit, and uses the subject, predicate and object combination scheduling method to realize the agent's command control over the controlled unit.

[0066] 3. Intelligent decision model training subsystem Specifically, the intelligent decision model training subsystem is used to support virtual self-game training of intelligent agent models, such as synchronous symmetric self-game, asynchronous symmetric self-game, synchronous asymmetric self-game, asynchronous asymmetric self-game, etc., to improve the level of intelligent agents. Among them, the intelligent decision model training subsystem includes an intelligent agent training architecture design module and a typical intelligent agent training method module. The decision algorithm model trained by this subsystem will support sequential decision-making and output decision instructions based on the situation information received in real time. The intelligent agent also supports variable frequency decision-making and adjusts the decision frequency according to the task situation to meet the task requirements. Among them, (1) Agent training architecture design module The agent training architecture design module supports agent training, effectively utilizes resources without sacrificing data efficiency and resource utilization rate, and realizes functions such as generation, collection, feature processing, sequential decision-making of the agent, frequency conversion decision-making, and prediction of task success rate. Among them, the agent training architecture design module includes a data generation engine sub-module, a continuous learning engine sub-module, a prediction and inference engine sub-module, and an intelligent engine master control sub-module; ① The data generation engine sub-module is mainly used for interacting with the high-concurrency simulation environment to generate high-quality sample data. The sample data is sampled through multiple simulation environments and supplied to the continuous learning engine module for use.

[0067] ② The continuous learning engine sub-module is used to consume massive data and optimize the agent using relevant algorithms; ③ The prediction and inference engine sub-module is used to quickly respond and drive the simulation environment to generate data.

[0068] ④ The intelligent engine master control sub-module is used to save the state of the reinforcement learning task and the workflow, and is also responsible for the life cycle and resource management of the intelligent engine.

[0069] The traditional distributed reinforcement learning engine adopts a two-layer architecture of data generation - training. Among them, the data generation engine module is responsible for continuously generating training samples, and the continuous learning engine module is responsible for continuously updating the decision model parameters according to the training samples. The sampling engine of the traditional reinforcement learning engine runs entirely on the CPU, and the learning engine and the prediction engine run on the GPU. A large amount of data communication will occur between the CPU and the GPU. Due to the bottleneck of the network bandwidth, a large amount of data transmission and parameter synchronization of the neural network will reduce the working efficiency of the engine. In addition, the larger the parameters of the neural network model, the more obvious the speed reduction of forward inference using the CPU. Finally, the sampling engine needs to schedule CPU resources to perform simulation deduction and neural network inference simultaneously, and cannot fully improve the utilization rate of CPU multi-threaded resources.

[0070] Therefore, the agent training architecture designed in this solution adds an additional prediction and inference engine sub-module, adopts a three-layer architecture of data generation - training - prediction, uses GPU resources to accelerate the neural network inference speed, effectively solves several problems of the traditional distributed reinforcement learning engine, converts the large-scale computing power into large-scale data processing power, and realizes the efficient reinforcement learning training of the decision model under distributed computing power.

[0071] And the running data generation engine sub-module interacts with multiple simulation environments and performs high-speed network transmission through the virtual network management component to obtain simulation situation information. After completing the format conversion of the simulation environment data to the input data of the agent neural network (including deletion of useless data, data normalization, OneHot encoding, etc.), the neural network features are input into the prediction and inference engine module, which calculates and outputs the neural network decision output. After completing the mapping from the neural network decision output to the simulation environment instructions, the data generation engine module sends specific instructions to the simulation environment. The training data formed during this process, namely the <State, Action, Reward> triple data stream, is input into the training data buffer pool for use by the subsequent continuous learning engine module.

[0072] After the training data is formed, the continuous learning engine module asynchronously reads the training data buffer pool. Here, "asynchronously" means that the containerized deployed continuous learning engine module reads the data in real time, rather than waiting for the slowest training data to arrive and then reading it uniformly. The continuous learning engine module calls the reinforcement learning algorithm, calculates the gradient through the loss function (Loss), determines the update direction of the agent neural network, performs gradient backpropagation, and updates the neural network parameters. During the update process of the neural network model parameters, the Ring-AllReduce method is adopted based on the virtual network management component. After the update is completed, the neural network model parameters are synchronized to the prediction and inference engine module.

[0073] (2) The typical agent training method module is used to combine with the agent training architecture design module and train the agent using different training methods. Among them, the training methods include 4 methods such as synchronous symmetric self-play, asynchronous symmetric self-play, synchronous asymmetric self-play, and asynchronous asymmetric self-play. The specific content of these multiple methods is as follows: ① Synchronous symmetric self-play, that is, synchronous symmetric self-play adversarial training, supports that both the red and blue sides of the training are deep reinforcement agents and the neural network models are exactly the same. First, complete the sampling task of this round for the working node. It needs to wait for all other sampling nodes to complete their respective sampling tasks of this round. That is, after all working nodes complete all sampling tasks of this round, the working node pushes its sampling results of this round to the parameter server or the sample cache pool, starts training, updates the model, and then starts the next sampling task of this round.

[0074] ② Asynchronous symmetric self - game, that is, asynchronous symmetric self - game adversarial training, supports that both the red and blue sides in the training are deep reinforcement agents and the neural network models are exactly the same. The working node that first completes the current round of sampling task does not need to wait for all other sampling nodes to complete their respective current - round sampling tasks. That is, after all working nodes complete all current - round sampling tasks, the working node that has completed first directly pushes its current - round sampling results to the parameter server or the sample cache pool, starts training, updates the model, and then starts the next round of sampling task.

[0075] ③ Synchronous asymmetric self - game, that is, synchronous asymmetric self - game adversarial training, supports that both the red and blue sides in the training are deep reinforcement agents, but the neural network models of the red side and the blue side are not the same. The working node that first completes the current round of sampling task needs to wait for all other sampling nodes to complete their respective current - round sampling tasks. That is, after all working nodes complete all current - round sampling tasks, the working node pushes its current - round sampling results to the parameter server or the sample cache pool, starts training, updates the model, and then starts the next round of sampling task.

[0076] ④ Asynchronous asymmetric self - game, that is, asynchronous asymmetric self - game adversarial training, supports that both the red and blue sides in the training are deep reinforcement agents, but the neural network models of the red side and the blue side are not the same. The working node that first completes the current round of sampling task does not need to wait for all other sampling nodes to complete their respective current - round sampling tasks. That is, after all working nodes complete all current - round sampling tasks, the working node that has completed first directly pushes its current - round sampling results to the parameter server or the sample cache pool, starts training, updates the model, and then starts the next round of sampling task.

[0077] 4. Intelligent Decision - Making Model Simulation Verification and Evaluation Sub - system Specifically, the intelligent decision - making model simulation verification and evaluation sub - system is used to provide simulation deduction, agent ability verification, and human - machine confrontation ability, and evaluate according to the data of simulation deduction, agent ability verification, and human - machine confrontation ability. It supports functions such as scenario editing, air task scenario construction, digital earth, 3D visualization display, digital earth, etc. It pre - sets typical anti - UAV system, manned aircraft, and UAV simulation models, realizes the interactive docking with the air game confrontation intelligent algorithm library, and conducts demonstration verification on the agent model, supporting the human - machine confrontation function. The intelligent decision - making model simulation verification and evaluation sub - system also includes the following modules: ① Scenario editing module, which supports users to design air game confrontation scenarios, supports constructing air task scenarios such as short - range, medium - range, and long - range, as well as the opposite - target penetration and strike task scenarios, and supports two / three - dimensional synchronous editing of scenario files.

[0078] The scenario editing module completes the conversion process from "mission scenario" to "simulation scenario". By constructing the basic elements of the behavior logic of the simulation and deduction platform, it realizes the decision-making behavior models of various units, forms a simulation scenario and stores it in a formal standard format. At the same time, the stored simulation scenario can be conveniently managed and used. It can flexibly define the application principles, the processes and rules of behavior decision-making, and can achieve dynamic modification during the simulation process, achieving the separation of the behavior model from the simulation engine and the physical model of the equipment, thereby improving the flexibility, reusability and development efficiency of simulation modeling.

[0079] ② The digital earth module provides real-time simulation and deduction functions based on the three-dimensional digital earth display environment, supports the deployment of local map servers, and users can conduct air mission game confrontation deductions in the digital scene based on the real terrain, landforms, and landscapes, including functions such as global elevation images, oblique photography, refined local scene display, and real sea surface rendering.

[0080] The digital earth supports the deployment of local map servers and supports the import of terrain, images, and three-dimensional models in standard formats (such as TMS, WMS, 3dtiles, etc.) into the digital earth, thus having the ability to quickly build scenario scenarios according to requirements. Traditional simulation and deduction systems are mainly based on two-dimensional digital earth, while traditional three-dimensional digital earth is mainly used for visualization display. This system combines simulation and deduction with three-dimensional digital earth for the first time, enabling users to conduct simulation experiments or simulation training in the digital scene based on the real terrain, landforms, and landscapes.

[0081] ③ The simulation model module deploys simulation models (including three-dimensional models and performance simulation models) of typical anti-drone systems, manned aircraft, unmanned aircraft, etc., to support simulation and deduction in typical air combat scenarios.

[0082] The platform pre-sets more than 31 simulation models (including three-dimensional models and performance simulation models) of typical anti-drone systems, manned aircraft, and unmanned aircraft, supports the import of third-party manned aircraft and unmanned aircraft simulation models (including three-dimensional physical models and performance simulation models) according to standard file formats, and supports the editing and replacement of payload models. Simulation models and rules are the basis for game confrontation simulation and deduction, and are also the key to reflecting the realism and complexity of the mission scenario. Build a simulation model system covering three categories: mission environment, entity units, and action rules to support typical air game confrontation simulation and deduction.

[0083] ④ The simulation and deduction module provides simulation and deduction and man-machine confrontation functions in typical air game confrontation scenarios, and at the same time meets the rapid training requirements of intelligent game algorithms, realizing accelerated simulation under the conditions of multi-sample large-scale complex confrontation. It includes functions such as deduction acceleration, man-machine confrontation, and confrontation situation display.

[0084] It has time advancement, model scheduling operation, simulation control, and data management during the simulation process, and provides interaction for the situation display and guidance control during the operation process, which is the basis for carrying out air combat confrontation simulation deduction. The simulation deduction engine follows the design of generality and flexibility and provides the following modular functions: The simulation process driving sub-module drives the appropriate model or simulation system to complete the simulation of the scenario plan and generates corresponding results.

[0085] The simulation operation control sub-module completes the initialization data loading and setting of the simulation, the management and control of the simulation operation, and the simulation time advancement according to the equipment information, scenario data, and action plan in the simulation input data. The system supports the deduction acceleration function, and the acceleration ratio is not less than 10 times the speed, and the maximum can reach 100 times the speed.

[0086] The simulation status monitoring sub-module monitors the states of various entities and entity interactions during the simulation process, collects information such as the states of simulation entities, model data, and entity interactions, and is used to establish simulation logs and simulation process data, providing necessary support for post-event analysis, replay, and situation display.

[0087] ⑤ The intelligent decision-making algorithm evaluation module provides the function of evaluating the training effect of the intelligent agent, supports the user to customize the training effect evaluation index, provides the function of automatically generating the training effect curve of the intelligent agent, and gradually converges as the intelligent agent training progresses. All algorithms in the algorithm library support the user to customize the training effect evaluation index. The system will comprehensively evaluate the intelligent agent model, algorithm design, AI capabilities, etc. according to the game confrontation data and deduction results, and feedback the evaluation results to the users and algorithm R & D personnel to realize the iterative optimization of the intelligent agent. The intelligent decision-making algorithm evaluation module includes two parts: the intelligent evaluation index management sub-module and the intelligent agent evaluation feedback calculation sub-module.

[0088] The evaluation index management sub-module is responsible for constructing an evaluation index system for the intelligent agent around the task objectives and content according to the evaluation requirements, forming the calculation method of the index, and the judgment standard for the decision-making ability of the intelligent agent of the decision-making intelligent agent model, supporting the iterative optimization of the intelligent agent.

[0089] The intelligent agent evaluation feedback calculation sub-module is responsible for receiving the intelligent game simulation confrontation deduction data, and realizing functions such as the conversion and processing of the deduction data and the calculation of the deduction index results according to the requirements of the evaluation index. It supports the calculation of effectiveness evaluation indexes such as the task completion situation and the AI decision-making battle loss ratio from the aspect of the intelligent agent decision-making ability. The evaluation result visualization interface is as Figure 5 shown.

[0090] As Figure 2 shown, the present invention also provides a usage process of an integrated system for the management, training, and verification of the intelligent decision-making algorithm for unmanned aerial vehicle air combat. The steps include: S1. Simulation Environment Docking Provide a simulation environment for the training of the intelligent decision-making model. The main tasks in this stage are the containerization transformation of the simulation environment, the docking of the simulation platform interface, and the design of the air mission scenario. Package the simulation environment into a container and deploy it to the adversarial knowledge graph platform and the intelligent game algorithm library system, and provide the interface docking between the adversarial knowledge graph platform and the intelligent game algorithm library system and the typical true 3D air combat simulation software, including the deduction control interface, the situation interface, and the control instruction interface. Users can conduct scenario design based on the actual tasks of the air combat business scenario.

[0091] S2. Agent Design S201. Neural Network Architecture Design Design the neural network structure and initial parameters of the agent according to the task requirements, control units, and equipment performance characteristics of the scenario. This step supports directly importing typical intelligent decision-making models from the algorithm library and adjusting the parameter configuration, and also supports the re-design of the model.

[0092] S202. Intelligent Optimization Algorithm Design Design and define the architecture form and optimization objectives of the agent's reinforcement learning algorithm according to the task requirements of the scenario, the characteristics of the control unit strength, and the requirements of the agent's working mode.

[0093] S3. Agent Training S301. Design of Evolutionary Training Method; Design the agent's evolutionary training method according to the task constraints and the designed intelligent game algorithm; S302. Distributed Training of Agents; Conduct large-scale distributed deep reinforcement learning training on the agent to improve the agent's ability level; S303. Migration and Deployment of Agents; Migrate and deploy the trained agent to the typical true 3D air combat simulation environment for adversarial deduction.

[0094] S4. Agent Verification and Evaluation S401. Management of Evaluation Index System According to the evaluation requirements, around the task objectives and content, construct an evaluation index system and model for the agent, form the calculation method of the index, and the judgment criteria for the decision-making ability, intelligence level, and agent effectiveness of the agent model.

[0095] S402. Agent Evaluation Conduct multiple rounds of adversarial simulations, collect the adversarial simulation data generated during the simulation process and the adversarial training process, analyze and evaluate the results of the adversarial simulations according to the management of the evaluation index system, analyze the intelligence level and effectiveness of the intelligent agents, and feedback the evaluation and optimization results to the users and algorithm R & D personnel to optimize the design of the intelligent agent architecture and training algorithms.

[0096] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.

Claims

1. An integrated system for management, training and verification of intelligent decision-making algorithms for aerial confrontation of unmanned aerial vehicles, characterized in that: include: The intelligent decision-making algorithm library management subsystem is used to support users to carry out secondary development of algorithms in the air game confrontation intelligent algorithm library according to their needs, and to carry out distributed training management, deployment and operation scheduling management of the algorithms in the algorithm library; The UAV decision-making agent model design subsystem provides a standardized UAV air confrontation decision-making agent development architecture, supports algorithm developers to develop agents for different air mission scenarios, supports the construction of single-agent and multi-agent models based on deep neural networks, perceives simulation situations in real time, and issues instructions to control the behavior of simulation entities; Intelligent decision-making model training subsystem, used to support virtual self-game training of intelligent agent models and improve the level of intelligent agents; The intelligent decision-making model simulation verification and evaluation subsystem is used to provide simulation deduction, intelligent agent capability verification and human-machine confrontation capability, and to perform evaluation based on simulation deduction, intelligent agent capability verification and human-machine confrontation capability data.

2. A system according to claim 1, characterized in that: The intelligent decision algorithm library management subsystem includes: The secondary development module is used to support users to carry out secondary development of the algorithms in the air game confrontation intelligent algorithm library according to their needs, and provides programming call interface, neural network call component interface, intelligent agent training and prediction interface and algorithm secondary development data flow interface; Among them, the programming call interface can open the algorithm editing template in the air game confrontation intelligent algorithm library to complete the call of reinforcement learning algorithm, transfer learning algorithm, incremental learning algorithm, small sample learning algorithm, meta-learning algorithm and machine learning algorithm; The neural network calling component interface is used to quickly build the components of the neural network. The components of the neural network include: fully connected network MLP, convolutional neural network CNN, deep residual network ResNet, recurrent neural network RNN, long short-term memory network LSTM, recurrent neural network GRU, attention mechanism Attention, feature conversion network Transformer, pointer network Pointer Net; The agent training and prediction interface allows users to set training parameters, start agent training, predict situations, output raw action information, and convert it into instructions before transmitting it to the simulation environment. The algorithm secondary development data flow interface is used to input data and parameters to complete data interaction; the algorithm secondary development data flow interface includes: situation information processor interface, intelligent agent training interface, intelligent agent prediction interface, action command conversion interface and reward value calculation interface; The algorithm training management module is used to manage and deploy the algorithms in the UAV air game confrontation intelligent algorithm library; The algorithm deployment and scheduling module is used to schedule and manage the deployment and operation of the drone aerial game confrontation intelligent algorithm library, realize the connection with the third-party system, and deploy and run the algorithm library and related operation support environment on private clouds and typical domestic public cloud platforms, supporting the unified management and call of the algorithm library by third-party systems.

3. A system according to claim 2, characterized in that: The algorithms in the air game confrontation intelligent algorithm library all support containerized deployment and scheduling; The algorithm training management module manages the algorithm training tasks of the algorithms in the drone air game confrontation intelligent algorithm library, including: supporting start, stop, view, log view, selection, edit, and delete operations during the training process, and supporting the selection of local services, GPUs servers, and Kubernetes clusters for training tasks. Among them, the Kubernetes cluster supports stand-alone and distributed training methods, and also supports the selection of multiple algorithms for training.

4. A system according to claim 1, characterized in that: The UAV decision-making intelligent agent model design subsystem includes: The intelligent agent network structure design module provides users with a framework for quickly building intelligent agent programming for different air mission scenarios. Users can customize the neural network structure of the intelligent agent based on the framework. The intelligent agent network structure design module includes a feature encoder, a feature aggregator, and a feature decoder. The observation situation input design module, by establishing a variety of standard situation data structures, supports R&D personnel to input the environment situation data structure of the intelligent game model according to unified standards, and realizes the entity state reading and overall situation acquisition of the simulation platform in the intelligent decision-making algorithm library management subsystem; the observation situation input design module includes an intelligent agent strategy network situation input submodule, an intelligent agent command decision training structure submodule and an observation situation input interface submodule; wherein, the intelligent agent strategy network situation input submodule obtains the state of our unit, the state of the enemy unit and the deduced global situation through the simulation environment docking interface, and inputs it to the situation data preprocessing unit to perform situation fusion, data completion, missing value repair, outlier processing and normalization operations, the intelligent agent command decision training structure submodule performs situation fusion on the global situation, the local situation of our unit and the local situation of the enemy unit, and uses the processed feature information as the input of the neural network, and the observation situation input interface submodule realizes the situation data interaction with the simulation environment; The action space output module supports R&D personnel to model and develop actions according to the interface by establishing multiple standard decision data structures; The action space output module includes an agent output action generation submodule and an action space output interface submodule; the action space output module includes an agent output action generation submodule and an action space output interface submodule; wherein the agent output action generation submodule includes a target selection unit, a sensor selection unit and an action selection unit, which can decode into specific target selection instructions, sensor selection instructions and action selection instructions according to the reasoning decision information, use a fully connected network to generate meta-actions, use an attention mechanism to select red units and enemy targets respectively, use a fully connected network to generate the x and y coordinates of the target position, output through the simulation environment interface and control the simulation entity, and the action space output interface submodule realizes interaction with the task instructions in the simulation environment.

5. A system according to claim 4, characterized in that: The feature encoder is used for feature encoding and feature extraction of various categories, and the feature encoder includes: a spatial feature encoder, an entity feature encoder and a universal feature encoder, wherein the spatial feature encoder is used to extract feature information of high-dimensional matrices such as images, the entity feature encoder is used to extract entity feature information, and the universal feature encoder is used to extract universal feature information; The feature aggregator is used to aggregate the features processed by different feature encoders and splice them into complete information; the feature aggregator includes: a Dense feature aggregator, an LSTM feature aggregator and a GRU feature aggregator; The feature decoder is used to decode and output the features aggregated by the feature aggregator; the feature decoder includes: a discrete action decoder, an ordered unit selection decoder, an unordered unit selection decoder and a single unit selection decoder, the discrete action decoder is used for making decisions on discrete actions, the ordered unit selection decoder is used for making ordered selections of multiple units, the unordered unit selection decoder is used for making unordered selections of multiple units, and the single unit selection decoder is used for making decisions on single unit selections; In the observation situation input design module, the multiple standard situation data structures include spatial situation representation, temporal situation representation, and statistical situation representation; In the action space output module, the multiple standard decision data structures include a multiple-choice data structure, a single-choice data structure, a multi-classification data structure, a binary classification data structure, and a continuous value prediction data structure.

6. A system according to claim 1, characterized in that: The intelligent decision model training subsystem includes: The intelligent agent training architecture design module is used to support intelligent agent training and realize the generation, collection, feature processing, sequential decision-making, variable frequency decision-making, and task success rate prediction of sample data without sacrificing data efficiency and resource utilization; The typical intelligent agent training method module is used to combine the intelligent agent training architecture design module to train the intelligent agent using different training methods, where the different training methods include synchronous symmetric self-game, asynchronous symmetric self-game, synchronous asymmetric self-game, and asynchronous asymmetric self-game.

7. A system according to claim 6, characterized in that: The agent training architecture design module includes: The data generation engine submodule is used to interact with the high-concurrency simulation environment to generate high-quality sample data; The continuous learning engine submodule is used to consume massive data and optimize the agent; Prediction and inference engine submodule, used to quickly respond and drive the simulation environment to generate data; The intelligent engine master control submodule is used to preserve the state and workflow of reinforcement learning tasks, and is responsible for the life cycle and resource management of the intelligent engine.

8. A system according to claim 1, characterized in that: The intelligent decision-making model simulation verification and evaluation subsystem includes: The scenario editing module supports users to design air game confrontation scenarios, build short-range, medium-range and long-range air mission scenarios and target penetration and strike mission scenarios, and simultaneously edit the two-dimensional and three-dimensional scenario files; The digital earth module provides real-time simulation and deduction functions based on the three-dimensional digital earth display environment, and supports the deployment of localized map servers; The simulation model module is used to deploy typical anti-UAV systems, manned aircraft, and UAV simulation models, build a simulation model system covering mission environment, entity units, and action rules, and support typical simulation deductions in air confrontation scenarios; The simulation module is used to provide simulation and man-machine confrontation functions in typical air game confrontation scenarios, while meeting the needs of rapid training of intelligent game algorithms; Wherein, the simulation deduction module includes: The simulation process driving submodule is used to drive the model or simulation system to complete the simulation of the scenario and generate corresponding results; The simulation operation control submodule completes the initialization data loading and setting of the task simulation, the management and control of the simulation operation, and the advancement of the simulation time according to the equipment information, assumption data, and action plan in the input data; The simulation status monitoring submodule is used to monitor the status of each entity and entity interaction during the simulation process, collect simulation entity status, model data, and entity interaction information, and is used to establish simulation logs and simulation process data; The intelligent decision-making algorithm evaluation module supports users to customize training effect evaluation indicators, automatically generates intelligent agent training effect curves, and gradually converges as the intelligent agent training proceeds.

9. A system according to claim 8, characterized in that: The intelligent decision-making algorithm evaluation module includes: The agent evaluation feedback calculation submodule is responsible for building an evaluation index system for the agent based on the evaluation requirements, around the task objectives and content, forming the calculation method of the index, the judgment criteria of the decision-making ability of the agent decision-making body, and supporting the iterative optimization of the agent; The intelligent agent evaluation feedback calculation submodule is responsible for receiving intelligent game simulation confrontation deduction data, and according to the evaluation index requirements, realizes the deduction data conversion processing and deduction index result calculation, and supports the calculation of intelligent agent task capabilities, task completion status and AI decision-making cost-effectiveness evaluation indicators.

10. A method for using the integrated system for management, training and verification of intelligent decision-making algorithms for aerial confrontation of unmanned aerial vehicles according to claim 1, characterized in that: The method comprises: S1. Simulation environment docking; the specific process of the simulation environment docking includes: S101. Containerize and package the simulation environment, and deploy the packaged information to the adversarial knowledge graph platform and intelligent game algorithm library system; S102. The simulation platform interface docking provides the interface docking between the confrontation knowledge graph platform and the intelligent game algorithm library system and the typical true three-dimensional air confrontation simulation software; S103. Air mission scenario design: based on the actual mission of air confrontation business scenario, conduct scenario design; S2. Agent design, the specific process of the agent design includes: S201. Neural network architecture design. The neural network architecture design process includes: designing the intelligent agent neural network structure and initial parameters according to the mission requirements, control units, and equipment performance of the assumed scenario; S202. Intelligent optimization algorithm design. The intelligent optimization algorithm design process includes: designing and defining the architecture and optimization objectives of the intelligent agent reinforcement learning algorithm according to the task requirements of the assumed scenario, the characteristics of the control unit force, and the requirements of the intelligent agent working mode; S3. Agent training, the agent training process includes: S301. Design of evolutionary training method. The design process of evolutionary training method includes: designing an intelligent agent evolutionary training method according to task constraints and designed intelligent game algorithm; S302. Agent distributed training, the agent distributed training process includes: large-scale distributed deep reinforcement learning training of the agent; S303. Agent migration deployment, the agent migration deployment process is: the trained agent is migrated and deployed to a typical true three-dimensional aerial confrontation simulation environment for confrontation deduction; S4. Agent verification and evaluation, the agent verification and evaluation process includes: S401. Evaluation index system management. The evaluation index system management process includes: building an evaluation index system and model for the intelligent agent based on the evaluation requirements, focusing on the task objectives and content, forming the calculation method of the index, and the evaluation criteria of the decision-making ability, intelligence level and effectiveness of the intelligent agent model; S402. Intelligent agent evaluation. The intelligent agent evaluation process includes: conducting multiple rounds of adversarial simulations, collecting adversarial simulation data generated by the simulation process and adversarial training process, analyzing and evaluating the adversarial simulation results according to the evaluation index system management, analyzing the intelligent ability level and intelligent agent performance of the intelligent agent, and feeding back the evaluation optimization results to users and algorithm developers.