An Optimization Training Method, System, Device and Medium for UAV Air Combat Decision Making

By presetting multiple enemy maneuvering strategies and using deep reinforcement learning models to train the drone, the problems of slow convergence speed and ineffective attempts during the training process are solved, and more efficient training and optimization effects are achieved.

CN115062888BActive Publication Date: 2025-05-27NAT UNIV OF DEFENSE TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210264535.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-17
Publication Date
2025-05-27
Estimated Expiration
2042-03-17

AI Technical Summary

Technical Problem

When training drones, the algorithm converges slowly, cannot effectively accumulate rewards, and it is difficult to find the optimal solution, resulting in a large number of ineffective explorations and useless attempts by the drone.

Method used

Preset a variety of enemy maneuver strategies for enemy drones, train our drones to fight against simple to complex enemy maneuver strategies through deep reinforcement learning models, and optimize training effects and efficiency.

Benefits of technology

The convergence speed of deep reinforcement learning algorithms has been accelerated, the evaluation reward return value has been improved, and the training efficiency of the drone and the quality of maneuvering strategies have been improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115062888B_ABST
    Figure CN115062888B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, system, device and medium for optimizing the training of unmanned aerial vehicle (UAV) air combat decision-making. The method presets a variety of enemy maneuver strategies for enemy UAVs; among them, various enemy maneuver strategies are related but have different complexities; through a deep reinforcement learning model, the UAV of our side is trained to make maneuver decisions against enemy maneuver strategies from simple to complex. The present invention can improve the training effect of the UAV of our side and obtain better maneuver strategies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of unmanned aerial vehicles, and particularly to a method, a system, a device and a medium for optimizing the training of unmanned aerial vehicle air combat decision-making. Background Art

[0002] One of the core ideas of reinforcement learning is to learn by continuously trying to accumulate experience. Just like the human brain learning knowledge, it is very important to learn step by step. When directly trying to solve particularly complex problems without a certain foundation, the desired effect is often not obtained. When training our own unmanned aerial vehicle, the convergence speed of the algorithm is slow. Through analyzing the visual demonstration of the training process, it is found that this is because the enemy unmanned aerial vehicle that has been using the MINMAX strategy in the initial stage of training is too powerful, leaving very few opportunities for our own unmanned aerial vehicle to "learn by trial and error". Therefore, it is difficult to effectively accumulate rewards and find the optimal solution. Our own unmanned aerial vehicle will make a large number of ineffective explorations such as staying away from the combat site, crashing to the ground, flying into the sky, etc., and unhelpful attempts such as still decelerating after being rear-ended or turning away after being rear-ended, and cannot obtain a better maneuvering strategy. Summary of the Invention

[0003] To solve the problems existing in the prior art, the present invention provides a method, a system, a device and a medium for optimizing the training of unmanned aerial vehicle air combat decision-making, which can optimize the training effect and efficiency on the basis of training the unmanned aerial vehicle to make autonomous maneuvering decisions to obtain combat advantages.

[0004] To achieve the above object, the technical solution of the present invention is realized as follows:

[0005] In a first aspect, the present invention provides a method for optimizing the training of unmanned aerial vehicle air combat decision-making, including the steps of:

[0006] Presetting a variety of enemy maneuvering strategies for enemy unmanned aerial vehicles; wherein, the various enemy maneuvering strategies are related but have different complexities;

[0007] Training the maneuvering strategies of our own unmanned aerial vehicle against the enemy maneuvering strategies from simple to complex through a deep reinforcement learning model.

[0008] Compared with the prior art, the present invention has the following beneficial effects:

[0009] The present application presets a variety of enemy maneuvering strategies for enemy unmanned aerial vehicles, and the various enemy maneuvering strategies are related but have different complexities. The maneuvering strategies of our own unmanned aerial vehicle trained through a deep reinforcement learning model against the enemy maneuvering strategies from simple to complex can accelerate the convergence speed of the deep reinforcement learning algorithm, obtain a higher evaluation reward return value faster, and improve the training efficiency of our own unmanned aerial vehicle.

[0010] Further, the maneuver strategy of our UAV trained by the deep reinforcement learning model to counter the enemy maneuver strategies from simple to complex includes:

[0011] Preset multiple groups of training rounds;

[0012] Train the maneuver strategy of our UAV to counter at least one of the enemy maneuver strategies within each group of training rounds; wherein, at least one of the enemy maneuver strategies countered within each group of training rounds is selected from all the enemy maneuver strategies according to a preset selection rule.

[0013] Further, the multiple enemy maneuver strategies include at least one of the following:

[0014] The enemy UAV maintains its course and moves in a uniform straight line;

[0015] The enemy UAV randomly selects an action from the maneuver action library;

[0016] The enemy UAV makes a maneuver decision according to the MINMAX strategy.

[0017] Further, it includes:

[0018] Couple the maneuver strategy of our UAV trained by the deep reinforcement learning model with the preset rules to obtain the maneuver strategy of our UAV.

[0019] Further, the coupling of the maneuver strategy of our UAV trained by the deep reinforcement learning model with the preset rules to obtain the maneuver strategy of our UAV includes:

[0020] Based on the experience of fighter pilots and the research results of air combat experts, summarize combat skills and form preset rules;

[0021] Couple the maneuver strategy of our UAV trained by the deep reinforcement learning model with the preset rules to obtain the maneuver strategy of our UAV.

[0022] In a second aspect, the present invention provides an optimized training system for UAV air combat decision-making, including:

[0023] A maneuver strategy preset unit for presetting multiple enemy maneuver strategies of the enemy UAV; wherein, the various enemy maneuver strategies are related but have different complexities;

[0024] A maneuver strategy learning unit for training the maneuver strategy of our UAV to counter the enemy maneuver strategies from simple to complex by means of a deep reinforcement learning model.

[0025] Compared with the prior art, the present invention has the following beneficial effects:

[0026] This application uses a maneuver strategy preset unit to preset various enemy maneuver strategies for enemy unmanned aerial vehicles (UAVs). The various enemy maneuver strategies are related but have different complexities. A maneuver strategy learning unit is used to train our own maneuver strategies to counter the enemy maneuver strategies from simple to complex through a deep reinforcement learning model, which can accelerate the convergence speed of the deep reinforcement learning algorithm, obtain a higher evaluation reward return value faster, and improve the training efficiency of our UAVs.

[0027] In a third aspect, the present invention provides a UAV air combat decision optimization training device, including at least one control processor and a memory communicatively connected to the at least one control processor; the memory stores instructions executable by the at least one control processor, and when the instructions are executed by the at least one control processor, the at least one control processor is enabled to execute a UAV air combat decision optimization training method as described above.

[0028] In a fourth aspect, the present invention provides a computer-readable storage medium storing computer-executable instructions for causing a computer to execute a UAV air combat decision optimization training method as described above. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the following description of embodiments in conjunction with the accompanying drawings, where:

[0030] Figure 1 is a flowchart of a UAV air combat decision optimization training method provided by an embodiment of the present invention;

[0031] Figure 2 is a simulation result diagram of a scenario transfer optimization training method provided by an embodiment of the present invention;

[0032] Figure 3 is a training time comparison diagram of a scenario transfer optimization training method provided by an embodiment of the present invention;

[0033] Figure 4 is a simulation result diagram of a soft update scenario transfer optimization method provided by an embodiment of the present invention;

[0034] Figure 5 is a simulation trajectory diagram of an actual air combat between our and enemy UAVs provided by an embodiment of the present invention;

[0035] Figure 6 is a schematic diagram of the actions of a barrel roll maneuver strategy provided by an embodiment of the present invention;

[0036] Figure 7The simulation trajectory diagram of the rule coupling optimization training method provided by an embodiment of the present invention;

[0037] Figure 8 The structural diagram of an unmanned aerial vehicle air combat decision-making optimization training system provided by an embodiment of the present invention. Detailed implementation manners

[0038] Next, the technical solutions of the embodiments of the present disclosure will be described clearly and completely with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present disclosure. It should be noted that, without conflict, the embodiments of the present disclosure and the features in the embodiments may be combined with each other. In addition, the accompanying drawings are used to supplement the description of the text part of the specification, enabling people to visually and vividly understand each technical feature and the overall technical solution of the present disclosure, but they should not be construed as limiting the protection scope of the present disclosure.

[0039] In the description of the present invention, the features defined as "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise specified, the meaning of "a plurality" is two or more.

[0040] In the description of the present invention, unless otherwise clearly defined, words such as "set", "installed", and "connected" should be understood in a broad sense. Those skilled in the art can reasonably determine the specific meanings of the above words in the present invention in combination with the specific content of the technical solution.

[0041] Next, the technical solutions in the embodiments of the present invention will be described clearly and completely with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0042] When training our own unmanned aerial vehicle, the speed of algorithm convergence is very slow. Through analyzing the visual demonstration of the training process, it is found that this is because the enemy unmanned aerial vehicle that has been using the MINMAX strategy is too powerful in the initial stage of training, leaving very few opportunities for our own unmanned aerial vehicle to "learn from mistakes", so it is impossible to effectively accumulate rewards and it is difficult to find the optimal solution. Our own unmanned aerial vehicle will make a large number of ineffective explorations such as staying away from the battlefield, crashing to the ground, and flying into the sky, and unhelpful attempts such as still decelerating after being rear-ended or turning away after being rear-ended. Therefore, it is impossible to obtain a good maneuvering strategy.

[0043] To solve the above problems, this application presets multiple enemy maneuver strategies for enemy drones, and the complexities of the multiple enemy maneuver strategies are different. Our drone trains our maneuver decisions to counter the enemy maneuver strategies from simple to complex through a deep reinforcement learning model. Our drone is trained based on multiple enemy maneuver strategies, which can accelerate the convergence speed of the deep reinforcement learning algorithm, obtain a relatively high evaluation reward return value, and improve the training effect of our drone.

[0044] Referring to Figures 1 to 7 , an embodiment of the present invention provides a method for optimizing the training of drone air combat decisions, including the steps:

[0045] (1) Preset multiple enemy maneuver strategies for enemy drones; among them, various enemy maneuver strategies are related but have different complexities.

[0046] Specifically, this embodiment designs a scenario transfer optimization training method. This method presets three training scenarios from simple to complex, which respectively correspond to three enemy maneuver strategies of enemy drones from simple to complex. That is, the complexities of the three enemy maneuver strategies are different according to different scenarios. The simple scenario corresponds to the first enemy maneuver strategy of the enemy maneuver strategy. The initial state of the enemy drone is random, and it maintains a constant heading and moves in a straight line at a constant speed after starting the simulation; the transition scenario corresponds to the second enemy maneuver strategy of the enemy maneuver strategy, and the drone randomly selects an action from the maneuver action library at each decision moment; the complex scenario corresponds to the third enemy maneuver strategy of the enemy maneuver strategy, and the enemy drone makes a maneuver decision according to the minimax strategy (that is, the MINMAX strategy). Of course, the scenario transfer optimization training method can preset multiple other enemy maneuver strategies from simple to complex according to requirements. That is, the scenario transfer optimization training method reasonably decomposes a complex scenario (that is, a scenario that always uses the MINMAX strategy) into multiple scenarios from simple to complex. After training with the scenarios from simple to complex, the complex scenario using the MINMAX strategy can obtain a faster convergence speed and more efficiently obtain a better maneuver strategy.

[0047] (2) Train our drone's maneuver decisions to counter the enemy maneuver strategies from simple to complex through a deep reinforcement learning model.

[0048] Specifically, our drone is trained using a deep reinforcement learning model. Based on the enemy maneuver strategies of the enemy drone from simple to complex, our drone's maneuver decisions to counter the enemy maneuver strategies from simple to complex are trained through the deep reinforcement learning model.

[0049] For better illustration, a simulation comparison analysis is carried out in this embodiment. The total number of simulation training rounds is 30,000 times. The first method is the direct training method. In the 30,000 training rounds, the maneuver strategy of the enemy UAV always selects the minimax strategy; the second method is the scenario transfer optimization training method. The 30,000 training rounds are divided into three groups of training rounds. The enemy UAV adopts the first enemy maneuver strategy in the first 10,000 training rounds, the second enemy maneuver strategy in the 10,000 - 20,000 training rounds, and the third enemy maneuver strategy in the 20,000 - 30,000 training rounds. The simulation results are as Figure 2 shown. The average reward return of the scenario transfer optimization training method for each group of training rounds increases rapidly. After 20,000 training rounds, the minimax strategy used by the scenario transfer optimization training method begins to converge rapidly, and the training speed is better than that of the direct training. At the same time, it can be very intuitively observed that after the scenario transfer, our UAV does not start learning from scratch when facing the new scenario, and the previous experience is inherited. The starting reward value of the training is relatively high, and the training effect improves rapidly after a period of time.

[0050] This embodiment also compares the training times of the two training methods. As Figure 3 shown, on the same device environment, the training time of the direct training method is longer than that of the scenario transfer optimization training method. The scenario transfer optimization training method shortens the training time by more than 50% compared with the direct training method. Analyzing the reasons, it is mainly because the calculation amount of the minimax strategy is relatively large, and always using the minimax strategy makes the calculation time long.

[0051] In one embodiment, since the scenario transfer optimization training method is a mutation - type scenario transformation, and the mutation - type scenario transformation is not conducive to stable training. To further optimize the scenario transfer optimization training method, this embodiment designs a soft - update scenario transfer optimization training method. This method trains our UAV's maneuver decision - making against at least one enemy maneuver strategy within each group of training rounds; among them, the enemy maneuver strategies against which in each group of training rounds are selected from all enemy maneuver strategies according to a preset selection rule.

[0052] Specifically, this embodiment uses the soft - update scenario transfer optimization training method for training. Based on the soft - update scenario transfer optimization training method, when the enemy UAV makes a maneuver decision each time, different enemy maneuver strategies are selected according to a preset selection rule. The preset selection rule is:

[0053] Assume that the number of training rounds is 30,000 rounds. The 30,000 training rounds are divided into three groups of training rounds. Then, when the enemy UAV is in the first 10,000 training rounds, the probability of selecting the first enemy maneuver strategy is P a , and the probability of selecting the second enemy maneuver strategy is 1 - Pa ; During 10,000 to 20,000 training rounds, the probability of selecting the second enemy maneuver strategy is P b , and the probability of selecting the third enemy maneuver strategy is 1 - P b ; During 20,000 to 30,000 training rounds, the third enemy maneuver strategy is selected.

[0054]

[0055]

[0056] Among them, s represents the current number of training rounds. As the number of rounds increases, the strategy selected by the blue - side UAV gradually becomes more complex. This gradual process from easy to difficult is more suitable for the deep reinforcement learning model of our UAVs to train.

[0057] For better illustration, this embodiment conducts a simulation comparison analysis. As Figure 4 shown, in terms of both training speed and stability effect, the training effect of the soft - update scenario transfer optimization method is superior to that of the scenario transfer optimization training method and the direct training method. The soft - update scenario transfer optimization method has a faster convergence speed. It begins to converge at about 15,000 rounds and can obtain a higher average reward return value faster. The results prove that the soft - update scenario transfer optimization method is effective.

[0058] In one embodiment, since the maneuver strategy learned by our UAVs is not suitable for actual air combat situations. As Figure 5 shown, r represents our UAV, and b represents the enemy UAV. At the beginning of the battle, our UAV is in an unfavorable state of being chased by the enemy UAV. The maneuver strategy of our UAV is to dive and then pull up while decelerating, quickly shortening the distance between the two sides, forcing the enemy UAV to overtake our UAV. Although the total reward return of our UAV is relatively large during the entire battle process, after choosing this maneuver strategy for a period of time, it is chased and locked by the enemy UAV and is at a disadvantage. In actual air combat, it is very likely to be shot down during this period. Therefore, the maneuver strategy learned by our UAVs in the above - mentioned embodiment does not meet the actual combat requirements.

[0059] To enable our UAVs to learn maneuver strategies that conform to actual air combat situations, this embodiment designs a rule - coupling optimization training method. The rule - coupling optimization training method couples the deep reinforcement learning model of our UAVs with preset rules.

[0060] Specifically, this embodiment refers to the experience of fighter pilots and the research results of air combat experts, summarizes some reasonable and effective combat skills, and forms preset rules, couples the deep reinforcement learning model with the preset rules, and couples the maneuvering strategy trained by the deep reinforcement learning model with the maneuvering strategy selected by the preset rules, and autonomously makes maneuvering decisions. The rule coupling optimization training method can use the deep reinforcement learning model to train the maneuvering strategy of our UAV, and can also use the preset rules to select our maneuvering strategy, that is, when encountering some specific environments, the preset rules can be triggered to select our maneuvering strategy to make the combat strategy more flexible and efficient.

[0061] For example, when our drone is locked by the enemy drone from the rear hemisphere, we can use some classic maneuvers to get rid of it, such as the snake maneuver strategy, tail-drop maneuver strategy, and roller maneuver strategy. Take the roller maneuver strategy as an example. Figure 6 As shown in the figure, the principle of the rolling maneuver strategy is that the fighter suddenly changes direction when being tracked, and keeps moving in a circular motion in the longitudinal plane, and moves at a constant speed in the horizontal direction, so that the opponent's tracking diverges. Therefore, this embodiment presets the rule: when the enemy drone is less than 300 meters behind our drone and the flight speed direction points to our drone, our drone selects the rolling maneuver strategy until the distance between our drone and the enemy drone is greater than 500 meters, and then switches to the deep reinforcement learning model. The simulation trajectory is shown in Figure 7 As shown in the figure, r represents our drone and b represents the enemy drone. From the simulation trajectory and data, it can be seen that our drone has increased the distance from the initial 300 meters to 500 meters through two roller maneuvers. During this period, the enemy drone has never been able to lock the red drone.

[0062] Reference Figure 8 The embodiment of the present invention further provides a UAV air combat decision optimization training system, comprising:

[0063] A maneuver strategy preset unit, used to preset a variety of enemy maneuver strategies of enemy UAVs; wherein the various enemy maneuver strategies are related but have different complexities;

[0064] The maneuver strategy learning unit is used to train our UAV's maneuver strategies against enemy maneuvers ranging from simple to complex through a deep reinforcement learning model.

[0065] It should be noted that since the UAV air combat decision optimization training system in this embodiment and the above-mentioned UAV air combat decision optimization training method are based on the same inventive concept, the corresponding contents in the method embodiment are also applicable to the system embodiment and will not be described in detail here.

[0066] An embodiment of the present invention further provides an unmanned aerial vehicle air combat decision optimization training device, including: at least one control processor and a memory communicatively connected to the at least one control processor.

[0067] The memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0068] The non-transitory software programs and instructions required to implement an unmanned aerial vehicle air combat decision optimization training method of the above embodiment are stored in the memory. When executed by the processor, the method for an unmanned aerial vehicle air combat decision optimization training in the above embodiment is executed. For example, the method steps S100 to step S200 described above are executed. Figure 1 in the method steps S100 to step S200.

[0069] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0070] An embodiment of the present invention further provides a computer-readable storage medium. The computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are executed by one or more control processors, the one or more control processors can be caused to execute an unmanned aerial vehicle air combat decision optimization training method in the above method embodiment. For example, the functions of the method steps S100 to step S200 described above are executed. Figure 1 in the method steps S100 to step S200.

[0071] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform. Those skilled in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0072] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "schematic embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0073] The above has specifically described the preferred embodiments of the present invention, but the present invention is not limited to the above embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention. These equivalent deformations or substitutions are all included within the scope defined by the claims of the present invention.

Claims

1. An optimization training method for UAV air combat decision-making, characterized in that, it includes the steps: Preset a variety of enemy maneuver strategies for enemy UAVs; among them, various of the enemy maneuver strategies are related but have different complexities; among them, the variety of enemy maneuver strategies includes at least one of the following: The enemy UAV maintains its heading and moves in a uniform straight line; The enemy UAV randomly selects an action from the maneuver action library; The enemy UAV makes a maneuver decision according to the MINMAX strategy; Train the maneuver strategy of our UAV to counter the enemy maneuver strategies from simple to complex through a deep reinforcement learning model, specifically: Preset multiple groups of training rounds; Train the maneuver strategy of our UAV to counter at least one of the enemy maneuver strategies within each group of training rounds; among them, at least one of the enemy maneuver strategies countered within each group of training rounds is selected from all the enemy maneuver strategies according to a preset selection rule.

2. An optimization training method for UAV air combat decision-making according to claim 1, characterized in that, it includes: Couple the maneuver strategy of our UAV trained by the deep reinforcement learning model with a preset rule to obtain the maneuver strategy of our UAV.

3. An optimization training method for UAV air combat decision-making according to claim 2, characterized in that, The coupling of the maneuver strategy of our UAV trained by the deep reinforcement learning model with a preset rule to obtain the maneuver strategy of our UAV includes: Based on the experience of fighter pilots and the research results of air combat experts, summarize combat skills and form a preset rule; Couple the maneuver strategy of our UAV trained by the deep reinforcement learning model with a preset rule to obtain the maneuver strategy of our UAV.

4. An optimization training system for UAV air combat decision-making, characterized in that, it includes: A maneuver strategy preset unit for presetting a variety of enemy maneuver strategies for enemy UAVs; among them, various of the enemy maneuver strategies are related but have different complexities; among them, the variety of enemy maneuver strategies includes at least one of the following: The enemy UAV maintains its heading and moves in a uniform straight line; The enemy UAV randomly selects an action from the maneuver action library; The enemy UAV makes a maneuver decision according to the MINMAX strategy; A maneuver strategy learning unit for training the maneuver strategy of our UAV to counter the enemy maneuver strategies from simple to complex through a deep reinforcement learning model, specifically: Preset multiple groups of training rounds; Train the maneuver strategy of our UAV to counter at least one of the enemy maneuver strategies within each group of training rounds; among them, at least one of the enemy maneuver strategies countered within each group of training rounds is selected from all the enemy maneuver strategies according to a preset selection rule.

5. An optimization training device for UAV air combat decision-making, characterized in that: It includes at least one control processor and a memory for communicating with the at least one control processor; The memory stores instructions executable by the at least one control processor, and the instructions are executed by the at least one control processor to enable the at least one control processor to execute a method for optimizing the training of an unmanned aerial vehicle air combat decision-making as described in any one of claims 1 to 3.

6. A computer-readable storage medium, characterized in that: the computer-readable storage medium stores computer-executable instructions for causing a computer to execute a method for optimizing the training of an unmanned aerial vehicle air combat decision-making as described in any one of claims 1 to 3.

Citation Information

Patent Citations

  • Robot navigation method and system based on multimode perception and reinforcement learning

    CN111367282A