Air battle intelligent decision-making method, device and equipment based on course reinforcement learning and medium

Through the method based on course reinforcement learning, the training samples and mission goals are dynamically adjusted, and the problems of slow convergence and high resource requirements in the training process of UAVs are solved, and rapid and effective tactical decision-making capabilities are improved.

CN120295129AActive Publication Date: 2025-07-11SICHUAN UNIV

Patent Information

Application Number
CN202510444767.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-11
Estimated Expiration
2045-04-10

AI Technical Summary

Technical Problem

In the prior art, the unmanned fighter jets have problems such as slow training convergence speed, high computing resource requirements, and poor adaptability of complex tactical environments when conducting combat based on reinforcement learning, resulting in the inability to effectively accelerate the training process, reduce computing costs, and improve decision-making and adaptability.

Method used

A curriculum-based reinforcement learning method is adopted to determine combat advantage indicators for the initial state data, generate a difficulty-guided state pool, and dynamically adjust the training samples using the progressive sliding buffer mechanism, gradually introduce more complex combat scenarios, dynamically adjust the difficulty of the task target, and gradually guide the strategy network to approach the final task target.

Benefits of technology

It optimizes learning efficiency, reduces training costs, significantly accelerates the training process, reduces the demand for computing resources, and improves the decision-making and adaptability of unmanned fighters in complex air combat environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120295129A_ABST
    Figure CN120295129A_ABST
Patent Text Reader

Abstract

The invention provides an air battle intelligent decision-making method and device based on course reinforcement learning, equipment and a medium, and relates to the field of reinforcement learning. Comprising the steps of determining corresponding combat advantage indexes based on multiple groups of initial state data; sorting the initial state data based on a combat advantage index, and storing the sorted initial state data into a difficulty-guided state pool; updating the initial state data in the progressive sliding buffer area from the difficulty-guided state pool according to a difficulty increasing mode; dynamically formulating an intermediate task target until a final task target is formulated; sampling current initial state data from the progressive sliding buffer area, and inputting the current initial state data into a strategy network of the SAC model to obtain a current action; and the SAC model is trained at least based on the intermediate task target, the current action and the current initial state data until the final task target is reached, and the trained SAC model is obtained for making an intelligent decision, so that the decision-making ability and the adaptive ability of the aircraft in a complex environment are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of reinforcement learning, and particularly to an air combat intelligent decision-making method, device, equipment and medium based on curriculum reinforcement learning. Background Art

[0002] In recent years, with the continuous development of UCAV (Unmanned Combat Aerial Vehicle) technology, how to enable it to make autonomous decisions in complex air combat environments has become the focus of research. Reinforcement Learning (RL) has shown great potential in realizing the autonomous control of UCAV. Especially for high-risk tasks such as air combat that require rapid adaptation and complex tactical maneuvers, although RL-based methods have achieved remarkable success in fields such as path planning, obstacle avoidance, and autonomous flight control, their application in UCAV combat missions faces major challenges due to the large amount of computing resources and time required for training. The RL training process in UCAV is particularly slow and computationally intensive, mainly because of the inherent complexity of the air combat environment. These scenarios require exploring high-dimensional action spaces, adapting to adversarial dynamics in real time, and precisely coordinating tactical maneuvers. The computational cost of achieving operational proficiency under these constraints poses a major obstacle to scalability and deployment readiness. In addition, the dynamic and adversarial nature of air combat exacerbates these challenges, requiring the system to quickly adapt to changing threats and make real-time decisions. Achieving this ability usually requires extending the training time, which highlights the urgent need for more efficient RL training paradigms that need to be tailored to the unique requirements of UCAV combat scenarios.

[0003] It can be seen that the related technologies for training UCAV to conduct combat based on reinforcement learning currently have technical defects such as slow training convergence speed, high computing resource requirements, and poor adaptability to complex tactical environments, resulting in the current inability to improve the decision-making ability and adaptability of UCAV in complex air combat environments on the premise of effectively accelerating the training process and reducing the computing cost. Summary of the Invention

[0004] Based on the above technical problems, the present invention provides an air combat intelligent decision-making method, device, equipment and medium based on curriculum reinforcement learning, aiming to overcome the above problems or at least partially solve the above problems.

[0005] The first aspect of the present invention provides an air combat intelligent decision-making method based on curriculum reinforcement learning, and the method includes:

[0006] Based on multiple groups of initial state data, determine the combat advantage index corresponding to each group of initial state data, and the combat advantage index characterizes the difficulty of the combat scenario for the first own aircraft;

[0007] Sort the multiple groups of initial state data based on the combat advantage metrics, and store the sorted multiple groups of initial state data in a difficulty-guided state pool;

[0008] Update the initial state data in the progressive sliding buffer in ascending order of difficulty from the difficulty-guided state pool;

[0009] Dynamically formulate intermediate task objectives for the policy network in the deep reinforcement learning SAC model to be trained. The intermediate task objectives guide the policy network to transition from simple tasks to complex tasks until the final task objective is formulated for the policy network;

[0010] Sample the current initial state data from the progressive sliding buffer, and input the current initial state data into the policy network of the deep reinforcement learning SAC model to be trained to obtain the current action of the first friendly aircraft;

[0011] Train the deep reinforcement learning SAC model to be trained based on at least the intermediate task objectives, the current action of the first friendly aircraft, and the current initial state data until the first friendly aircraft reaches the final task objective, and obtain the trained deep reinforcement learning SAC model;

[0012] Input the current air combat environment information into the policy network of the trained deep reinforcement learning SAC model, and output the action information of the second friendly aircraft.

[0013] The second aspect of the present invention provides an air combat intelligent decision-making device based on curriculum reinforcement learning. The device includes:

[0014] An index determination module for determining the combat advantage index corresponding to each group of initial state data based on multiple groups of initial state data. The combat advantage index characterizes the difficulty of the combat scenario for the first friendly aircraft;

[0015] A state pool storage module for sorting the multiple groups of initial state data based on the combat advantage index and storing the sorted multiple groups of initial state data in a difficulty-guided state pool;

[0016] A buffer update module for updating the initial state data in the progressive sliding buffer in ascending order of difficulty from the difficulty-guided state pool;

[0017] An objective formulation module for dynamically formulating intermediate task objectives for the policy network in the deep reinforcement learning SAC model to be trained. The intermediate task objectives guide the policy network to transition from simple tasks to complex tasks until the final task objective is formulated for the policy network;

[0018] An action output module, configured to sample current initial state data from the progressive sliding buffer, input the current initial state data into a policy network in the to-be-trained deep reinforcement learning SAC model, and obtain a current action of the first own aircraft;

[0019] A model training module, configured to train the to-be-trained deep reinforcement learning SAC model at least based on the intermediate task objective, the current action of the first own aircraft, and the current initial state data until the first own aircraft reaches the final task objective, and obtain a trained deep reinforcement learning SAC model;

[0020] A decision output module, configured to input current air combat environment information into a policy network in the trained deep reinforcement learning SAC model and output action information of the second own aircraft.

[0021] A third aspect of the present invention provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the computer program is executed by the processor, it implements the air combat intelligent decision-making method based on curriculum reinforcement learning in the first aspect of the embodiments of the present invention.

[0022] A fourth aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored, where when the computer program is executed by a processor, it implements the air combat intelligent decision-making method based on curriculum reinforcement learning in the first aspect of the embodiments of the present invention.

[0023] Through the air combat intelligent decision-making method based on curriculum reinforcement learning provided by the present invention, automatic curriculum learning is integrated into the air combat training process of the aircraft: determining the corresponding combat index advantages for the initial state data, sorting the initial state data based on the combat index advantages to generate a difficulty-guided state pool, and then updating the initial state data in the progressive sliding buffer in the order of increasing difficulty from the difficulty-guided state pool, so as to dynamically adjust the progressive sliding buffer as the state pool of training samples, gradually introducing more complex combat scenarios during the training process to ensure the diversity of training samples, and at the same time avoiding the influence of overly simple samples on the training effect, thereby supporting effective curriculum learning in air combat missions; at the same time, the task objective difficulty of the present invention gradually increases, starting from simple tasks to dynamically adjust the difficulty of the task objective, and using intermediate task objectives to gradually guide the policy network to approach the final task objective. This dynamic adjustment mechanism ensures that the policy network can gradually master complex tactical behaviors during the training process and improve the tactical decision-making ability. In this way, the present invention adopts a progressive difficulty method, starting from simplified tasks, and advancing to complex scenarios through automatic complexity measurement and adaptive goal setting. This adaptive curriculum optimizes the learning efficiency, reduces the training cost, and significantly accelerates the training process while ensuring the decision-making quality of the deep reinforcement learning SAC model, reduces the computational resource requirements, and enhances the decision-making ability and adaptability of the second friendly aircraft in complex air combat environments, enabling the second friendly aircraft to effectively and quickly obtain advanced tactical strategies. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments of the present invention. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0025] Figure 1 is a flowchart of the steps of an air combat intelligent decision-making method based on curriculum reinforcement learning shown in an embodiment of the present invention;

[0026] Figure 2 is a schematic diagram of air combat geometric relationships shown in an embodiment of the present invention;

[0027] Figure 3 is a schematic flowchart of combat advantage index calculation and state pool generation shown in an embodiment of the present invention;

[0028] Figure 4 is a schematic flowchart of progressive task generation shown in an embodiment of the present invention;

[0029] Figure 5It is a schematic diagram showing the update of a progressive sliding buffer shown in an embodiment of the present invention;

[0030] Figure 6 It is a schematic diagram of a weapon engagement area shown in an embodiment of the present invention;

[0031] Figure 7 It is a schematic diagram of an automatic engagement course learning framework shown in an embodiment of the present invention;

[0032] Figure 8 It is a schematic diagram of the overall process of an air combat intelligent decision-making method based on curriculum reinforcement learning shown in an embodiment of the present invention;

[0033] Figure 9 It is a structural block diagram of an air combat intelligent decision-making device based on curriculum reinforcement learning provided in an embodiment of the present invention. Detailed implementation manners

[0034] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0035] At present, there are several key technical problems faced in training unmanned combat aircraft based on Deep Reinforcement Learning (DRL) for combat, specifically including the following three aspects: 1. Slow training convergence speed: Traditional DRL methods usually rely on a large number of samples and repeated computational iterations during the training process. Especially in the highly complex and dynamically changing environment of air combat, the convergence speed of model training is often slow, making it difficult to meet the real-time requirements of practical applications. The long training cycle not only increases the development cost but also affects the timeliness of actual combat deployment. 2. High computational resource requirements: Air combat missions involve an extremely large and high-dimensional state space and action space, and UCAV needs to make immediate and accurate decisions in these complex spaces. However, traditional DRL methods require a large amount of computational resources and data samples during the training stage, resulting in an extremely large computational overhead, which has become a bottleneck for application promotion. Especially in an environment with limited resources or computing power, this problem is particularly prominent, seriously affecting the scalability and practical operation ability of the system. 3. Poor adaptability to complex tactical environments: The air combat environment is highly dynamic and adversarial, and tactical strategies need to be adjusted in real time according to the actions of the opponent and changes in the battlefield environment. However, traditional reinforcement learning methods, especially the curriculum learning model based on fixed task difficulty, lack flexible dynamic adjustment capabilities and cannot adaptively adjust the training difficulty or optimize the strategy according to the training progress and environmental changes. This leads to the difficulty of continuously improving the intelligent decision-making model in complex and rapidly changing air combat scenarios, affecting its effectiveness and robustness in actual combat.

[0036] Due to the dynamic, adversarial, and high-dimensional characteristics of air combat, the challenges of reinforcement learning training are exacerbated. Different from standard drone tasks such as obstacle avoidance or autonomous landing, the UCAV combat scenario requires continuous adaptation to intelligent opponents, where the environmental changes are unpredictable and the reward signals are sparse, requiring extensive exploration and extended convergence time, making the training process both time-consuming and resource-consuming. In addition, the risk of catastrophic forgetting during incremental learning and the computational burden of high-fidelity simulation further hinder the development of an efficient training process. Currently, there is an urgent need for a new method that can effectively accelerate the training process, reduce the computational cost, and at the same time enhance the decision-making ability and adaptability of UCAV in complex air combat environments, so as to better meet the actual needs of unmanned combat aircraft combat missions.

[0037] Therefore, in order to at least partially solve one or more of the above problems and other potential problems, embodiments of the present invention propose an air combat intelligent decision-making method based on curriculum reinforcement learning. In this method, by determining the corresponding combat advantage index for each group of initial state data and sorting the initial state data based on the combat index advantage to generate a difficulty-guided state pool, automated combat complexity assessment is achieved; by adopting a progressive sliding buffer mechanism, the initial state data in the progressive sliding buffer is updated in an increasing difficulty manner from the difficulty-guided state pool to dynamically adjust the state pool of the progressive sliding buffer as training samples, and more complex combat scenarios are gradually introduced during the training process to ensure the diversity of training samples; by using intermediate task goals to gradually guide the policy network to approach the final task goal, goal-driven progressive task generation is realized. This dynamic adjustment mechanism ensures that the policy network can gradually master complex tactical behaviors during the training process and improves the tactical decision-making ability. Thus, the present invention adopts a progressive difficulty method, starting from a simplified task, and advancing to complex scenarios through automatic complexity measurement and adaptive goal setting. This adaptive curriculum optimizes the learning efficiency, reduces the training cost, accelerates the training process, reduces the computational resource requirements, and enhances the decision-making ability and adaptability of the second friendly aircraft in a complex air combat environment, enabling the second friendly aircraft to effectively and quickly obtain advanced tactical strategies.

[0038] Please refer to Figure 1 , Figure 1 is a flowchart of the steps of an air combat intelligent decision-making method based on curriculum reinforcement learning shown in an embodiment of the present invention. As Figure 1 shown, the air combat intelligent decision-making method based on curriculum reinforcement learning provided in this embodiment at least includes the following steps:

[0039] Step S11: Based on multiple groups of initial state data, determine the corresponding combat advantage index for each group of initial state data, where the combat advantage index characterizes the difficulty of the combat scenario for the first friendly aircraft.

[0040] In this embodiment, multiple groups of initial state data can be obtained, and each group of initial state data characterizes the multi-dimensional initial situation information of both sides in the air combat training process. After obtaining multiple groups of initial state data in this embodiment, based on the multiple groups of initial state data, a combat advantage index corresponding to each group of initial state data can be determined. Among them, the first own aircraft in this embodiment is the own aircraft (the own side among the two sides in the combat) in the model training process, and the combat advantage index characterizes the difficulty of the combat scenario for the first own aircraft (abbreviation: combat scenario difficulty). This embodiment can dynamically evaluate multiple key factors in the combat scenario through each group of initial state data, quantify the difficulty of the combat scenario, and obtain the combat advantage index corresponding to each group of initial state data. In an alternative embodiment, the multiple key factors in the combat scenario at least include: the engagement angle between the two sides, energy advantage (ability consumption), and positional relationship.

[0041] In an alternative embodiment, the initial state data of this embodiment can be processed from scenario data. Among them, for the scenario data, this embodiment can randomly initialize the air combat training scenario and randomly generate multiple groups of scenario data. Among them, a scenario is an assumption about the basic situation, combat intention, and combat development of both sides in the combat; for example, each group of scenario data can at least include one or more of the following: the initial longitude and latitude, initial speed, initial altitude, and initial yaw angle of the first own aircraft, and the relative distance, relative speed, relative altitude, relative heading angle, and relative azimuth angle of the opponent aircraft relative to the first own aircraft.

[0042] Step S12: Sort the multiple groups of initial state data based on the combat advantage index, and store the sorted multiple groups of initial state data in the difficulty-guided state pool.

[0043] In this embodiment, based on the combat advantage index corresponding to each group of initial state data, the multiple groups of initial state data can be sorted, and the sorted multiple groups of initial state data can be stored in the difficulty-guided state pool to ensure that the deep reinforcement learning SAC model to be trained gradually adapts to increasingly complex combat situations during training, and finally generates a difficulty-guided state pool. The difficulty-guided state pool of this embodiment stores multiple groups of initial state data sorted according to the combat advantage index, and each group of stored initial state data carries the corresponding combat advantage index.

[0044] Step S13: Update the initial state data in the progressive sliding buffer in an increasing order of difficulty from the difficulty-guided state pool.

[0045] In this embodiment, a progressive sliding buffer is provided, and the progressive sliding buffer is used as the state pool of the training samples for air combat training. Through the progressive sliding buffer mechanism, in this embodiment, the initial state data in the progressive sliding buffer is progressively and slidably updated in the order of increasing combat scenario difficulty from the difficulty-guided state pool, so as to dynamically adjust the state pool of the training samples. It can be understood that the initial state of the progressive sliding buffer is multiple groups of initial state data with the lowest combat scenario difficulty in the difficulty-guided state pool.

[0046] Step S14: Dynamically formulate an intermediate task objective for the policy network in the deep reinforcement learning SAC model to be trained, where the intermediate task objective guides the policy network to transition from simple tasks to complex tasks until a final task objective is formulated for the policy network.

[0047] In this embodiment, the difficulty of the task objective (the difficulty of the combat scenario for the first friendly aircraft) formulated for the policy network in the deep reinforcement learning SAC model to be trained increases gradually. An intermediate task objective is dynamically formulated for the policy network in the deep reinforcement learning SAC model to be trained. Specifically, the difficulty of the intermediate task objective can be dynamically adjusted according to the training progress of the deep reinforcement learning SAC model to be trained. The intermediate task objective guides the policy network to transition from simple tasks to complex tasks until a final task objective is formulated for the policy network in the deep reinforcement learning SAC model to be trained.

[0048] It should be noted that this embodiment does not specifically limit the execution order between step S14 and step S13. Step S13 can be executed first and then step S14, or step S14 can be executed first and then step S13, or step S13 and step S14 can be executed simultaneously.

[0049] Step S15: Sample the current initial state data from the progressive sliding buffer, and input the current initial state data into the policy network in the deep reinforcement learning SAC model to be trained to obtain the current action of the first friendly aircraft.

[0050] In this embodiment, the deep reinforcement learning SAC algorithm is used for training, and the deep reinforcement learning SAC model to be trained is initialized. The deep reinforcement learning SAC model to be trained is a component of the intelligent agent Agent in deep reinforcement learning. The deep reinforcement learning SAC model to be trained includes at least a policy network, and the policy network is used to output the action decision of the first friendly aircraft according to the current state information (such as position, speed, and altitude, etc.) of the first friendly aircraft and the opponent aircraft.

[0051] In this embodiment, a set of initial state data can be sampled from the progressive sliding buffer as the current initial state data, and the current initial state data is input into the policy network in the deep reinforcement learning SAC model to be trained, so as to obtain the current action of the first friendly aircraft output by the policy network for the intermediate task target. That is to say, the current initial state data in this embodiment is the current set of initial state data. In an optional example, the current action of the first friendly aircraft output by the policy network may include action values such as the elevator, rudder, aileron, and throttle of the first friendly aircraft. In an example, the aircraft (such as the first friendly aircraft, the second friendly aircraft) may include: unmanned fighter aircraft.

[0052] Step S16: Train the deep reinforcement learning SAC model to be trained at least based on the intermediate task target, the current action of the first friendly aircraft, and the current initial state data until the first friendly aircraft reaches the final task target, so as to obtain the trained deep reinforcement learning SAC model.

[0053] In this embodiment, the deep reinforcement learning SAC model to be trained can be trained at least based on the intermediate task target formulated for the policy network in the deep reinforcement learning SAC model to be trained, the current action of the first friendly aircraft output by the policy network in the deep reinforcement learning SAC model to be trained, and the current initial state data input into the policy network in the deep reinforcement learning SAC model to be trained until it is determined that the first friendly aircraft reaches the final task target, so as to obtain the trained deep reinforcement learning SAC model.

[0054] In the training process of this embodiment, a reward mechanism will also be designed: a positive reward is given if the first friendly aircraft successfully pursues the opponent aircraft, otherwise a negative reward is given; through multiple training iterations, the pursuit strategy of the first friendly aircraft is gradually optimized until the first friendly aircraft reaches the final task target, so as to obtain the trained deep reinforcement learning SAC model.

[0055] Step S17: Input the current air combat environment information into the policy network in the trained deep reinforcement learning SAC model, and output the action information of the second friendly aircraft.

[0056] After obtaining the trained deep reinforcement learning SAC model in this embodiment, the policy network in the trained deep reinforcement learning SAC model can be used to make intelligent decisions for the friendly aircraft in actual air combat and output the actions of the friendly aircraft. Specifically, the current air combat environment information can be input into the policy network in the trained deep reinforcement learning SAC model to output the action information of the second friendly aircraft. Herein, the second friendly aircraft in this embodiment is the friendly aircraft (the friendly side in the two combat sides) during model application, and the current air combat environment information is a set of air combat environment information, including multiple initial situation information of the two combat sides during the actual air combat process.

[0057] Aiming at the technical defect that traditional curriculum learning relies on static curriculum settings and lacks the ability to dynamically adjust the difficulty according to the model progress and actual situation. Especially in an air combat environment with strong confrontation and fast dynamic changes, fixed curriculum design cannot effectively cope with complex combat scenarios. In this embodiment, automatic curriculum learning is integrated into the air combat training process of the aircraft: determining the corresponding combat index advantages for the initial state data, sorting the initial state data based on the combat index advantages to generate a difficulty-guided state pool, and then updating the initial state data in the progressive sliding buffer in the order of increasing difficulty from the difficulty-guided state pool, so as to dynamically adjust the progressive sliding buffer as the state pool of training samples, gradually introducing more complex combat scenarios during the training process to ensure the diversity of training samples, while avoiding the influence of overly simple samples on the training effect, thereby supporting effective curriculum learning in air combat tasks; at the same time, the task objective difficulty of the present invention gradually increases. Starting from simple tasks, the difficulty of the task objective is dynamically adjusted, and intermediate task objectives are used to gradually guide the policy network to approach the final task objective. This dynamic adjustment mechanism ensures that the policy network can gradually master complex tactical behaviors during the training process and improve the tactical decision-making ability. In this way, the present invention adopts a progressive difficulty method, starting from simplified tasks and advancing to complex scenarios through automatic complexity measurement and adaptive target setting. This adaptive curriculum optimizes the learning efficiency, reduces the training cost, and significantly accelerates the training process while ensuring the decision-making quality of the deep reinforcement learning SAC model, reduces the computing resource requirements, and improves the decision-making ability and adaptability of the second friendly aircraft in complex air combat environments, enabling the second friendly aircraft to effectively and quickly obtain advanced tactical strategies.

[0058] Combined with the above embodiments, in one implementation manner, in order to make the dynamic curriculum learning better utilize the training samples, reduce ineffective exploration, and improve the learning efficiency and stability, this embodiment can design different curriculum scenarios according to the advantages and disadvantages of the initial spatial positions of the first friendly aircraft and the opposing aircraft. This design can simulate various situations that may be encountered in actual air combat, enabling the intelligent agent to train in a diverse environment, thereby possessing the ability to cope with various initial situations (initial state data).

[0059] Specifically, when randomly generating multiple sets of scenario data, the range of random variables in the scenario can be designed. For example, from the initial yaw angle of the first friendly aircraft the initial altitude h0, the initial velocity v0, the relative distance d, the relative course angle θ of the opposing aircraft relative to the first friendly aircraft heading and the relative azimuth angle θ azimuth jointly determine. In one example, by changing the values of the relative course angle and the relative azimuth angle, the mission courses are designed according to the principle from simple to difficult, and a total of four sets of difficulty courses are generated, respectively meeting the following conditions:

[0060] ① θ azimuth ∈[-Π / 2, Π / 2], θ heading = 0, d ∈ [1000, 5000] m, h0 ∈ [100, 12000] m, v0 ∈ [120, 340] m / s

[0061] ② θ azimuth ∈[-Π, Π], θ heading = 0, d ∈ [1000, 5000] m, h0 ∈ [100, 12000] m, v0 ∈ [120, 340] m / s

[0062] ③ θ azimuth ∈[-Π, Π], θ heading ∈[-Π, 0], d ∈ [1000, 5000] m, h0 ∈ [100, 12000] m, V0 ∈ [120, 340] m / s

[0063] ④ θ azimuth ∈[-Π, Π], θ heading ∈[-Π, Π], d ∈ [1000, 5000] m, h0 ∈ [100, 12000] m, v0 ∈ [120, 340] m / s;

[0064] In this embodiment, based on the set multiple sets of difficulty courses (such as the above four sets of difficulty courses), multiple sets of scenario data belonging to the multiple sets of difficulty courses can be randomly generated, so as to obtain multiple sets of initial state data based on the multiple sets of scenario data. In this way, when determining the combat advantage indicators corresponding to the initial state data and subsequent sorting based on the combat advantage indicators, the initial state data, calculating the combat advantage indicators and performing subsequent sorting are obtained in sequence according to the difficulty levels of the multiple sets of difficulty courses, so that the sorting algorithm can be simplified and computing resources can be saved.

[0065] Combined with the above embodiments, in one implementation, the present invention further provides an air combat intelligent decision-making method based on curriculum reinforcement learning. In this method, the step of "determining the combat advantage index corresponding to each set of initial state data based on multiple sets of initial state data" in the above step S11 may specifically include steps S21 to S24:

[0066] Step S21: Determine the angle advantage value based on the engagement angle in each set of initial state data.

[0067] In this embodiment, each set of initial state data at least includes: the engagement angle between the first own aircraft and the opponent aircraft, the relative distance, the relative flight altitude, and the flight speed of the first own aircraft; wherein, the opponent aircraft is the opponent aircraft in the air combat (the opponent in both combat sides). In this embodiment, the calculation of the combat advantage index mainly considers three aspects: angle advantage, distance advantage, and energy advantage. Specifically, for the angle advantage, the angle advantage value corresponding to each set of initial state data can be determined based on the engagement angle in each set of initial state data.

[0068] Step S22: Determine the distance advantage value based on the relative distance in each set of initial state data.

[0069] In this embodiment, for the distance advantage, the distance advantage value corresponding to each set of initial state data can be determined based on the relative distance in each set of initial state data.

[0070] Step S23: Determine the energy advantage value based on the relative flight altitude and flight speed in each set of initial state data.

[0071] In this embodiment, for the energy advantage, the energy advantage value corresponding to each set of initial state data can be determined based on the relative flight altitude and the flight speed of the first own aircraft in each set of initial state data.

[0072] Step S24: Determine the combat advantage index corresponding to each set of initial state data based on at least the angle advantage value, the distance advantage value, and the energy advantage value, where the larger the value of the combat advantage index, the lower the difficulty of the combat scenario for the first own aircraft.

[0073] In this embodiment, after obtaining the angle advantage value, the distance advantage value, and the energy advantage value corresponding to each set of initial state data respectively, the combat advantage index corresponding to each set of initial state data can be determined based on at least the angle advantage value, the distance advantage value, and the energy advantage value corresponding to each set of initial state data respectively. In this embodiment, the larger the value of the combat advantage index, the higher the advantage of the first own aircraft, the lower the difficulty of the combat scenario for the first own aircraft, and it also represents that the difficulty of the combat scenario of the mission is low.

[0074] Combined with the above embodiments, in one implementation, the present invention further provides an air combat intelligent decision-making method based on curriculum reinforcement learning. In this method, the above step S21 may specifically include steps S31 to S34, the above step S22 may specifically include steps S35 to S37, and the above step S23 may specifically include steps S38 to S310:

[0075] Step S31: Determine the detection advantage value of the attack angle based on at least the maximum detection angle of the radar and the attack angle.

[0076] In this embodiment, the angles consider the attack angle δ ATA and the escape angle σ AA , which are the angles between the velocity directions v Figure 2 of the first own aircraft (such as the red aircraft in r ) and the velocity direction v Figure 2 of the opposing aircraft (such as the blue aircraft in b ) and the line of sight (LOS line) between the two aircraft, as shown in Figure 2 . Figure 2 It is a schematic diagram of the air combat geometric relationship shown in an embodiment of the present invention. That is to say, the engagement angles in this embodiment include: the attack angle and the escape angle.

[0077] In this embodiment, based on the angle theory and energy theory in air combat, the situation advantage in the joint state can be captured and evaluated from the detection advantage and threat advantage, and finally the combat advantage index can be obtained. Specifically, the detection advantage value of the attack angle can be determined based on at least the maximum detection angle of the radar and the attack angle in the initial state data.

[0078] In an optional example, the detection advantage function of the attack angle can be determined, and the detection advantage value of the attack angle can be determined based on the detection advantage function of the attack angle and the attack angle in each set of initial state data. For example, the detection advantage function of the attack angle is shown in the following formula (1):

[0079]

[0080] where δ Rmax represents the maximum detection angle of the radar, φ0 is an adjustable parameter, and is the attack angle in each set of initial state data.

[0081] Step S32: Determine the threat advantage value of the attack angle based on at least the angle corresponding to the maximum engagement area, the missile lock angle, the non-escape attack area angle, and the attack angle.

[0082] In this embodiment, the threat advantage value of the attack angle can be determined based on at least the angle corresponding to the maximum engagement area, the missile lock-on angle, the inescapable attack area angle, and the attack angle in each set of initial state data.

[0083] In an alternative example, a threat advantage function of the attack angle can be determined, and based on the threat advantage function of the attack angle and the attack angle in each set of initial state data, the threat advantage value of the attack angle can be determined. For example, the threat advantage function of the attack angle is shown in the following formula (2):

[0084]

[0085] where δ R represents the angle corresponding to the maximum engagement area, δ MK represents the missile lock-on angle, δ MO represents the inescapable attack area angle, φ1 and φ2 are adjustable parameters, and φ is the attack angle in each set of initial state data.

[0086] Step S33: Determine the threat advantage value of the escape angle based on at least the escape angle.

[0087] In this embodiment, the threat advantage value of the escape angle can be determined based on at least the escape angle in each set of initial state data.

[0088] In an alternative example, a threat advantage function of the escape angle can be determined, and based on the threat advantage function of the escape angle and the escape angle in each set of initial state data, the threat advantage value of the escape angle can be determined. For example, the threat advantage function of the escape angle is shown in the following formula (3):

[0089]

[0090] where φ3 is an adjustable parameter, and φ is the escape angle in each set of initial state data, where

[0091] Step S34: Determine the angle advantage value based on at least the detection advantage value of the attack angle, the threat advantage value of the attack angle, and the threat advantage value of the escape angle.

[0092] In this embodiment, after determining the detection advantage value of the attack angle, the threat advantage value of the attack angle, and the threat advantage value of the escape angle, the angle advantage value can be determined based on at least the detection advantage value of the attack angle, the threat advantage value of the attack angle, and the threat advantage value of the escape angle.

[0093] In an alternative example, the combined threat advantage function of two angles is shown in the following formula (4):

[0094]

[0095] Among them, α ATA and α AA respectively represent the influencing factors of the attack angle and the escape angle, and the constraint condition is: α ATA +α AA = 1.

[0096] In this way, the total angle advantage function is shown in the following formula (5), and the angle advantage value can be obtained based on this total angle advantage function.

[0097]

[0098] Among them, α det and α thr respectively represent the influencing factors of the detection advantage and the threat advantage, and the constraint condition is: α det +α thr = 1.

[0099] Step S35: Determine the detection advantage value of the distance based on at least the maximum detection distance of the radar, the missile lock-on distance, and the relative distance.

[0100] In this embodiment, for the detection advantage and threat advantage of the distance, the influences of the radar and the missile are mainly considered. This embodiment can determine the detection advantage value of the distance based on at least the maximum detection distance of the radar, the missile lock-on distance, and the relative distance in each group of initial state data.

[0101] In an optional example, the detection advantage function of the distance can be determined, and the detection advantage value of the distance can be determined based on the detection advantage function of the distance and the relative distance in each group of initial state data. For example, the detection advantage function of the distance is shown in the following formula (6):

[0102]

[0103] Among them, D Rmax represents the maximum detection distance of the radar, D Lock represents the missile lock-on distance, φ4 is an adjustable parameter, and d is the relative distance in each group of initial state data.

[0104] Step S36: Determine the threat advantage value of the distance based on at least the maximum attack distance of the missile, the non-escape distance, the minimum attack distance, and the relative distance.

[0105] In this embodiment, the threat advantage value of the distance can be determined based on at least the maximum attack distance of the missile, the non-escape distance, the minimum attack distance, and the relative distance in each group of initial state data.

[0106] In an optional example, a threat advantage function based on distance can be determined, and a threat advantage value based on distance can be determined based on the threat advantage function based on distance and the relative distance in each set of initial state data. For example, the threat advantage function based on distance is shown in the following formula (7):

[0107]

[0108] where D Mmax represents the maximum attack distance of the missile, D Mnoesp represents the non-escape distance, D Mmin represents the minimum attack distance, φ5 is an adjustable parameter, and d is the relative distance in each set of initial state data.

[0109] Step S37: Determine the distance advantage value based on at least the detection advantage value based on distance and the threat advantage value based on distance.

[0110] In this embodiment, after determining the detection advantage value based on distance and the threat advantage value based on distance, the distance advantage value can be determined based on at least the detection advantage value based on distance and the threat advantage value based on distance.

[0111] In an optional example, the total distance advantage function is shown in the following formula (8), and the distance advantage value can be obtained based on this total distance advantage function.

[0112]

[0113] where and respectively represent the influencing factors of detection advantage and threat advantage, and the constraint condition is:

[0114] Step S38: Determine the potential energy advantage value based on at least the optimal flight altitude, the maximum flight altitude, and the relative flight altitude.

[0115] In this embodiment, the energy in the air combat process is manifested as potential energy (such as represented by the relative flight altitude) and kinetic energy (such as represented by the relative flight speed). For the potential energy advantage, this embodiment can determine the potential energy advantage value based on at least the optimal flight altitude, the maximum flight altitude, and the relative flight altitude in each set of initial state data.

[0116] In an optional example, a potential energy advantage function can be determined, and the potential energy advantage value can be determined based on the potential energy advantage function and the relative flight altitude in each set of initial state data. For example, the potential energy advantage function is shown in the following formula (9):

[0117]

[0118] where H opt represents the optimal flight altitude, Hmax represents the maximum flight altitude, and h is the relative flight altitude in each set of initial state data.

[0119] Step S39: Determine the kinetic energy advantage value based on at least the kinetic energy of the other aircraft and the kinetic energy of the first own aircraft determined based on the flight speed of the first own aircraft.

[0120] In this embodiment, for the kinetic energy advantage, this embodiment can determine the kinetic energy advantage value based on at least the kinetic energy of the other aircraft and the kinetic energy of the first own aircraft; wherein, the kinetic energy E of the first own aircraft is determined based on the flight speed v of the first own aircraft in each set of initial state data: E = v 2 / 2g. In an alternative example, the kinetic energy of the other aircraft is determined based on the flight speed of the other aircraft in each set of initial state data.

[0121] In an alternative example, a kinetic energy advantage function can be determined, and based on the kinetic energy advantage function, the kinetic energy advantage value can be determined. For example, the kinetic energy advantage function is shown in the following formula (10):

[0122]

[0123] where E T represents the kinetic energy of the other aircraft, E represents the kinetic energy of the first own aircraft, and φ6 is an adjustable parameter.

[0124] Step S310: Determine the energy advantage value based on at least the potential energy advantage value and the kinetic energy advantage value.

[0125] In this embodiment, after determining the potential energy advantage value and the kinetic energy advantage value, the energy advantage value can be determined based on at least the potential energy advantage value and the kinetic energy advantage value.

[0126] In an alternative example, the total energy advantage function is shown in the following formula (11), and the energy advantage value can be obtained based on this total energy advantage function.

[0127]

[0128] where β h and β v respectively represent the influencing factors of potential energy advantage and kinetic energy advantage, and the constraint condition is: β h +β v = 1.

[0129] Combining the above embodiments, in an alternative implementation manner, the above step S24 can be specifically obtained through the following formula (12):

[0130]

[0131] where is the combat advantage index, is the angular advantage value, is the distance advantage value, is the energy advantage value, and λ0, λ1, and λ2 are the weighting coefficients that balance the contributions of each component in the final combat advantage index.

[0132] In one embodiment, as Figure 3 shown, Figure 3 is a schematic flowchart of the calculation of the combat advantage index and the generation of the state pool shown in an embodiment of the present invention. In Figure 3 , the combat advantage index is calculated at least through the speed, altitude, and distance in each set of initial state data, and the difficulty score ψ of each set of initial state data i (i.e., the combat advantage index) can be obtained. Then, according to the combat advantage index ψ i , multiple sets of initial state data are automatically sorted from simple to complex to form a difficulty-guided state pool: This difficulty-guided state pool promotes the training process by gradually introducing the deep reinforcement learning SAC model to be trained into increasingly complex initial state data, thereby optimizing the learning efficiency.

[0133] Combined with the above embodiments, in one implementation manner, the present invention also provides an air combat intelligent decision-making method based on curriculum reinforcement learning. In this embodiment, the above step S14 may specifically include steps S41 to S43:

[0134] Step S41: The target generator generates an intermediate target angle and formulates an intermediate task target for the policy network to reach the intermediate target angle.

[0135] This embodiment maintains a pair of target generators and target discriminators with adaptive angles. In an air combat scenario, a key parameter affecting the task difficulty is the attack angle δ ATA of the aircraft, which has been proven to significantly affect the task complexity. In this embodiment, the task target is to control the opponent's aircraft within the weapon engagement zone. Therefore, in this embodiment, the attack angle is used as the range angle of the weapon engagement zone as the final task target. Considering that directly allowing the intelligent agent to learn to reach the range angle of the attack angle within the weapon engagement zone is usually challenging and time-consuming. Therefore, in this embodiment, an intermediate target angle is introduced as an accessible intermediate task target, so as to more effectively guide the intelligent agent and accelerate the learning process. Therefore, in this embodiment, the target generator generates an intermediate target angle and formulates an intermediate task target for the policy network of the deep reinforcement learning SAC model to be trained to reach the intermediate target angle.

[0136] Step S42: During the training process of the to-be-trained deep reinforcement learning SAC model, when the target discriminator determines that the attack angle of the first friendly aircraft reaches the intermediate target angle, the target generator determines the updated intermediate target angle based on at least the intermediate target angle and the difference between the intermediate target angle and the range angle of the weapon engagement zone.

[0137] In this embodiment, during the training process of the to-be-trained deep reinforcement learning SAC model, when the target discriminator determines that the attack angle of the first friendly aircraft reaches the intermediate target angle, the target generator can update the intermediate target angle: determine the updated intermediate target angle based on at least the intermediate target angle and the difference between the intermediate target angle and the range angle of the weapon engagement zone. Among them, the intermediate target angle is greater than the range angle of the weapon engagement zone.

[0138] In an optional example, the target generator can obtain the updated intermediate target angle based on the following formula (13)

[0139]

[0140] where α is a decay factor that gradually reduces the impact of each adjustment is the intermediate target angle, f(Δδ) represents the angle reduction Δδ is the difference between the intermediate target angle and the range angle of the weapon engagement zone, β is a decay constant, and k is the iteration index determined by the target discriminator, which can ensure a smooth reduction of the angle adjustment amount when the agent approaches the target.

[0141] In an example, the target discriminator can be evaluated during the training process of the to-be-trained deep reinforcement learning SAC model by collecting the deduction data in the air combat simulation environment (such as the replay data of the recent k rounds) to determine whether the attack angle of the first friendly aircraft reaches the intermediate target angle. If the intermediate target angle is reached, the target generator will input the updated intermediate target angle into the policy network of the to-be-trained deep reinforcement learning SAC model to formulate a new intermediate task target.

[0142] Step S43: The target generator formulates a new intermediate task target for the policy network with the attack angle reaching the updated intermediate target angle until the final task target with the attack angle reaching the range angle of the weapon engagement zone is formulated for the policy network.

[0143] In this embodiment, after the target generator obtains the updated intermediate target angle, it can formulate a new intermediate task target for the policy network of the deep reinforcement learning SAC model to be trained, so that the attack angle reaches the updated intermediate target angle, guiding the intelligent agent to transition from simple tasks to complex tasks. The target generator continuously adjusts until it converges to the range angle of the weapon engagement zone, that is, until it formulates a final task target for the policy network so that the attack angle reaches the range angle of the weapon engagement zone, representing the ultimate challenge faced by the intelligent agent.

[0144] In this embodiment, the difficulty of the target task is dynamically adjusted according to the training progress of the model, and the intermediate task target is used to gradually guide the model to approach the final target. Among them, this embodiment refines the task parameters (adjust the target attack angle) to achieve effective curriculum learning, and the target generator gradually increases the task difficulty in a controlled manner according to the performance of the policy network. This dynamic adjustment mechanism ensures that the model can gradually master complex tactical behaviors during training, improve tactical decision-making capabilities, and aims to use the incremental difficulty adjustment strategy to guide the intelligent agent to complete increasingly challenging tasks.

[0145] In one embodiment, as Figure 4 shown, Figure 4 is a schematic flowchart of progressive task generation shown in an embodiment of the present invention. In Figure 4 , an adaptive target angle generation module is proposed. In this module, the target discriminator obtains replay data through the air combat simulation environment, and judges whether the attack angle has reached the intermediate target angle based on the replay data The In the case of reaching the intermediate target angle , the target generator generates a new intermediate target angle The new intermediate target angle is input to the UCAV intelligent agent for continued training to continuously adjust until the attack angle converges to the final target angle (that is, the range angle of the weapon engagement zone ), representing the ultimate challenge faced by the UCAV intelligent agent.

[0146] Combined with the above embodiments, in one implementation manner, the present invention also provides an air combat intelligent decision-making method based on curriculum reinforcement learning. In this method, step S13 above may specifically include step S51 and step S52:

[0147] Step S51: Based on the first K current actions of the first own aircraft output by the policy network in the deep reinforcement learning SAC model to be trained, perform simulation in the air combat simulation environment to obtain the winning rate of the first own aircraft in the last K rounds.

[0148] In this embodiment, after obtaining the current action of the first friendly aircraft output by the policy network in the deep reinforcement learning SAC model to be trained, the current action of the first friendly aircraft output by the policy network in the deep reinforcement learning SAC model to be trained can be input into the air combat simulation environment to perform a simulation, and a simulation result is obtained. The simulation result can at least include any one of the following: policy return, the win-loss result of the first friendly aircraft and the opposing aircraft in this game.

[0149] In this embodiment, based on the most recent K current actions of the first friendly aircraft output by the policy network in the deep reinforcement learning SAC model to be trained, a simulation is performed in the air combat simulation environment to obtain the winning rate of the first friendly aircraft in the most recent K games. For example, by simulating the most recent 10 current actions, it is determined that the first friendly aircraft won 6 out of the most recent 10 games, and the winning rate of these most recent 10 games is 60%.

[0150] Step S52: When the winning rate of the first friendly aircraft in the most recent K games exceeds the winning rate threshold, the progressive sliding buffer reads and stores M groups of new initial state data from the difficulty-guided state pool in the order of increasing difficulty, and discards M groups of old initial state data in the order of increasing difficulty.

[0151] In this embodiment, a winning rate threshold (which can be freely set) is preset, and this winning rate threshold determines the update frequency of the progressive sliding buffer. In this embodiment, the update frequency of the progressive sliding buffer can be controlled by the winning rate threshold. Whenever the winning rate of the first friendly aircraft in the most recent K games exceeds the winning rate threshold, the initial state data in the progressive sliding buffer is updated: the progressive sliding buffer reads and stores M groups of new initial state data from the difficulty-guided state pool in the order of increasing difficulty (such as the order from large to small of the corresponding combat advantage indicators), and discards M groups of old initial state data in the order of increasing difficulty, so as to ensure that there are always L groups of initial state data stored in the progressive sliding buffer; L and M are natural numbers greater than 0, and L is greater than M, and can be freely set.

[0152] This embodiment designs a progressive sliding buffer mechanism for dynamically adjusting the training sample state pool: during the training process, according to the learning state of the current model, more complex combat scenarios are gradually introduced to ensure the diversity of training samples, and at the same time, avoid the influence of overly simple samples on the training effect. This mechanism optimizes the quality of samples during training through a sliding window method, reduces the influence of invalid samples, and improves the training efficiency.

[0153] In one embodiment, as Figure 5 shown, Figure 5 is the update schematic diagram of the progressive sliding buffer shown in an embodiment of the present invention. In Figure 5A progressive sliding buffer module is proposed. In the training scheduler, the progressive sliding buffer module maintains a dynamic state buffer (i.e., the progressive sliding buffer), whose function is similar to a sliding window. The capacity of the progressive sliding buffer module is fixed at N, and each update integrates k s new and more difficult data points (i.e., k s sets of new initial state data with higher representational difficulty), while discarding k′ s (which can be the same as or different from k s ) of the oldest and less difficult entries (i.e., k′ s sets of the oldest initial state data with lower representational difficulty). Specifically, the state of the progressive sliding buffer at time t is represented as The update of the progressive sliding buffer is controlled by a sliding trigger. When the win rate W t of the first own aircraft exceeds the win rate threshold θ s1 , the sliding trigger will be activated: triggering the selection of new and more challenging states from the difficulty-guided state pool to ensure that the policy network gradually encounters increasingly difficult scenarios, thus supporting effective curriculum learning in the air combat mission.

[0154] Combined with the above embodiments, in one implementation, the present invention also provides an air combat intelligent decision-making method based on curriculum reinforcement learning. In this method, "sampling the current initial state data from the progressive sliding buffer" in the above step S15 may specifically include step S61:

[0155] Step S61: Sampling the current initial state data from the progressive sliding buffer or a random sampling generator through a dynamic hybrid sampling strategy.

[0156] In this embodiment, a dynamic hybrid sampling strategy is used to balance exploration and exploitation during training: instead of only sampling the current initial state data from the progressive sliding buffer, the current initial state data is sampled from the progressive sliding buffer or a random sampling generator through a dynamic hybrid sampling strategy.

[0157] In an optional example, the current initial state data can be sampled from the sliding buffer with probability p and from the random sampling generator (i.e., the random state sampler oracle(·)) with probability 1 - p, where 0 < p < 1. This method allows the agent to explore new states while exploiting the stored experience from the progressive sliding buffer which is dynamically adjusted to adapt to the agent's evolving skill level, thus promoting a steady improvement in performance.

[0158] In this embodiment, a hybrid sampling strategy is adopted, which combines sampling from a progressive sliding buffer and random sampling to improve the balance between exploration and exploitation. In this way, the agent can not only maintain the exploitation of known tasks but also explore new combat scenarios, quickly improving its tactical response ability.

[0159] Combined with the above embodiments, in one implementation, the present invention further provides an air combat intelligent decision-making method based on curriculum reinforcement learning. In this method, step S16 above may specifically include steps S71 to S74:

[0160] Step S71: Input the current action and current initial state data of the first friendly aircraft into the air combat simulation environment to perform simulation, determine whether the first friendly aircraft reaches the intermediate task target, and update the network parameters of the deep reinforcement learning SAC model to be trained after each simulation.

[0161] In this embodiment, the current action of the first friendly aircraft output by the policy network of the deep reinforcement learning SAC model to be trained and the corresponding current initial state data can be input into the air combat simulation environment to perform simulation, determine whether the first friendly aircraft reaches the corresponding intermediate task target, and update the network parameters of the deep reinforcement learning SAC model to be trained after each simulation.

[0162] Step S72: In the case where the first friendly aircraft does not reach the intermediate task target, sample new first initial state data from the progressive sliding buffer, input the new first initial state data into the policy network in the deep reinforcement learning SAC model to be trained, obtain a new first action of the first friendly aircraft, until it is determined that the first friendly aircraft reaches the intermediate task target based on the new first initial state data and the new first action for simulation.

[0163] In this embodiment, in the case where it is determined that the first friendly aircraft does not reach the intermediate task target, new first initial state data is sampled from the progressive sliding buffer. The new first initial state data is a new set of initial state data. The new first initial state data is input into the policy network in the deep reinforcement learning SAC model to be trained, and a new first action of the first friendly aircraft is obtained. The new first action of the first friendly aircraft output by the policy network of the deep reinforcement learning SAC model to be trained and the corresponding new first initial state data are input into the air combat simulation environment to perform simulation, determine whether the first friendly aircraft reaches the corresponding intermediate task target, until it is determined that the first friendly aircraft reaches the corresponding intermediate task target.

[0164] Step S73: When the first friendly aircraft reaches the intermediate mission objective, determine a new intermediate mission objective, sample new second initial state data from the progressive sliding buffer, and input the new second initial state data into the policy network in the deep reinforcement learning SAC model to be trained, so as to obtain a new second action of the first friendly aircraft.

[0165] In this embodiment, when it is determined that the first friendly aircraft reaches the corresponding intermediate mission objective, a new intermediate mission objective is determined (for example, the target generator formulates a new intermediate mission objective for the policy network of the deep reinforcement learning SAC model to be trained), and new second initial state data is sampled from the progressive sliding buffer. The new second initial state data is a new set of initial state data, which is different from the new first initial state data. The new second initial state data is input into the policy network in the deep reinforcement learning SAC model to be trained, so as to obtain a new second action of the first friendly aircraft.

[0166] Step S74: Input the new second action of the first friendly aircraft and the new second initial state data into the air combat simulation environment to perform simulation, and determine whether the first friendly aircraft reaches the new intermediate mission objective until it is determined through simulation that the first friendly aircraft reaches the final mission objective, so as to obtain the trained deep reinforcement learning SAC model.

[0167] In this embodiment, after obtaining the new second action of the first friendly aircraft output by the policy network of the deep reinforcement learning SAC model to be trained, the new second action of the first friendly aircraft and the corresponding new second initial state data are input into the air combat simulation environment to perform simulation, and it is determined whether the first friendly aircraft reaches the corresponding new intermediate mission objective. When it is determined that the first friendly aircraft does not reach the new intermediate mission objective, new second initial state data is resampled from the progressive sliding buffer, and the resampled new second initial state data is input into the policy network in the deep reinforcement learning SAC model to be trained to obtain a new second action of the first friendly aircraft again. The newly obtained new second action of the first friendly aircraft and the corresponding resampled new second initial state data are input into the air combat simulation environment to perform simulation, and it is determined whether the first friendly aircraft reaches the corresponding new intermediate mission objective until it is determined that the first friendly aircraft reaches the corresponding new intermediate mission objective. And so on, through multiple simulation training iterations, until it is determined through simulation that the first friendly aircraft reaches the final mission objective, so as to obtain the trained deep reinforcement learning SAC model.

[0168] Combined with any of the above embodiments, in one embodiment, the deep reinforcement learning SAC model established in this embodiment is an algorithm based on policy gradient, which realizes stable learning and better exploration by maximizing the expected return of the policy and the entropy of the policy. The objective function of the deep reinforcement learning SAC model mainly consists of two parts: policy entropy and expected return This algorithm introduces a soft Q function, which is summarized as follows:

[0169] where α is the temperature parameter, which is used to balance the influence of return and entropy; s t is the state variable, that is, the current initial state data; a t is the action vector, that is, the current action of the first own aircraft.

[0170] The deep reinforcement learning SAC model of this embodiment includes two value networks (basic value network and target value network) with the same network structure but different parameters and a policy network. The structure of this policy network is a multi-layer perceptron (MLP) with a hidden layer of [128, 128], and the activation function uses ReLU. The value network is similar to the policy network structure. In the model training of this embodiment, the Adam optimizer can be used, and the parameter update adopts the gradient descent method to update the network parameters by minimizing the loss function.

[0171] In one embodiment, the state variable s t can be constructed based on the six-degree-of-freedom (6-DOF) motion characteristics of the aircraft, including translational motion (position and velocity) and rotational motion (Euler angles and angular velocity). These variables enable the agent to accurately capture the spatial position and orientation of the aircraft, so that the state space can effectively represent and process the key state information in the dynamic air combat scenario. Although the continuous state space is usually represented by basic motion parameters, these parameters are affected by translational and rotational effects, so it is difficult to accurately capture the relative relationship between agents. Therefore, the state variable of this embodiment is mainly defined by the basic state variable of the aircraft and the relative variable between the aircrafts, so as to more accurately describe their relative positions and orientations. In this embodiment, the macro state consists of two consecutive frames, and the motion state vector in each frame is defined as follows:

[0172]

[0173] where h is the height of the first own aircraft, v n , v e and v d respectively represent the velocity components of the first own aircraft in the north, east and opponent directions, a x , a y and a zis the acceleration component of the first friendly aircraft, ψ is the yaw angle of the first friendly aircraft, θ is the pitch angle of the first friendly aircraft, φ is the roll angle of the first friendly aircraft, α is the angle of attack of the first friendly aircraft, and β is the sideslip angle of the first friendly aircraft. respectively represent the rates of change of the yaw angle, pitch angle, and roll angle. represents the relative variables between aircraft, including the relative distance d rel , the attack angle δ ATA , the escape angle σ AA and the relative velocity υ rel . As Figure 2 shown, the center of mass of the aircraft is connected by the line of sight (LOS line), and the LOS line is also used to determine the relative distance d rel . The escape angle σ AA is defined as the angle between the LOS line and the velocity vector of the target aircraft ( Figure 2 the blue aircraft in), while the attack angle δ ATA represents the angle between the LOS line and the velocity vector of the attacking aircraft ( Figure 2 the red aircraft in). The following equations define the calculation of these two angles:

[0174]

[0175] where the vector r rel represents the relative position, and the total velocity υ of the aircraft is characterized as the 2-norm of the velocity vector v = [υ n , υ e , υ d , where v T , v n , v e , and v d respectively represent the velocity components in the north, east, and opposite directions of the local, and can be expressed as:

[0176] In an alternative embodiment, the Weapons Engagement Zone (WEZ) is defined as a conical region with a cone angle of 30 degrees and a slant range of 1 kilometer. Aircraft entering this region will be locked and attacked by the opponent's missiles, as Figure 6 shown, Figure 6 is a schematic diagram of the weapons engagement zone shown in an embodiment of the present invention. Figure 6 Considering the relative distance r between the two aircraft, the thresholds of the WEZ are d1 = 200m and d2 = 1000m, and the calculation process is as follows:

[0177]

[0178] In an embodiment, the action vector a tIt includes four consecutive control commands: aileron deflection command elevator deflection command rudder deflection command and throttle command The first three commands respectively control the roll, pitch and yaw of the aircraft, so as to achieve precise attitude adjustment. At the same time, the throttle command adjusts the engine power, so as to manage the speed and energy of the aircraft. Therefore, the action vector at time step t can be expressed as

[0179] In one embodiment, in an autonomous air combat mission, the reward function provides feedback for the agent's actions, guiding it to achieve the mission objectives; the main objective is to quickly and strategically maneuver the unmanned fighter aircraft to force the opposing aircraft into the WEZ area to ensure the best target engagement conditions. The reward function design of this embodiment is divided into process reward (short-term reward) and terminal reward (long-term reward), which serve as signals for the agent to learn and make decisions at different stages of the mission. The process reward is calculated at each decision step to evaluate the current position and engagement situation of the own aircraft. This reward combines factors such as the relative distance, angle of attack, escape angle and time penalty with the opposing aircraft to encourage efficient task completion.

[0180] In a specific example, the process reward is shown as follows, where c1, c2, c3 and c4 are weighting coefficients that balance the contributions of each sub-reward term:

[0181]

[0182] Among them, is the distance reward term, which encourages the aircraft to shorten the distance from the opposing aircraft and maintain a tactical advantage by maintaining an effective engagement range: Among them, ρ1 is the sensitivity coefficient, d′ t represents the distance in the previous frame, d t represents the distance at the current time step t, and the function is bounded. As the distance approaches the optimal engagement range, the provided reward decreases.

[0183] is the angle of attack reward term, and the goal is to minimize the angle of attack, encouraging the aircraft to effectively align relative to the opposing aircraft: Among them, ρ2 is the sensitivity coefficient, δ′ t represents the angle of attack in the previous frame, and δ t is the angle of attack at the current time step t.

[0184] is the escape angle reward term, which promotes a favorable relative azimuth and helps the aircraft establish an effective tactical position: Among them, ρ3 is the sensitivity coefficient, σ′ tis the escape angle of the previous frame, σ t is the escape angle at the current time t.

[0185] tanh(-t) is the time penalty term to penalize the agent for prolonging the task time, so as to encourage it to quickly complete the task goal. The longer the time the agent spends, the heavier the penalty, thus cultivating its preference for taking prompt and decisive actions.

[0186] The terminal reward of this embodiment is determined according to the final state of the round task, and high rewards or penalties are provided based on the specific results of each round. The terminal reward includes the following conditions:

[0187] a. If our own aircraft successfully enters the WEZ area first, a reward of +50 is obtained;

[0188] b. If the opponent's aircraft enters the WEZ area first, a penalty of -50 is imposed;

[0189] c. If the distance between our own aircraft and the opponent's aircraft exceeds the predefined threshold, indicating separation, a penalty of -50 is assigned;

[0190] d. If the altitude of our own aircraft exceeds the allowed flight range, a penalty of -50 is imposed;

[0191] e. If the task duration exceeds the final time threshold, a penalty of -50 is imposed.

[0192] In this embodiment, the combination of short-term rewards (tactical positioning and engagement metrics) and long-term rewards (task end state) can guide the UCAV agent to balance immediate actions and overall task goals. This dual structure can achieve optimal decision-making, thereby improving task efficiency and tactical effectiveness.

[0193] Combined with the above embodiments, in one embodiment, this embodiment can train the deep reinforcement learning SAC model to be trained based on the above state variables, action vectors, and rewards. The training process at least includes the following steps:

[0194] 1. Initialize the basic value network Q(s * , a t |ω t ), the target value network Q Q (s * , a t |ω t ), and the policy network π(a Q* |s t ) with random network parameters ω, ω t ; among them, the policy network π(a t |s t ): The input is the state s t , and the output is the action at ; The basic value network Q(s t , a t |ω Q ): The input is the state s t and the action a t , and the output is the Q value; The target value network Q * (s t , a t |ω Q* ): It has the same structure as the basic value network and is used for stable training. It is synchronized with the basic value network through soft update. Initialize the experience replay buffer R to store the experience data (current state, current action, reward, next state) sampled by the agent in the environment;

[0195] 2. Obtain the state vector s of the air combat simulation environment at time t t . If it is the initial time, then t = 0, and the initial state s t is obtained from the progressive sliding buffer through the hybrid sampler or the random sampling oracle(·);

[0196] 3. Input the initial state s t into the policy network π(a t |s t ), and obtain the output a t . In an optional embodiment, the action a t ~Π θ (·|s t )+ε can be selected, where ε is the noise term for exploring the environment; that is, the action a t |s t is obtained by adding the noise term to the output of the policy network π(a t );

[0197] 4. Input the obtained aircraft control action (i.e., a t ) into the environment and perform a simulation once to obtain the next frame state s t+1 , the reward value r t , and store (s t , a t , r t , s t+1 ) into the experience replay buffer R;

[0198] 5. If the number of samples in the experience replay buffer R is greater than the batch size, randomly sample an N-tuple {(s i , a i , r i , s i+1 )} i=1,...,NLearn, calculate the target Q value through the target value network, and the calculation process is as follows:

[0199] a i+1 ~Π θ (·|s s+1 );

[0200] Among them, r i is the reward, γ is the influence factor, and α is the temperature coefficient of the entropy regularization term, which is used to control the balance between exploration and exploitation. Update the basic value network through the loss function . Subsequently, use reparameterization to sample actions Update the policy network through the loss function , and perform a soft update on the target value network: ω * ←τω+(1-τ)ω * , where τ is the soft update coefficient;

[0201] 6. Update the environmental state to the next state s t+1 , and then return to step 2;

[0202] 7. If the termination state is reached (completing the current intermediate task goal, the aircraft altitude exceeding the boundary, the aircraft distance exceeding the boundary, or the time step exceeding the maximum value, etc.), then reset the simulation environment, update the progressive sliding buffer and the intermediate task goal, and start the next round of training until the training ends.

[0203] In one embodiment, this embodiment proposes an air combat intelligent decision-making method based on curriculum reinforcement learning, and designs an automatic engagement curriculum learning framework, as Figure 7 shown, Figure 7 is a schematic diagram of an automatic engagement curriculum learning framework shown in an embodiment of the present invention. This embodiment combines dynamic curriculum learning and deep reinforcement learning technologies, and can effectively perform flight control and intelligent decision-making on a six-degree-of-freedom unmanned combat aircraft. In Figure 7 , the automatic engagement curriculum learning framework includes:

[0204] Automated combat complexity assessment: This embodiment designs a combat superiority index (CSI), which quantifies the difficulty (complexity) of the combat scenario by dynamically evaluating multiple key factors in the combat scenario (such as engagement angle, energy advantage, the positional relationship between the two sides, etc.). The CSI index includes the relative distance, attack angle, energy consumption, etc. of the aircraft on both sides of the operation, ensuring that the model gradually adapts to increasingly complex combat situations during training, and finally generates a difficulty-guided state pool (that is, a state pool that can be sorted by complexity), so that the training difficulty can be dynamically adjusted according to the model feedback.

[0205] Progressive Sliding Buffer Mechanism: In this embodiment, a progressive sliding buffer mechanism is designed to dynamically adjust the training sample state pool (i.e., the progressive sliding buffer). During the training process, according to the learning state of the current model, more complex combat scenarios are gradually introduced to ensure the diversity of training samples while avoiding the influence of overly simple samples on the training effect. This mechanism optimizes the quality of samples during training through a sliding window method, reduces the impact of invalid samples, and improves training efficiency.

[0206] Goal-Driven Progressive Task Generation: Under this framework, the difficulty of tasks increases gradually. The system dynamically adjusts the difficulty of the target tasks according to the training progress of the model, and uses intermediate goals (such as adjusting the target attack angle) to gradually guide the model towards the final goal (the endgame goal). By designing an intermediate goal guidance mechanism, the difficulty of combat task goals (such as attack angle) is dynamically adjusted, and the model is gradually guided to complete complex tactical decision-making tasks, ensuring that the model can gradually master complex tactical behaviors during training and improve tactical decision-making ability.

[0207] Hybrid Sampling Strategy: This embodiment adopts a hybrid sampling strategy that combines sampling from the progressive sliding buffer and random sampling to improve the balance between exploration and exploitation. In this way, the model can both maintain the utilization of known tasks and explore new combat scenarios, quickly improving its tactical response ability.

[0208] Through this automatic engagement curriculum learning framework, this embodiment includes integrated optimization in multiple aspects such as automated combat complexity assessment, task goal generation, progressive buffer mechanism, and hybrid sampling strategy. It can significantly accelerate the training process, reduce computational resource requirements, ensure the efficiency and adaptability of the training process, and effectively improve air combat decision-making ability while ensuring the decision-making quality of the policy network.

[0209] In one embodiment, as Figure 8 shown, Figure 8 is the overall process schematic diagram of an air combat intelligent decision-making method based on curriculum reinforcement learning shown in an embodiment of the present invention. In Figure 8 it, the method includes the following steps:

[0210] Step 1: Randomly initialize the training scenario to generate randomly initialized training scenario data, which at least includes: the initial longitude and latitude, initial speed, initial altitude, and initial yaw angle of the first friendly aircraft, as well as the relative distance, relative speed, relative altitude, relative heading angle, and relative azimuth angle of the enemy aircraft relative to the first friendly aircraft. For example, 10,000 groups of training scenario data are generated;

[0211] Step 2: Calculate the CSI metrics of each set of initial state data corresponding to each set of training scenarios in the above steps, including relevant angles, distances, energy consumption, etc. of the aircraft on both sides of the combat. Sort the 10,000 generated CSI metrics in descending order. The larger the value, the higher the advantage of the first own aircraft and the lower the difficulty of the combat scenario of the mission. The 10,000 sets of initial state data sorted according to the 10,000 CSI metrics are stored in the state pool guided by difficulty;

[0212] Step 3: Generate a progressive sliding buffer with a capacity of k, design a sliding trigger update mechanism, calculate the winning rate of the first own aircraft in the recent k rounds, control the update frequency of the buffer through the winning rate threshold. Whenever the winning rate reaches the specified threshold, gradually slide and update the initial state data in the progressive sliding buffer in ascending order of difficulty from the state pool guided by difficulty. The initialization of the progressive sliding buffer is the k sets of initial state data with the highest ranking in the state pool guided by difficulty;

[0213] Step 4: Design a mission objective generator, including an objective generator and an objective discriminator. Select the replay data of the recent k rounds from the air combat simulation environment, calculate and determine whether the first own aircraft has reached the specified intermediate mission objective. If the objective is reached, the objective generator will input a new intermediate mission objective into the intelligent agent model to guide the intelligent agent to transition from simple tasks to complex tasks until the final mission objective is input into the intelligent agent model;

[0214] Step 5: Train the deep reinforcement learning SAC algorithm. Its input includes the current state information of the aircraft on both sides of the combat, that is, the current initial state data (such as position, speed, and altitude, etc.), and the output is the policy decision of the first own aircraft, that is, the current actions of the first own aircraft, including action values such as elevator, rudder, aileron, and throttle; at the same time, design a reward mechanism, giving a positive reward for successfully chasing the opponent's aircraft, otherwise giving a negative reward; during the training process, the opponent's aircraft keeps flying straight, and through multiple training iterations, gradually optimize the chasing strategy of the first own aircraft;

[0215] Step 6: Use the objective discriminator to judge the policy performance of the current first own aircraft. If the pre-specified final mission objective has not been reached, return to Step 4 to continue training. If the final mission objective has been reached, proceed to the next step;

[0216] Step 7: Deploy the policy model optimized through training to the decision-making production unit. The decision-making production unit will call the optimal policy in real time according to the current air combat situation and make intelligent decisions.

[0217] It should be noted that for method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that the embodiments of the present invention are not limited by the described action sequences, because according to the embodiments of the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential for the embodiments of the present invention.

[0218] Based on the same inventive concept, an embodiment of the present invention provides an air combat intelligent decision-making device based on curriculum reinforcement learning. Refer to Figure 9 , Figure 9 which is a structural block diagram of an air combat intelligent decision-making device based on curriculum reinforcement learning provided by an embodiment of the present invention. As Figure 9 shown, the air combat intelligent decision-making device based on curriculum reinforcement learning in this embodiment may include:

[0219] An index determination module, configured to determine a combat advantage index corresponding to each group of initial state data based on multiple groups of initial state data, where the combat advantage index characterizes the difficulty of the combat scenario for the first own aircraft;

[0220] A state pool storage module, configured to sort the multiple groups of initial state data based on the combat advantage index and store the sorted multiple groups of initial state data into a difficulty-guided state pool;

[0221] A buffer update module, configured to update the initial state data in the progressive sliding buffer in an increasing order of difficulty from the difficulty-guided state pool;

[0222] A target formulation module, configured to dynamically formulate intermediate task targets for the policy network in the deep reinforcement learning SAC model to be trained, where the intermediate task targets guide the policy network to transition from simple tasks to complex tasks until the final task target is formulated for the policy network;

[0223] An action output module, configured to sample the current initial state data from the progressive sliding buffer, input the current initial state data into the policy network in the deep reinforcement learning SAC model to be trained, and obtain the current action of the first own aircraft;

[0224] A model training module, configured to train the deep reinforcement learning SAC model to be trained at least based on the intermediate task targets, the current action of the first own aircraft, and the current initial state data until the first own aircraft reaches the final task target, and obtain a trained deep reinforcement learning SAC model;

[0225] A decision output module, configured to input the current air combat environment information into a policy network in the trained deep reinforcement learning SAC model, and output action information of a second own aircraft.

[0226] Optionally, each group of initial state data at least includes: the engagement angle, relative distance, relative flight altitude between the first own aircraft and the opponent aircraft, and the flight speed of the first own aircraft; the index determination module includes:

[0227] A first determination module, configured to determine an angle advantage value based on the engagement angle in each group of initial state data;

[0228] A second determination module, configured to determine a distance advantage value based on the relative distance in each group of initial state data;

[0229] A third determination module, configured to determine an energy advantage value based on the relative flight altitude and flight speed in each group of initial state data;

[0230] A fourth determination module, configured to determine a combat advantage index corresponding to each group of initial state data based on at least the angle advantage value, the distance advantage value, and the energy advantage value, where the larger the value of the combat advantage index, the lower the difficulty of the combat scenario for the first own aircraft.

[0231] Optionally, the engagement angle includes: an attack angle and an escape angle; the first determination module includes:

[0232] A first advantage determination module, configured to determine a detection advantage value of the attack angle based on at least the maximum detection angle of the radar and the attack angle;

[0233] A second advantage determination module, configured to determine a threat advantage value of the attack angle based on at least the angle corresponding to the maximum engagement area, the missile lock angle, the non-escape attack area angle, and the attack angle;

[0234] A third advantage determination module, configured to determine a threat advantage value of the escape angle based on at least the escape angle;

[0235] A fourth advantage determination module, configured to determine the angle advantage value based on at least the detection advantage value of the attack angle, the threat advantage value of the attack angle, and the threat advantage value of the escape angle;

[0236] The second determination module includes:

[0237] A fifth advantage determination module, configured to determine a detection advantage value of the distance based on at least the maximum detection distance of the radar, the missile lock distance, and the relative distance;

[0238] A sixth advantage determination module, configured to determine a threat advantage value of distance based on at least the maximum attack distance, non-escape distance, minimum attack distance of the missile, and the relative distance;

[0239] A seventh advantage determination module, configured to determine the distance advantage value based on at least the detection advantage value of the distance and the threat advantage value of the distance;

[0240] A third determination module includes:

[0241] An eighth advantage determination module, configured to determine a potential energy advantage value based on at least the optimal flight altitude, maximum flight altitude, and the relative flight altitude;

[0242] A ninth advantage determination module, configured to determine a kinetic energy advantage value based on at least the kinetic energy of the other aircraft and the kinetic energy of the first own aircraft determined based on the flight speed of the first own aircraft;

[0243] A tenth advantage determination module, configured to determine the energy advantage value based on at least the potential energy advantage value and the kinetic energy advantage value.

[0244] Optionally, intermediate task objectives are dynamically formulated for the policy network in the deep reinforcement learning SAC model to be trained, and the intermediate task objectives guide the policy network to transition from simple tasks to complex tasks until the final task objective is formulated for the policy network, including:

[0245] An angle generation module, configured to generate an intermediate target angle for the target generator and formulate an intermediate task objective for the policy network to reach the intermediate target angle for the attack angle;

[0246] An angle update module, configured to, during the training process of the deep reinforcement learning SAC model to be trained, when the target discriminator determines that the attack angle of the first own aircraft reaches the intermediate target angle, the target generator determines an updated intermediate target angle based on at least the intermediate target angle and the difference between the intermediate target angle and the range angle of the weapon engagement zone;

[0247] An angle formulation module, configured to the target generator formulate a new intermediate task objective for the policy network to reach the updated intermediate target angle for the attack angle until the final task objective for the policy network to reach the range angle of the weapon engagement zone is formulated;

[0248] Wherein, the intermediate target angle is greater than the range angle of the weapon engagement zone.

[0249] Optionally, a buffer update module includes:

[0250] A winning rate determination module, configured to perform simulation in an air combat simulation environment based on the last K current actions of the first own aircraft output by the policy network in the to-be-trained deep reinforcement learning SAC model, and obtain the winning rate of the first own aircraft in the last K games;

[0251] An update processing module, configured to, when the winning rate of the first own aircraft in the last K games exceeds the winning rate threshold, the progressive sliding buffer reads and stores M groups of new initial state data from the difficulty-guided state pool in an increasing difficulty manner, and discards M groups of old initial state data in a manner from low to high difficulty;

[0252] Wherein, K and M are natural numbers greater than 0.

[0253] Optionally, the action output module includes:

[0254] A hybrid sampling module, configured to sample the current initial state data from the progressive sliding buffer or a random sampling generator through a dynamic hybrid sampling strategy.

[0255] Optionally, the model training module includes:

[0256] A simulation execution module, configured to input the current action and the current initial state data of the first own aircraft into the air combat simulation environment to perform simulation, determine whether the first own aircraft reaches the intermediate task target, and update the network parameters of the to-be-trained deep reinforcement learning SAC model after each simulation;

[0257] A task judgment module, configured to, when the first own aircraft does not reach the intermediate task target, sample new first initial state data from the progressive sliding buffer, input the new first initial state data into the policy network in the to-be-trained deep reinforcement learning SAC model, and obtain a new first action of the first own aircraft until it is determined that the first own aircraft reaches the intermediate task target based on the new first initial state data and the new first action through simulation;

[0258] A task update module, configured to, when the first own aircraft reaches the intermediate task target, determine a new intermediate task target, sample new second initial state data from the progressive sliding buffer, and input the new second initial state data into the policy network in the to-be-trained deep reinforcement learning SAC model to obtain a new second action of the first own aircraft;

[0259] A training result determination module is configured to input the new second action and the new second initial state data of the first own aircraft into the air combat simulation environment to perform simulation, and determine whether the first own aircraft reaches the new intermediate mission objective until it is determined through simulation that the first own aircraft reaches the final mission objective, so as to obtain the trained deep reinforcement learning SAC model.

[0260] Based on the same inventive concept, another embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the air combat intelligent decision-making method based on curriculum reinforcement learning as described in any one of the above embodiments of the present invention are implemented.

[0261] Based on the same inventive concept, another embodiment of the present invention provides an electronic device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes, the steps in the air combat intelligent decision-making method based on curriculum reinforcement learning as described in any one of the above embodiments of the present invention are implemented.

[0262] For the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, refer to the partial description of the method embodiments. Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the embodiments, refer to each other.

[0263] The above has introduced in detail a method, device, equipment and medium for air combat intelligent decision-making based on curriculum reinforcement learning provided by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. An air combat intelligent decision-making method based on curriculum reinforcement learning, characterized in that The method includes: Based on multiple sets of initial state data, determining a combat advantage index corresponding to each set of initial state data, where the combat advantage index characterizes the difficulty of the combat scenario for the first friendly aircraft; Sorting the multiple sets of initial state data based on the combat advantage index, and storing the sorted multiple sets of initial state data in a difficulty-guided state pool; Updating the initial state data in the progressive sliding buffer in an increasing difficulty manner from the difficulty-guided state pool; Dynamically formulating intermediate task objectives for the policy network in the deep reinforcement learning SAC model to be trained, where the intermediate task objectives guide the policy network to transition from simple tasks to complex tasks until the final task objective is formulated for the policy network; Sampling the current initial state data from the progressive sliding buffer, and inputting the current initial state data into the policy network in the deep reinforcement learning SAC model to be trained to obtain the current action of the first friendly aircraft; Training the deep reinforcement learning SAC model to be trained at least based on the intermediate task objectives, the current action of the first friendly aircraft, and the current initial state data until the first friendly aircraft reaches the final task objective to obtain a trained deep reinforcement learning SAC model; Inputting the current air combat environment information into the policy network in the trained deep reinforcement learning SAC model to output the action information of the second friendly aircraft.

2. The air combat intelligent decision-making method based on curriculum reinforcement learning according to claim 1, wherein Each set of initial state data at least includes: the engagement angle, relative distance, relative flight altitude, and flight speed of the first friendly aircraft and the opposing aircraft; determining the combat advantage index corresponding to each set of initial state data based on multiple sets of initial state data includes: Determining an angle advantage value based on the engagement angle in each set of initial state data; Determining a distance advantage value based on the relative distance in each set of initial state data; Determining an energy advantage value based on the relative flight altitude and flight speed in each set of initial state data; Determining the combat advantage index corresponding to each set of initial state data at least based on the angle advantage value, the distance advantage value, and the energy advantage value, where the larger the value of the combat advantage index, the lower the difficulty of the combat scenario for the first friendly aircraft.

3. The air combat intelligent decision-making method based on curriculum reinforcement learning according to claim 2, wherein, The engagement angle includes: the attack angle and the escape angle; determining the angle advantage value based on the engagement angle in each set of initial state data includes: Determining the detection advantage value of the attack angle at least based on the maximum detection angle of the radar and the attack angle; Determining the threat advantage value of the attack angle at least based on the angle corresponding to the maximum engagement area, the missile lock angle, the non-escape attack area angle, and the attack angle; Determining the threat advantage value of the escape angle at least based on the escape angle; Determining the angle advantage value at least based on the detection advantage value of the attack angle, the threat advantage value of the attack angle, and the threat advantage value of the escape angle; Determining the distance advantage value based on the relative distance in each set of initial state data includes: Determine the detection advantage value of the distance based at least on the maximum detection distance of the radar, the missile lock - on distance, and the relative distance; Determine the threat advantage value of the distance based at least on the maximum attack distance of the missile, the non - escapable distance, the minimum attack distance, and the relative distance; Determine the distance advantage value based at least on the detection advantage value of the distance and the threat advantage value of the distance; Determine the energy advantage value based on the relative flight altitude and flight speed in each set of initial state data, including: Determine the potential energy advantage value based at least on the optimal flight altitude, the maximum flight altitude, and the relative flight altitude; Determine the kinetic energy advantage value based at least on the kinetic energy of the opposing aircraft and the kinetic energy of the first own aircraft determined based on the flight speed of the first own aircraft; Determine the energy advantage value based at least on the potential energy advantage value and the kinetic energy advantage value; 4. The air combat intelligent decision-making method based on curriculum reinforcement learning according to claim 1, characterized in that Dynamically formulate intermediate task objectives for the policy network in the deep reinforcement learning SAC model to be trained. The intermediate task objectives guide the policy network to transition from simple tasks to complex tasks until the final task objective is formulated for the policy network, including: The target generator generates an intermediate target angle and formulates an intermediate task objective for the policy network to achieve the attack angle reaching the intermediate target angle; During the training process of the deep reinforcement learning SAC model to be trained, when the target discriminator determines that the attack angle of the first own aircraft reaches the intermediate target angle, the target generator determines the updated intermediate target angle based at least on the intermediate target angle and the difference between the intermediate target angle and the range angle of the weapon engagement zone; The target generator formulates a new intermediate task objective for the policy network to achieve the attack angle reaching the updated intermediate target angle until the final task objective for the policy network to achieve the attack angle reaching the range angle of the weapon engagement zone; Wherein, the intermediate target angle is greater than the range angle of the weapon engagement zone.

5. The air combat intelligent decision-making method based on curriculum reinforcement learning according to any one of claims 1 to 4, characterized in that, Update the initial state data in the progressive sliding buffer in an increasing - difficulty manner from the difficulty - guided state pool, including: Based on the last K current actions of the first own aircraft output by the policy network in the deep reinforcement learning SAC model to be trained, perform simulation in the air combat simulation environment to obtain the winning rate of the first own aircraft in the last K games; When the winning rate of the first own aircraft in the last K games exceeds the winning rate threshold, the progressive sliding buffer reads and stores M groups of new initial state data from the difficulty - guided state pool in an increasing - difficulty manner and discards M groups of old initial state data in a low - to - high difficulty manner; Wherein, K and M are natural numbers greater than 0.

6. The air combat intelligent decision-making method based on curriculum reinforcement learning according to any one of claims 1 to 4, characterized in that Sample the current initial state data from the progressive sliding buffer, including: Sample the current initial state data from the progressive sliding buffer or the random sampling generator through a dynamic hybrid sampling strategy.

7. The air combat intelligent decision-making method based on curriculum reinforcement learning according to any one of claims 1 to 4, characterized in that Train the deep reinforcement learning SAC model to be trained at least based on the intermediate task objective, the current action of the first friendly aircraft, and the current initial state data until the first friendly aircraft reaches the final task objective, to obtain a trained deep reinforcement learning SAC model, including: Input the current action and current initial state data of the first friendly aircraft into the air combat simulation environment to perform simulation, determine whether the first friendly aircraft reaches the intermediate task objective, and update the network parameters of the deep reinforcement learning SAC model to be trained after each simulation; In the case that the first friendly aircraft does not reach the intermediate task objective, sample new first initial state data from the progressive sliding buffer, input the new first initial state data into the policy network in the deep reinforcement learning SAC model to be trained, to obtain a new first action of the first friendly aircraft, until it is determined that the first friendly aircraft reaches the intermediate task objective based on the new first initial state data and the new first action for simulation; In the case that the first friendly aircraft reaches the intermediate task objective, determine a new intermediate task objective, sample new second initial state data from the progressive sliding buffer, input the new second initial state data into the policy network in the deep reinforcement learning SAC model to be trained, to obtain a new second action of the first friendly aircraft; Input the new second action and new second initial state data of the first friendly aircraft into the air combat simulation environment to perform simulation, determine whether the first friendly aircraft reaches the new intermediate task objective, until it is determined through simulation that the first friendly aircraft reaches the final task objective, to obtain the trained deep reinforcement learning SAC model.

8. An air combat intelligent decision-making device based on curriculum reinforcement learning, characterized in that, The device includes: An index determination module, configured to determine a combat advantage index corresponding to each group of initial state data based on multiple groups of initial state data, where the combat advantage index characterizes the difficulty of the combat scenario for the first friendly aircraft; A state pool storage module, configured to sort the multiple groups of initial state data based on the combat advantage index, and store the sorted multiple groups of initial state data into a difficulty-guided state pool; A buffer update module, configured to update the initial state data in the progressive sliding buffer in an increasing difficulty manner from the difficulty-guided state pool; An objective formulation module, configured to dynamically formulate intermediate task objectives for the policy network in the deep reinforcement learning SAC model to be trained, where the intermediate task objectives guide the policy network to transition from simple tasks to complex tasks until the final task objective is formulated for the policy network; An action output module, configured to sample current initial state data from the progressive sliding buffer, input the current initial state data into the policy network in the deep reinforcement learning SAC model to be trained, to obtain the current action of the first friendly aircraft; A model training module, configured to train the deep reinforcement learning SAC model to be trained at least based on the intermediate task objective, the current action of the first own aircraft, and the current initial state data until the first own aircraft reaches the final task objective, so as to obtain a trained deep reinforcement learning SAC model; A decision output module, configured to input the current air combat environment information into a policy network in the trained deep reinforcement learning SAC model and output the action information of the second own aircraft.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the computer program is executed by the processor, it implements the air combat intelligent decision-making method based on curriculum reinforcement learning according to any one of claims 1 to 7.

10. A computer-readable storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by the processor, it implements the air combat intelligent decision-making method based on curriculum reinforcement learning according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Unmanned aerial vehicle maneuvering decision-making method based on deep reinforcement learning

    CN113093802A

  • Multi-agent reinforcement learning-based multi-machine air combat decision-making method

    CN113791634A

  • Robot strategy training system and training method based on course reinforcement learning

    CN115454096A

  • Aircraft close-range air combat situation assessment adaptive weight design method

    CN115688268A

  • Unmanned aerial vehicle air combat autonomous decision-making method based on SAC algorithm and alliance training

    CN116187777A

Cited By

  • Positioning method and device based on multi-image reasoning

    CN122049569A

  • Fixed-wing aircraft high-maneuverability flight control method based on curriculum-type reinforcement learning

    CN122331303A

  • Fixed-wing aircraft high-maneuver flight control method based on course-based reinforcement learning

    CN122331303B