Radar mode and maneuvering decision collaborative strategy generation method in simulation environment
By designing a deep reinforcement learning algorithm framework based on EDN-PPOA, the coupling problem between radar working mode and maneuvering behavior in aircraft simulation confrontation was solved, efficient coordination between radar and maneuvering strategy was achieved, and the stability and robustness of the strategy were improved.
Patent Information
- Application Number
- CN202510967658.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2025-10-14
AI Technical Summary
Existing reinforcement learning algorithms lack analysis and modeling of radar operating modes in aircraft simulation confrontation, resulting in reduced strategy reliability. In addition, there is a complex coupling relationship between radar switching and maneuvering behavior, which affects the accuracy of state information in the observation space.
A deep reinforcement learning algorithm framework based on EDN-PPOA is designed. Combined with the radar working mode and working rule constraints, the state space and hybrid action space are constructed. The entropy concept and noise network are introduced to improve the robustness of the model, and the collaborative decision-making of radar and maneuver strategy is realized.
Efficient coordination between radar and maneuverability is achieved in complex mixed action spaces, with strategy performance superior to traditional methods, achieving strategy stability and robustness in observation noisy environments.
Smart Images

Figure CN120779746A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer simulation and artificial intelligence, and in particular relates to a method for generating a radar mode and maneuver decision-making collaborative strategy in a simulated environment. Background Art
[0002] The aircraft simulation and confrontation system uses computer simulation to provide a detailed and realistic simulation of the entire aircraft mission execution process. To effectively enhance the user experience and the controllability of the confrontation game and simulation system, it is necessary to simulate and design the confrontation game and simulation system from the perspective of actual aircraft simulation and confrontation. More importantly, strategy simulation and convenient interaction design are required to restore the realism of aircraft simulation and confrontation while improving the user's control level in the confrontation game and simulation system.
[0003] Executing a rational, rapid, and efficient maneuvering strategy during simulated aircraft confrontations with targets can create a situational advantage, create a closed-loop time difference between the two sides, and thus seize a strategic advantage. This is crucial for seizing air superiority and defeating the enemy. Based on situational information collected by sensors and expert experience, mission execution objectives are converted into corresponding mathematical models, and the optimal strategy is discovered through intelligent algorithms. However, as the realism of the simulation environment increases, aircraft simulated confrontations also involve strategic behaviors such as radar mode selection and switching, making the strategic behavior space more complex. Therefore, it is necessary to study hybrid decision-making strategies for radar mode selection and maneuvering, taking into account the characteristics of different radar operating modes, to achieve more accurate and efficient strategic coordination in both time and space.
[0004] With AlphaGo achieving outstanding results, artificial intelligence algorithms, such as deep reinforcement learning, have been widely used in the field of simulated aircraft combat. Researchers have primarily sought to address issues such as reward function design and policy training convergence in simulated aircraft combat. The paper "Short-range air combat maneuver decision of UAV swarm based on multi-agent Transformer introducing virtual objects" (JIANG F, XU M, LI Y, et al. Short-range air combat maneuver decision of UAV swarm based on multi-agent Transformer introducing virtual objects [J]. Engineering Applications of Artificial Intelligence, 2023, 123: 106358.) uses a situation assessment function as a reward function for close-range aircraft combat scenarios. 《Aerial combat maneuvering policy learning based on confrontation demonstrations and dynamic quality replay》(HUD, YANG R, ZHANG Y, et al. Aerial combat maneuvering policy learning based on confrontation demonstrations and dynamic quality replay[J]. Engineering Applications of Artificial Intelligence, 2022, 111: 104767.) Considering the characteristics of beyond-visual-range aircraft simulation confrontation mainly based on attacks by autonomous aircraft, the reward function is designed in combination with the attack area of the autonomous aircraft.
[0005] However, there is currently little research on the action space in simulated aircraft combat. Depending on the type of action being performed, the action space can be divided into an executable strategy set, a maneuver set, and a posture control set. "Mastering air combat game with deep reinforcement learning" (ZHU J, KUANG M, ZHOU W, et al. Mastering air combat game with deep reinforcement learning [J]. Defence Technology, 2024, 34: 295-312.) models the maneuver space as discrete maneuvers: direct flight, pursuit, circling, somersault, attack, and evasion. Subsequently, to achieve more precise and complex maneuvers, a three-degree-of-freedom kinematic model and a six-degree-of-freedom kinematic model were used.
[0006] However, the above prior art has the following problems: Existing reinforcement learning research primarily focuses on simulated aircraft maneuvers during confrontations, lacking analysis and modeling of radar operating modes. During simulated confrontations, the choice of radar operating mode directly impacts detection effectiveness and accuracy, thus interfering with behavioral decision-making and reducing strategy reliability.
[0007] 2) Radar mode switching and maneuvering together form a complex hybrid action space that is both discrete and continuous. Existing reinforcement learning algorithms are limited to a single action space. Furthermore, radar switching and maneuvering interact with each other, creating a complex coupling relationship.
[0008] 3) Introducing a radar operating mechanism can introduce errors in the state information of the observation space. Most studies fail to consider the differences in situational information accuracy caused by different radar operating modes. The instability and volatility of the state observation space will directly affect the reinforcement learning strategy learning process and the effectiveness of the adversarial strategy. Summary of the Invention
[0009] To address the aforementioned issues in the prior art, the present invention provides a method for generating a coordinated strategy for radar patterns and maneuvering decisions in a simulated environment. The technical issues addressed by the present invention are achieved through the following technical solutions: An embodiment of the present invention provides a method for generating a radar mode and maneuver decision-making collaborative strategy in a simulated environment, the method comprising: Analyze the functional characteristics of different radar operating modes and design corresponding operating criteria constraints based on the analysis results; Constructing an attack reward function based on the three factors of angle, distance, and height, constructing a radar reward function based on the working criteria constraints, and constructing a total reward function based on the attack reward function and the radar reward function; Combined with the radar working mode and working criteria constraints, a state space consisting of aircraft state information, target state information, relative situation information, and sensor state information is defined, and a hybrid action space consisting of maneuver action space and radar action space is defined; A deep reinforcement learning algorithm framework based on EDN-PPOA is designed. Based on the total reward function, the state space, and the mixed action space, the network parameters are adjusted to output a mixed action policy model corresponding to the optimal network parameters. The designed deep reinforcement learning algorithm framework based on EDN-PPOA is as follows: on the basis of the traditional PPO algorithm, an advantage function combining GAE and the entropy of the policy distribution is adopted in the critic network processing, a Dueling network is introduced in the actor network processing to decouple the mixed action space policy and a noise network is introduced to improve the robustness of the model, and an overall objective optimization function is constructed based on the critic network processing and the actor network processing; The current state space is input into the hybrid action strategy model, and the corresponding hybrid action space is output to achieve autonomous decision-making on maneuvering and radar working mode selection.
[0010] Beneficial effects of the present invention: The proposed method for generating a coordinated strategy for radar mode and maneuvering decision-making in a simulated environment provides a reasonable and feasible solution for achieving efficient strategic coordination between radar and maneuvering strategies during confrontation. It implements hybrid action coordinated strategy generation for radar operating modes and maneuvering decisions. Specifically, the functional characteristics of different radar operating modes are analyzed, corresponding working criteria constraints are designed, and a total reward function consisting of attack reward function and radar reward function is designed to reduce ineffective actions of the agent. Then, the appropriate state space and hybrid action space are redefined according to the strategy generation requirements. Finally, a deep reinforcement learning algorithm framework based on EDN-PPOA is proposed. The robustness and exploration capability of the network model are improved by introducing the concept of entropy and a noise network. The coupling relationship between radar and maneuvering strategies is also considered, and the hybrid action space is decoupled by combining the Dueling network framework. Simulation experimental results show that the proposed method achieves hybrid action coordinated strategy generation for radar operating modes and maneuvering decisions. It can also achieve efficient coordination between radar and maneuvering in complex hybrid action spaces, and its strategy performance is superior to traditional empirical methods. Furthermore, the proposed method demonstrates significant strategy stability and robustness in an observation noise environment, providing important technical support for intelligent decision-making in simulated aircraft confrontation.
[0011] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 This is a flow chart of a method for generating a radar mode and maneuver decision-making collaborative strategy in a simulated environment provided by an embodiment of the present invention; Figure 2 Schematic diagram of functional modeling of radar operating modes provided by an embodiment of the present invention; Figure 3 This is a schematic diagram of the deep reinforcement learning algorithm framework based on EDN-PPOA designed in an embodiment of the present invention; Figure 4 2 is a schematic diagram comparing the reward changes of various algorithms during the training process provided by the embodiments of the present invention; Figure 5(a) to Figure 5(b) 2 is a schematic diagram showing the comparison results of the mean and variance of the cumulative rewards in the last 50 rounds of training for each algorithm provided in an embodiment of the present invention; Figure 6(a) to Figure 6(b) 1 is a schematic diagram of the robustness ablation experiment results of the attack intention strategy under random situations provided by an embodiment of the present invention; Figure 7 Schematic diagram of action visualization results of different algorithms under the attack strategy provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0013] The present invention will be further described in detail below with reference to specific examples, but the embodiments of the present invention are not limited thereto.
[0014] See Figure 1 The embodiment of the present invention provides a method for generating a radar mode and maneuver decision-making collaborative strategy in a simulated environment, which specifically includes the following steps: S10. Analyze the functional characteristics of different radar operating modes and design corresponding operating criteria constraints based on the analysis results.
[0015] The embodiment of the present invention analyzes the functional characteristics of different radar operating modes, utilizes the difference in detection data error to reflect the difference between the different operating modes, and designs corresponding operating criteria and constraints based on the radar operating modes. In the embodiment of the present invention, the radar operating modes include RWS (Range While Search) mode, TAS (Track And Search) mode, and STT (Single Target Tracking) mode. The functional characteristics of the three radar operating modes are analyzed accordingly, and corresponding operating criteria and constraints are designed based on the analysis results. The analysis includes: analyzing the functional characteristics of the three radar operating modes, and obtaining the analysis results that the detection accuracy corresponding to the RWS mode, TAS mode, and STT mode increases in sequence; the detection ranges corresponding to the RWS mode and STT mode are the same, and the detection range corresponding to the TAS mode is smaller than the detection range of the RWS mode; based on the analysis results, the corresponding operating criteria and constraints designed include: the execution time of the STT mode cannot exceed the maximum executable time; the RWS mode and STT mode, and the RWS mode and TAS mode can directly enter each other; executing an attack in the STT mode; and in the case of a single target, the resource consumption corresponding to the RWS mode, TAS mode, and STT mode increases in sequence. More specifically: First, based on the mission context of simulated aircraft confrontation, we selected three typical operating modes: RWS, TAS, and STT. Combining relevant expert experience and extensive documentation, we summarized the functional characteristics of different radar operating modes: detection accuracy is lowest in RWS mode, followed by TAS, and highest in STT. Detection range is the same in RWS and STT modes, but is reduced in TAS mode.
[0016] Secondly, by combining the simulated confrontation experience of relevant pilots and radar working principles, the three radar working modes were functionally modeled, that is, the working principle constraints were designed, which mainly include the following features: Figure 2 shown.
[0017] 1) In RWS mode, you can directly enter STT mode or TAS mode, but you need to ensure that a certain number of track points have been obtained.
[0018] 2) The STT mode execution time should not be too long and should be controlled within its maximum executable time. When in STT mode, the radar will block other targets. If it is in STT mode for a long time, it will lose the ability to predict simulated confrontation scenarios.
[0019] 3) Attack missions can only be performed in STT mode.
[0020] 4) The maximum detection range in RWS mode and STT mode is the same, but in TAS mode, the maximum detection range is smaller.
[0021] 5) In the case of single target, according to resource consumption, the three radar working modes can be sorted as follows: RWS mode <TAS模式<STT模式。
[0022] At the same time, to ensure that the simulation environment closely matches the actual aircraft simulation confrontation process, an early warning aircraft was also introduced. When the radar fails to intercept the target, information acquisition mainly relies on the detection of the early warning aircraft. The detection accuracy and effect are also modeled as Gaussian noise. In addition, the accuracy differences of different radar operating modes are mapped to Gaussian noise models with different variances. Finally, the Gaussian noise models are: , , , .
[0023] S20. Construct an attack reward function based on the three factors of angle, distance, and height, construct a radar reward function based on the working criteria constraints, and construct a total reward function based on the attack reward function and the radar reward function.
[0024] When attacking an enemy aircraft, the goal is to guide our aircraft to quickly meet the enemy aircraft, occupy a situational advantage, make the enemy aircraft target fall into our aircraft's attack zone, and make the attack zone as large as possible, taking into account the three factors of angle, distance, and height. Therefore, in an embodiment of the present invention, an attack reward function is constructed based on the consideration of the three factors of angle, distance, and height, including: constructing a target azimuth reward function and a target entry angle reward function based on the consideration of angle factors; constructing a distance reward function based on the consideration of distance factors; constructing an altitude reward function based on the consideration of altitude factors; constructing an attack reward function according to the target azimuth reward function, the target entry angle reward function, the distance reward function, and the altitude reward function. More specifically: The head-on situation requires that the target entry angle and azimuth angle simultaneously meet certain conditions. The reward model for the target azimuth angle is established based on the target attack zone, and the reward model for the target entry angle is established based on expert experience. The smaller the target azimuth angle, the better, and the larger the target entry angle, the better. The target azimuth angle reward function constructed in this embodiment of the present invention is expressed as follows: ; in, represents the target azimuth reward, represents the target azimuth, represents the maximum search azimuth, represents the maximum off-axis emission angle, represents the inescapable cone angle, ; The constructed goal enters the angular reward function, which is expressed as follows: ; in, represents the target entry corner reward, represents the target entry angle.
[0025] Furthermore, the higher the altitude, the larger the attack zone of the autonomous aircraft. However, due to the performance limitations of the aircraft and the fact that excessive altitude can lead to performance degradation of the autonomous aircraft, the concept of optimal aircraft simulation confrontation altitude is introduced. When the optimal altitude is exceeded, the probability of being attacked decreases and the reward value obtained decreases. The distance reward function constructed in this embodiment of the present invention is expressed as follows: ; in, Indicates distance reward, Indicates the distance between the carrier and the target. Indicates the maximum detection distance of the radar. 、 Respectively represent the minimum attack distance and maximum attack distance of the target. 、 They represent the minimum and maximum distances that cannot be escaped respectively; Furthermore, distance is the most intuitive indicator for measuring the aircraft's simulated confrontation situation. Here, a reward function is designed based on the attack zone and radar detection model. When the target enters the no-escape zone, the probability of hitting the target is maximized, and the reward is the highest. The height reward function constructed in this embodiment of the present invention is expressed as follows: ; in, Indicates high reward, Indicates the aircraft altitude. Indicates the target height, Indicates the optimal positioning height of the carrier aircraft; Furthermore, the attack reward function constructed in the embodiment of the present invention is expressed as follows: ; in, Indicates attack reward, 、 、 Respectively represent the importance weights of the three factors of angle, height and distance. 、 They represent the weights of azimuth and entry angle respectively.
[0026] Furthermore, to ensure the effectiveness of radar strategy execution while avoiding the introduction of a large amount of subjective experience, a radar strategy reward model was designed. When the radar detects a target and the current action meets the radar operating criteria, a positive reward is given. When an invalid detection is made, no reward is given. When the STT mode is executed for a long time, a penalty is imposed. The radar reward function constructed in this embodiment of the present invention is expressed as follows: ; in, Indicates radar reward, Indicates the mode type, Indicates RWS mode, Indicates TAS mode, Indicates STT mode, Indicates whether the target is within the radar detection range. Indicates that it is within the radar detection range. It is out of radar detection range. Indicates whether the radar working mode selection meets the working criteria constraints, Indicates that the radar working mode selection meets the working criteria constraints, Indicates that the radar working mode selection does not meet the working criteria constraints, Indicates the duration of STT mode execution. Indicates the maximum executable time in STT mode.
[0027] Furthermore, the embodiment of the present invention combines the attack reward function and the radar reward function to obtain the final total reward function, which is expressed as follows: .
[0028] in, Represents the total reward.
[0029] S30. In combination with the radar working mode and working criteria constraints, define a state space consisting of carrier state information, target state information, relative situation information, and sensor state information, and define a hybrid action space consisting of a maneuvering action space and a radar action space.
[0030] In the embodiment of the present invention, the carrier state information includes the carrier speed, the carrier speed change rate, the carrier pitch angle, the carrier pitch angle change rate, the carrier yaw angle, and the carrier yaw angle change rate; the target state information includes the target speed, the target speed change rate, the target pitch angle, the target pitch angle change rate, the target yaw angle, and the target yaw angle change rate; the relative situation information includes the horizontal relative distance between the target and the carrier, the vertical relative distance between the target and the carrier, the target entry angle, and the target azimuth; and the sensor state information includes the STT execution status and tracking stability. More specifically: The simulation environment studied in this embodiment of the present invention incorporates radar detection. Combining radar operating modes and operating criteria constraints, two sensor characteristic state quantities, STT execution state and tracking stability, are added to the situational characteristic state space. Ultimately, the state space defined in this embodiment of the present invention consists of four components. This component comprehensively reflects the spatial maneuverability advantage through the state information of the carrier and target, as well as relative situational information. The current radar execution status is determined based on sensor characteristic quantities, thereby providing more comprehensive indicative information to the agent. Specific definitions are as follows: ; ; ; ; ; in, Indicates the status information of the carrier, including the speed of the carrier , the speed change rate of the carrier aircraft , the pitch angle of the carrier aircraft and the pitch angle change rate of the carrier aircraft , the yaw angle of the carrier aircraft and the yaw angle change rate of the carrier aircraft ; Indicates target status information, including the target's speed , target speed change rate , the target's pitch angle and the target's pitch angle change rate , the target's yaw angle and the target's yaw angle change rate ; Relative situation information, including the horizontal relative distance between the target and the carrier aircraft and vertical relative distance , target entry angle and target azimuth ; Indicates sensor status information, including STT execution status ( ), tracking stability , respectively count the STT continuous execution time and the number of stable tracking times; Represents a state space defined by four parts, Represents vector merging.
[0031] The maneuvering action space of the embodiment of the present invention includes the overload information of the carrier aircraft in the tangential, lateral and normal directions respectively; the radar action space includes the execution probabilities corresponding to the RWS mode, TAS mode and STT mode respectively. More specifically: Compared with the discrete action space, the continuous action space of the overload control model established by the embodiment of the application is more in line with the real aircraft simulation confrontation process, and ensures the generation of complex strategy behavior. At the same time, the discrete radar working mode switching action is continuous in the form of execution probability, which is also more in line with the decision logic of the pilot, so as to finally build a six-dimensional mixed action space, which is expressed as: ; Among them, 、 、 represent the overload information of the carrier aircraft in the tangential direction, the lateral direction and the normal direction respectively; 、 、 represent the execution probabilities of the RWS mode, the TAS mode and the STT mode respectively.
[0032] S40, a deep reinforcement learning algorithm framework based on EDN-PPOA is designed, based on the total reward function, the state space and the mixed action space, the network parameters are adjusted to output the mixed action strategy model corresponding to the optimal network parameters; wherein the deep reinforcement learning algorithm framework based on EDN-PPOA designed is: on the basis of the traditional PPO algorithm, the advantage function combining the GAE and the entropy of the strategy distribution is adopted in the Critic network processing, the Dueling network is introduced in the Actor network processing to decouple the mixed action space and the noise network is introduced to improve the robustness of the model, and the total target optimization function is constructed based on the Critic network processing and the Actor network processing.
[0033] Based on the total reward function constructed by S20, and the state space and the mixed action space defined by S30, the embodiment of the application designs a deep reinforcement learning algorithm framework based on EDN-PPOA as shown in Figure 3 to generate a mixed action strategy. Specifically: The Critic network gives reward feedback according to the environment, iteratively optimizes the state value function, so as to more accurately evaluate the advantages and disadvantages of the current strategy. According to the characteristics of the PPO algorithm, the advantage function is used as the direction of gradient update, so that the variance is reduced and the convergence is more stable. The Critic network is updated based on the time difference, which is specifically as follows: ; Among them, represents the overall parameters of the Critic network, represents the target function of the Critic network, represents the state space at the moment , which includes 、 、 、 , Indicates The estimated value of express The target value of the moment, express Reward of the moment.
[0034] Then the advantage function can be defined as: ; ; in, express The advantage value at time, represents the discount factor, The value is ~T-1, T represents the selected time length, express The reward value at that moment, express The reward value calculated according to the total reward function at each moment, express The state space at time, Indicates The output strategy is: Representation Strategy The entropy of the distribution, Represents the temperature factor.
[0035] To achieve more accurate evaluation and reduce variance, the embodiment of the present invention adopts the generalized advantage estimation (GAE) method. In this case, the form of the advantage function can be changed to: ; in, represents the smoothing parameter, , express The state space at time, Indicates The estimated value of Indicates It can be seen that the embodiment of the present invention adopts the advantage function combining GAE and the entropy of policy distribution in the critic network processing.
[0036] Furthermore, the Actor network, also known as the policy network, learns through policy gradients and ultimately obtains the optimal action distribution. In order to achieve the decoupling of the radar working mode switching action and the overload control instruction action in the mixed action space, the embodiment of the present invention introduces the Dueling network architecture. This architecture uses a shared feature extraction layer and an independent action output layer to accurately establish a flexible nonlinear mapping from situation to action, avoiding mutual interference between mixed actions, thereby achieving joint judgment and independent decision-making. The decision output of the Actor network introduced with the Dueling network architecture is expressed as follows: ; in, Represents the overall network parameters of the Actor network, , represents the network parameters of the feature extraction layer, Represents the network parameters of the action output layer corresponding to the radar action, represents the network parameters of the action output layer corresponding to the maneuver action, express After introducing the Actor network of the Dueling network, the maneuvering action and radar action to be executed are selected. represents the radar action space, , express After passing through the feature extraction layer and the action output layer corresponding to the radar action, Select the radar action to be executed. represents the maneuvering action space, express After passing through the feature extraction layer and the action output layer corresponding to the maneuver, Choose the maneuver to perform, .
[0037] Furthermore, the Dueling network introduced in the Actor network processing in the embodiment of the present invention includes a shared feature extraction layer and an independent action output layer. Different action output layers are used to select the maneuvering action to be executed from the maneuvering action space and the radar action to be executed from the radar action space. The maneuvering action and the radar action are then output after adding noise through different noise networks. More specifically: To address the problem of high volatility in policy learning caused by noise disturbances, a noise network is introduced in the action output layer. By adding noise to the network weights, the anti-interference ability of the policy network is effectively improved, while further increasing the search range and achieving large-scale generalization. The parameter form of the action output layer can be expressed as and , at this time the output strategy of the Actor network can be expressed as: ; in, represents element-wise multiplication, represents the noise weight scale parameter, represents random weight noise, represents the deterministic bias term of the weight, represents the noise bias scale parameter, represents random bias noise, represents the deterministic bias term of the bias, i.e. The noise follows a normal distribution.
[0038] The final output strategy of the Actor network is expressed as follows: ; in, represents the overall parameters of the noise network, represents the noise parameters of the noise network corresponding to the radar action, , represents the noise parameters of the noise network corresponding to the maneuver, express The state space at time, represents the radar action space, express After passing through the feature extraction layer, the action output layer corresponding to the radar action, and the noise network corresponding to the radar action, Select the radar action to be executed. represents the maneuvering action space, express After passing through the feature extraction layer, the action output layer corresponding to the maneuver action, and the noise network corresponding to the maneuver action, Choose the maneuver to perform, express Select the maneuvers and radar actions to be executed at any time. express After introducing the Actor network of the noise network and the Dueling network, the maneuvering action and radar action to be executed are selected.
[0039] Furthermore, based on the traditional PPO algorithm, the embodiment of the present invention constructs an overall objective optimization function based on the Critic network processing and the Actor network processing, which is expressed as follows: ; in, represents the overall objective optimization function, Represents the overall network parameters of the Actor network, Indicates the overall network parameters after the Actor network is updated. represents the noise parameters of the noise network, express The state space at time, express Select the maneuvers and radar actions to be executed at any time. express After introducing the Actor network of the noise network and the Dueling network, the maneuvering action and radar action to be executed are selected. Indicates that after updating the overall network parameters of the Actor network, After introducing the Actor network of the noise network and the Dueling network, the maneuvering action and radar action to be executed are selected. represents the pruning function, Represents the cropping parameters, express The advantage value at time, Indicates The output strategy is: Representation Strategy The entropy of the distribution, Represents the temperature factor.
[0040] Furthermore, the embodiment of the present invention is based on the established simulation confrontation simulation environment, simulates and interacts with the network model, and optimizes the network performance by adjusting the network parameters to output a hybrid action strategy model corresponding to the optimal network parameters. The hybrid action strategy model is Figure 3 The overall network framework shown includes the Critic network, Actor network, noise network and Dueling network.
[0041] S50: Input the current state space into the hybrid action strategy model and output the corresponding hybrid action space to achieve autonomous decision-making for maneuvering and radar operating mode selection. It can be seen that the embodiment of the present invention not only achieves decision-making for maneuvering, but also for radar operating mode.
[0042] In order to verify the effectiveness of the radar mode and maneuver decision-making collaborative strategy generation method in a simulated environment provided by an embodiment of the present invention, the following experiments were conducted for verification.
[0043] All simulation experiments in this embodiment of the present invention were compiled and developed using Python 3.8 and PyCharm 2023.2 on a Windows 10 PC with an Intel i7 processor and 32GB of RAM. To ensure the opponent's flexibility, this embodiment of the present invention sets the target to adopt a horizontal S-escape strategy. Furthermore, to ensure the robustness of the algorithm and excellent performance, the initial training situation is unfavorable. The carrier and target aircraft models in the experiment are identical, with the specific relevant parameters shown in Table 1.
[0044] Table 1 Aircraft model related parameters
[0045] To validate the effectiveness of our proposed method, we compared it with cutting-edge algorithms such as SAC (Soft Actor-Critic) and DDPG (Deep Deterministic Policy Gradient). We also compared it with the advanced Dueling-Noisy net-Multi-steps (DNM-DQN) algorithm in discrete space, further demonstrating the policy flexibility and potential for improvement in continuous action spaces. The specific hyperparameter settings are shown in Table 2.
[0046] Table 2 Hyperparameter settings of each comparison algorithm
[0047] During the attack intention strategy training process, the target takes an S-shaped maneuver to accelerate and escape, and the initial situation is set to an unfavorable tail pursuit scenario to improve the adaptability of the strategy. The reward change comparison results of each algorithm during the training process are shown in the figure below. Figure 4 At the same time, a schematic diagram of the comparison results of the cumulative reward mean and variance for the last 50 rounds is also shown, as shown in Figure 5(a) to Figure 5(b) As shown, Figure 5(a) shows the comparison results of the cumulative reward mean, and Figure 5(b) shows the comparison results of the cumulative reward variance.
[0048] from Figure 4 It can be seen that the method proposed in this invention is significantly better than other advanced algorithms in terms of performance, and its convergence speed is also significantly faster than SAC and DDPG. The main reason for DDPG's weak performance is that the observation information in the adversarial environment is disturbed, which seriously affects the convergence of the deterministic strategy. At the same time, SAC is also highly volatile. This phenomenon also reflects the effectiveness of the method proposed in this invention from the side, which can solve the problems of instability and difficulty in convergence caused by noise. In addition, Figure 5(a) to Figure 5(b) This demonstrates that the proposed method can significantly reduce variance while effectively improving performance, thereby enhancing the stability of the algorithm. Although DNM-DQN has lower variance, the main reason is that the discrete action space limits the richness of the strategy, causing it to converge to a suboptimal solution.
[0049] In addition, the present invention conducted ablation experiments to remove four modules: the noise network (corresponding to N in EDN-PPOA), entropy regularization (corresponding to E in EDN-PPOA), the Dueling network (corresponding to D in EDN-PPOA), and the innovative total reward function design (corresponding to A in EDN-PPOA), and compared them with the traditional PPO algorithm.
[0050] Taking the training model with a random number seed of 10 as an example, the present invention counts the cumulative rewards of 100 rounds under the random initial state, and analyzes and illustrates the distribution of the cumulative rewards through box plots and violin plots. The results are as follows: Figure 6(a) to Figure 6(b) As shown, Figure 6(a) shows a violin plot and Figure 6(b) shows a box plot.
[0051] Violin plots and box plots are commonly used data visualization methods that can intuitively display the distribution characteristics and feature information of the data. Among them, the box plot mainly shows the distribution characteristics of the data, while the violin plot reflects the distribution and density information of the data on its basis. As can be seen from Figure 6(a), the distribution of EDN-PPOA is more concentrated, and the median in Figure 6(b) is also higher than that of other algorithms. In addition, after removing the innovative total reward function and the Dueling network, the performance of the algorithm will drop significantly, which indirectly reflects the role of the radar working mode in maneuver decision-making and the impact on the aircraft simulation confrontation process. The radar strategy based on expert experience lacks coordination with maneuvering, so the effect is weaker than the method model proposed in this invention. Figure 6(a) to Figure 6(b) It is clearly shown that under random situation, the method proposed in this invention has stronger robustness and better performance.
[0052] In addition, the present invention also designs three indicators: radar use success rate, situation cumulative error, and resource consumption to verify the excellent performance of the method proposed in the present invention. The radar use success rate refers to the proportion of the radar working mode currently executed during the confrontation process that meets the working conditions (including use conditions and whether it is within the detection range). For example, when the RWS mode does not accumulate enough track points or the target is not within the corresponding detection range, the strategy output STT mode or TAS mode is an invalid result. This indicator is used to judge the execution success rate of the radar strategy and whether it can accurately switch to the correct working mode according to the current situation. Since different radar working modes consume different resources, the present invention is simplified according to multiple factors such as radar power. Taking the RWS mode as the benchmark, the benchmark value is 1, and the resources consumed in the TAS mode and STT mode correspond to 2 and 3 respectively. The radar resource usage is based on the radar use success rate to calculate the total resource consumption during the confrontation process. The situation cumulative error mainly refers to the sum of the position errors caused by the corresponding radar working mode and detection situation at each moment during the confrontation process.
[0053] Finally, the cumulative reward mean (denoted as Reward), radar use success rate (denoted as Radar_success), cumulative error (denoted as Accumulate_error), and resource consumption (denoted as Resource) corresponding to all algorithms in 100 rounds under random situations are statistically analyzed, as shown in Table 3.
[0054] Table 3 Results of different indicators corresponding to different algorithms under random attack intention strategy
[0055] Table 3 shows that radar usage strategy affects maneuvering strategy and even the results of simulated aircraft confrontations. The proposed method generates radar strategies that surpass expert experience, not only flexibly selecting radar operating modes based on situational awareness but also effectively avoiding excessive waste of radar resources. Furthermore, the innovative total reward function design significantly improves the radar usage success rate and avoids ineffective action outputs. When this component is missing, the radar usage success rate significantly decreases. Furthermore, based on the results of other algorithms, new insights into radar usage strategies during simulated aircraft confrontations have emerged: excessive pursuit of precise situational information acquisition can, to a certain extent, sacrifice the flexibility and advantages of strategic maneuvers, thereby impacting the entire confrontation process.
[0056] Finally, by intercepting the trajectory fragments of the confrontation process, the action output results at a certain moment are selected for visualization, and then the causal relationship between the situation and the action and the rationality of the action output are analyzed, such as Figure 7 shown. Figure 7 In the paper, the method proposed in the present invention is compared with the Without Noisy algorithm and the Fix policy algorithm. It can be seen that the strategy based on expert experience is too rigid in terms of switching timing and switching type. When the target makes complex maneuvers and quickly increases the distance, it is easy to switch to the wrong mode, resulting in the loss of the target. The Without Noisy algorithm even mistakenly enters the TAS mode in the early stage of the confrontation. In comparison, the method proposed in the present invention has a greater advantage in target detection, and can reserve sufficient maneuvering space for the target, ensuring that the target is always within the detection range. This also further confirms that the method proposed in the present invention achieves the synergy between maneuvering decision-making and radar strategy.
[0057] In summary, the radar mode and maneuver decision-making collaborative strategy generation method proposed in the embodiment of the present invention in a simulated environment provides a reasonable and feasible solution for achieving efficient strategic coordination of radar strategy and maneuver strategy during confrontation. Specifically, the functional characteristics of different radar operating modes are analyzed, the corresponding working criteria constraints are designed, and a total reward function including attack reward function and radar reward function is designed to reduce the ineffective actions of the intelligent agent. Then, the appropriate state space and mixed action space are redefined according to the strategy generation requirements. Finally, on this basis, a deep reinforcement learning algorithm framework based on EDN-PPOA is proposed. By introducing the concept of entropy and noise network, the robustness and exploration ability of the network model are improved. At the same time, the coupling relationship between radar strategy and maneuver strategy is considered, and the strategy decoupling of the mixed action space is carried out by combining the Dueling network framework. Simulation experimental results show that the proposed method realizes the generation of mixed action collaborative strategy of radar operating mode and maneuver decision, and can also achieve efficient coordination of radar and maneuver in complex mixed action space. The strategy performance is better than traditional empirical methods. At the same time, the proposed method shows significant strategy stability and robustness in observation noise environment, providing important technical support for intelligent decision-making in simulated aircraft confrontation.
[0058] In the description of the present invention, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature specified as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, "plurality" means two or more, unless otherwise specifically defined.
[0059] Although the present invention is described herein in conjunction with various embodiments, those skilled in the art may understand and implement other variations of the disclosed embodiments by reviewing the specification and accompanying drawings in the process of implementing the claimed invention. In the specification, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude multiple components or steps. The fact that certain measures are described in different embodiments does not mean that these measures cannot be combined to produce good results.
[0060] The above is a further detailed description of the present invention in conjunction with specific preferred embodiments, and the specific implementation of the present invention should not be considered to be limited to these descriptions. For those skilled in the art of the present invention, without departing from the concept of the present invention, several simple deductions or substitutions can be made, which should be considered to fall within the scope of protection of the present invention.
Claims
1. A method for generating a radar pattern and maneuver decision-making collaborative strategy in a simulated environment, characterized in that: The method comprises: Analyze the functional characteristics of different radar operating modes and design corresponding operating criteria constraints based on the analysis results; Constructing an attack reward function based on the three factors of angle, distance, and height, constructing a radar reward function based on the working criteria constraints, and constructing a total reward function based on the attack reward function and the radar reward function; Combined with the radar working mode and working criteria constraints, a state space consisting of aircraft state information, target state information, relative situation information, and sensor state information is defined, and a hybrid action space consisting of maneuver action space and radar action space is defined; A deep reinforcement learning algorithm framework based on EDN-PPOA is designed. Based on the total reward function, the state space, and the mixed action space, the network parameters are adjusted to output a mixed action policy model corresponding to the optimal network parameters. The designed deep reinforcement learning algorithm framework based on EDN-PPOA is as follows: on the basis of the traditional PPO algorithm, an advantage function combining GAE and the entropy of the policy distribution is adopted in the critic network processing, a Dueling network is introduced in the actor network processing to decouple the mixed action space policy and a noise network is introduced to improve the robustness of the model, and an overall objective optimization function is constructed based on the critic network processing and the actor network processing; The current state space is input into the hybrid action strategy model, and the corresponding hybrid action space is output to achieve autonomous decision-making on maneuvering and radar working mode selection.
2. The method for generating a radar mode and maneuver decision-making collaborative strategy in a simulation environment according to claim 1, characterized in that: The radar operating modes include RWS mode, TAS mode and STT mode; The functional characteristics of the three radar operating modes are analyzed accordingly, and corresponding operating criteria constraints are designed based on the analysis results, including: Analysis of the functional characteristics of the three radar operating modes revealed that the detection accuracy of the RWS, TAS, and STT modes increases in sequence, and that the detection ranges of the RWS and STT modes are the same, while the detection range of the TAS mode is smaller than that of the RWS mode. Based on the analysis results, the corresponding working criteria constraints of the design include: The execution time of STT mode cannot exceed the maximum executable time; There is direct access between RWS mode and STT mode, and between RWS mode and TAS mode; Execute the attack in STT mode; In the single-target case, the resource consumption corresponding to the RWS mode, TAS mode, and STT mode increases in turn.
3. The method for generating a radar mode and maneuver decision-making collaborative strategy in a simulation environment according to claim 1, characterized in that: The attack reward function is constructed based on three factors: angle, distance, and height, including: Based on the consideration of angle factors, the target azimuth angle reward function and the target entry angle reward function are constructed; Construct a distance reward function based on distance considerations; Construct a height reward function based on height considerations; An attack reward function is constructed according to the target azimuth reward function, the target entry angle reward function, the distance reward function, and the height reward function.
4. The method for generating a radar mode and maneuver decision-making collaborative strategy in a simulation environment according to claim 3, characterized in that: The target azimuth reward function is constructed as follows: ; in, represents the target azimuth reward, represents the target azimuth, Indicates the maximum search azimuth, represents the maximum off-axis emission angle, represents the inescapable cone angle, ; The constructed goal enters the angular reward function, which is expressed as follows: ; in, represents the target entry corner reward, represents the target entry angle; The distance reward function constructed is expressed as follows: ; in, Indicates distance reward, Indicates the distance between the carrier and the target. Indicates the maximum detection distance of the radar. 、 Respectively represent the minimum attack distance and maximum attack distance of the target. 、 They represent the minimum and maximum distances that cannot be escaped respectively; The constructed height reward function is expressed as follows: ; in, Indicates high reward, Indicates the aircraft altitude. Indicates the target height, Indicates the optimal positioning height of the carrier aircraft; The constructed attack reward function is expressed as follows: ; in, Indicates attack reward, 、 、 Respectively represent the importance weights of the three factors of angle, height and distance. 、 They represent the weights of azimuth and entry angle respectively.
5. The method for generating a radar mode and maneuver decision-making collaborative strategy in a simulation environment according to claim 1, characterized in that: The constructed radar reward function is expressed as follows: ; in, Indicates radar reward, Indicates the mode type, Indicates RWS mode, Indicates TAS mode, Indicates STT mode, Indicates whether the target is within the radar detection range. Indicates whether the radar working mode selection meets the working criteria constraints, Indicates the duration of STT mode execution. Indicates the maximum executable time in STT mode.
6. The method for generating a radar mode and maneuver decision-making coordination strategy in a simulation environment according to claim 2, characterized in that: The carrier aircraft status information includes the carrier aircraft's speed, the carrier aircraft's speed change rate, the carrier aircraft's pitch angle, the carrier aircraft's pitch angle change rate, the carrier aircraft's yaw angle, and the carrier aircraft's yaw angle change rate; the target state information includes the target's speed, the target's speed change rate, the target's pitch angle, the target's pitch angle change rate, the target's yaw angle, and the target's yaw angle change rate; the relative situation information includes the horizontal relative distance between the target and the carrier aircraft, the vertical relative distance between the target and the carrier aircraft, the target entry angle, and the target azimuth; the sensor status information includes the STT execution status and tracking stability; The maneuvering action space includes the overload information of the carrier aircraft in the tangential, lateral and normal directions respectively; the radar action space includes the execution probabilities corresponding to the RWS mode, TAS mode and STT mode respectively.
7. The method for generating a radar mode and maneuver decision-making collaborative strategy in a simulation environment according to claim 1, characterized in that: The advantage function of the combination of GAE and entropy of policy distribution used in critic network processing is expressed as follows: ; in, express The advantage value at time, represents the discount factor, represents the smoothing parameter, , The value is ~T-1, T represents the selected time length, Represents the overall parameters of the Critic network, , express The reward value at that moment, express The reward value calculated according to the total reward function at each moment, express The state space at time, express The state space at time, Indicates The estimated value of Indicates The estimated value of Indicates The output strategy is: Representation Strategy The entropy of the distribution, Represents the temperature factor.
8. The method for generating a radar mode and maneuver decision-making coordination strategy in a simulation environment according to claim 1, characterized in that: The Dueling network introduced in the Actor network processing includes a shared feature extraction layer and an independent action output layer, so as to select a maneuvering action to be executed from the maneuvering action space and a radar action to be executed from the radar action space through different action output layers; and the maneuvering action and the radar action are output after adding noise through different noise networks.
9. The method for generating radar mode and maneuver decision-making coordination strategy in a simulation environment according to claim 8, characterized in that: The output strategy through the Actor network is expressed as follows: ; in, Represents the overall network parameters of the Actor network, , represents the network parameters of the feature extraction layer, Represents the network parameters of the action output layer corresponding to the radar action, represents the network parameters of the action output layer corresponding to the maneuver action, represents the overall parameters of the noise network, represents the noise parameters of the noise network corresponding to the radar action, , represents the noise parameters of the noise network corresponding to the maneuver, express The state space at time, represents the radar action space, express After passing through the feature extraction layer, the action output layer corresponding to the radar action, and the noise network corresponding to the radar action, Select the radar action to be executed. represents the maneuvering action space, express After passing through the feature extraction layer, the action output layer corresponding to the maneuver action, and the noise network corresponding to the maneuver action, Choose the maneuver to perform, express Select the maneuvers and radar actions to be executed at any time. express After introducing the Actor network with the noise network and the Dueling network, the maneuvering and radar actions to be executed are selected.
10. The method for generating a radar mode and maneuver decision-making collaborative strategy in a simulation environment according to claim 1, characterized in that: The overall objective optimization function constructed based on Critic network processing and Actor network processing is expressed as follows: ; in, represents the overall objective optimization function, Represents the overall network parameters of the Actor network, Indicates the overall network parameters after the Actor network is updated. represents the noise parameters of the noise network, express The state space at time, express Select the maneuvers and radar actions to be executed at any time. express After introducing the Actor network of the noise network and the Dueling network, the maneuvering action and radar action to be executed are selected. Indicates that after updating the overall network parameters of the Actor network, After introducing the Actor network of the noise network and the Dueling network, the maneuvering action and radar action to be executed are selected. represents the pruning function, Represents the cropping parameters, express The advantage value at time, Indicates The output strategy is: Representation Strategy The entropy of the distribution, Represents the temperature factor.