A loyal wingman target search and lock mission execution method
Patent Information
- Application Number
- CN202410070271.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-17
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2044-01-17
AI Technical Summary
[0003]为了避免现有技术的不足之处,本申请提供一种忠诚僚机目标搜索与锁定任务执行方法,用以解决现有技术中存在面临任务复杂度高、学习困难、执行特定任务时的稳定性和鲁棒性不足,如在目标锁定过程中一旦目标丢失,算法即会停止,必须切换到搜索模式;在接近探测区域边缘时,算法频繁切换,导致性能下降等问题的问题
[0065]本公开的实施例中,通过上述忠诚僚机目标搜索与锁定任务执行方法,一方面,底层控制模型负责控制飞机的油门杆、方向舵、升降舵和副翼来控制飞机飞行,达到期望航向、速度以及高度;先验区域模型利用先验化训练,去先验化执行的方法进行运算;顶层目标搜索任务模型基于底层控制模型构建,无需从头训练如何控制飞机飞行,且可迁移至不同的顶层任务,大大简化了任务训练复杂度;顶层目标锁定任务模型能够在一定范围内自动寻回丢失目标,降低了模型切换频率,且提升了控制系统稳定性。另一方面,该方法能够将复杂的忠诚僚机训练任务分而治之,缩短训练时间,且提供了高效的自主目标搜索与目标搜索方法,解决了算法模型切换时的不稳定性问题。
Smart Images

Figure CN117991806B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of loyal wingman maneuver control and autonomous mission execution technology, and in particular to a method for loyal wingman target search and locking mission execution. Background Technology
[0002] In modern air combat, the coordinated operation between lead and wingmen is crucial, and this formation is widely used in actual air combat due to its numerous advantages. Currently, promising intelligent air combat algorithms can be divided into two categories: those based on traditional intelligent algorithms and those based on reinforcement learning algorithms. With the continuous improvement of computing power in recent years, research on reinforcement learning has experienced explosive growth, showing immense potential. However, existing reinforcement learning algorithms face challenges such as high task complexity, learning difficulties, and insufficient stability and robustness when performing specific tasks. For example, if the target is lost during target locking, the algorithm will stop and must switch to search mode. Furthermore, frequent switching near the edge of the detection area leads to performance degradation. Summary of the Invention
[0003] To avoid the shortcomings of existing technologies, this application provides a method for performing a loyal wingman target search and locking task, which addresses the problems of high task complexity, learning difficulties, insufficient stability and robustness when performing specific tasks, such as the algorithm stopping and having to switch to search mode if the target is lost during the target locking process; and the frequent switching of the algorithm when approaching the edge of the detection area, leading to performance degradation.
[0004] According to an embodiment of this disclosure, a method for performing a loyal wingman target search and locking mission is provided, the method comprising:
[0005] Based on environmental complexity and combat mission, a low-level control model is constructed, and flight control is performed using the low-level control model;
[0006] Based on the radar detection area, a priori region model is constructed by introducing range ratio and angle ratio; the priori region model includes a priori region sub-model, a safe region sub-model, a danger region sub-model, and a termination region sub-model.
[0007] Based on the underlying control model, a top-level target search task model is constructed, and the target search is performed using the top-level target search task model.
[0008] Based on the underlying control model and the prior area model, a top-level target locking task model is constructed, and the target locking is performed using the top-level target locking task model.
[0009] Furthermore, the steps in constructing the underlying control model include:
[0010] Based on the environmental complexity and combat mission, several state variables are designed, and a low-level control model is constructed based on these state variables; the expression of the low-level control model is as follows:
[0011] [ΔH,Δψ,ΔV,H abs ,cosφ,sinφ,sinθ,cosθ,V x V y V z ,V]
[0012] In the formula, ΔH is the difference between the desired altitude and the wingman's current altitude, Δψ is the difference between the desired heading and the wingman's current heading, ΔV is the difference between the desired speed and the wingman's current speed, and H... abs The wingman's current altitude is given by φ, roll angle is given by θ, and pitch angle is given by V. x V represents the wingman's x-axis velocity. y V represents the y-axis velocity of the wingman. z Let V be the z-axis velocity of the wingman, and V be the resultant velocity of the wingman.
[0013] The action space of the underlying control model consists of a quadruple of four consecutive control variables:
[0014] [C a C e C r C t ]
[0015] In the formula, C a C is the throttle lever control quantity used to control the speed of the wingman; e C is the rudder control variable used to control the wingman's yaw angle; r C is the elevator control parameter used to control the pitch angle of the wingman; t Aileron control parameters are used to control the roll angle of the wingman.
[0016] Discretize the continuous action space to balance computational complexity and simulation accuracy.
[0017] Furthermore, the steps to discretize the continuous action space include:
[0018] Define the reward function:
[0019]
[0020] In the formula, R ψ R is the heading reward function. H For a high reward function, R V R is the speed reward function. φ For the rolling reward function;
[0021] The overall control reward function is calculated based on the heading reward function, altitude reward function, speed reward function, and roll reward function, and is used as the overall metric; the overall control reward function is expressed as:
[0022]
[0023] Set a height penalty constant R PH According to the height penalty constant R PH And the overall reward of the wingman's underlying control model calculated by overall metrics:
[0024] R1 = R C +R PH
[0025] In the formula, R1 is the overall reward of the wingman's underlying control model.
[0026] Furthermore, based on the characteristics of the task, the termination conditions of the underlying control model are defined.
[0027] Furthermore, the range ratio is the ratio of the target's distance within the area to the radar's detection radius, with a range of [0, 1.5]. The angle ratio is the ratio of the angle between the target and the wingman's velocity vector to the radius of the radar's detection area, with a range of [-1.5, 1.5]. The angle ratio is signed; a positive sign indicates that the target is to the left of the wingman's velocity vector, and a negative sign indicates that the target is to the right of the wingman's velocity vector.
[0028] Furthermore, the steps for target search using the top-level target search task model include:
[0029] The top-level target search task model processes the coordinates of the top-level target waypoints, converting them into the desired heading, speed, and altitude of the loyal wingman, and then subtracts these from the current heading, speed, and altitude of the main aircraft.
[0030] Converting latitude and longitude to the navigation coordinate system, the waypoint coordinates are obtained as (x i ,y i ,z i If the coordinates of the loyal wingman are (x, y, z), then the height difference in the vertical plane is:
[0031] Δh=z i -z
[0032] On the horizontal plane, the position vector of the target waypoint relative to the loyal wingman is obtained as follows:
[0033]
[0034] The magnitude of the heading angle difference can then be obtained as:
[0035]
[0036] In the formula, The velocity vector on the horizontal plane;
[0037] Determine the sign of the angle difference:
[0038]
[0039] The difference in heading angle is used as the input to the underlying control model to complete the waypoint-based target search task.
[0040] Furthermore, based on the underlying control model and the prior region model, a top-level target locking task model is constructed. The steps for target locking using the top-level target locking task model include:
[0041] Based on the underlying control model and the prior region model, a top-level target locking task model is constructed:
[0042] [ΔH,Δψ,ΔV x H abs ,cosφ,sinφ,sinθ,cosθ,V x V y V z ,V,R,ψ r ,F A ,F D ]
[0043] In the formula, R is the relative distance between the target enemy aircraft and the wingman, and ψ r Let F be the angle between the target enemy aircraft's velocity vector and its line of sight. A For the angle ratio, F D It is the distance ratio;
[0044] The action space of the top-level target locking control model consists of a triplet of three discrete error quantities, which controls the wingman's heading, speed, and altitude through five control levels:
[0045] [ΔH d ,Δψ d ,ΔV d ]
[0046] In the formula, ΔH d Let Δψ be the height error. d Let ΔV be the heading error. d This is the speed error amount;
[0047] Penalties for sub-models in different regions are calculated to obtain penalties for lost areas, safe areas, danger zones, and termination zones; among them,
[0048] The penalty for lost areas is:
[0049]
[0050] In the formula, P DL As a guidance penalty for distance loss, P AL As a penalty for lost angle guidance, P DAL A penalty for losing both distance and angle;
[0051] The safe zone penalty is:
[0052] P safe =0
[0053] In the formula, P safe Punishment in the safe zone;
[0054] The penalty for dangerous zones is:
[0055]
[0056] In the formula, P danger Punishment for dangerous zones;
[0057] The penalty for termination is:
[0058] P terminate =-5
[0059] In the formula, P terminate Penalty for termination zone;
[0060] The reward given for a target remaining within the detection area regardless of whether it is in the safe zone or the danger zone:
[0061] R detect =1
[0062] In the formula, R detect Rewards will be given for the exploration area;
[0063] The top-level target locking task model locks the target based on the penalties for lost zones, safe zones, danger zones, termination zones, and exploration zones.
[0064] The technical solutions provided by the embodiments of this disclosure may include the following beneficial effects:
[0065] In the embodiments of this disclosure, the aforementioned loyal wingman target search and locking task execution method, on the one hand, involves a low-level control model responsible for controlling the aircraft's throttle, rudder, elevator, and ailerons to control flight and achieve the desired heading, speed, and altitude; a priori region model utilizes priori training and de-priori execution methods for computation; a top-level target search task model is built upon the low-level control model, eliminating the need for retraining on how to control aircraft flight and allowing transfer to different top-level tasks, significantly simplifying task training complexity; and a top-level target locking task model can automatically retrieve lost targets within a certain range, reducing model switching frequency and improving control system stability. On the other hand, this method can divide and conquer complex loyal wingman training tasks, shortening training time, and provides efficient autonomous target search and target search methods, solving the instability problem during algorithm model switching. Attached Figure Description
[0066] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure. It is obvious that the drawings described below are merely some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.
[0067] Figure 1 This diagram illustrates the steps of a method for performing a loyal wingman target search and locking mission according to an exemplary embodiment of this disclosure.
[0068] Figure 2 This diagram illustrates a six-degree-of-freedom aircraft model in an exemplary embodiment of this disclosure.
[0069] Figure 3 This diagram illustrates a hierarchical control architecture in an exemplary embodiment of this disclosure.
[0070] Figure 4 The diagram illustrates the reward curves for the training rounds of the underlying model of the method and the PPO algorithm in an exemplary embodiment of this disclosure.
[0071] Figure 5 This diagram illustrates a priori region model in an exemplary embodiment of this disclosure.
[0072] Figure 6 This diagram illustrates a waypoint target search mission for a leader-chief formation in an exemplary embodiment of this disclosure.
[0073] Figure 7 This diagram shows the actual simulation trajectory of the top-level target search task model in an exemplary embodiment of this disclosure;
[0074] Figure 8This diagram illustrates the training round reward curve of the top-level target locking task model in an exemplary embodiment of this disclosure.
[0075] Figure 9 This illustration shows a side view of the target triangular maneuver and wingman target locking mission execution in an exemplary embodiment of this disclosure.
[0076] Figure 10 This illustration shows a top view of the target triangular maneuver and wingman target locking mission execution in an exemplary embodiment of this disclosure.
[0077] Figure 11 This illustration shows a side view of the target head-on maneuver and wingman target locking mission execution in an exemplary embodiment of this disclosure.
[0078] Figure 12 This diagram shows a top view of a target head-on maneuver and a wingman target locking mission being executed in an exemplary embodiment of this disclosure.
[0079] Figure 13 This illustration shows a side view of the target escape maneuver and wingman target locking mission execution in an exemplary embodiment of this disclosure.
[0080] Figure 14 This illustration shows a top view of a target escape maneuver and wingman target locking mission execution in an exemplary embodiment of this disclosure.
[0081] Figure 15 A comparison chart showing the switching frequency of algorithm models with a priori regions in exemplary embodiments of this disclosure is presented. Detailed Implementation
[0082] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided so that this disclosure will be more comprehensive and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0083] Furthermore, the accompanying drawings are merely illustrative diagrams of embodiments of this disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities.
[0084] This example implementation provides a method for performing a loyal wingman target search and locking mission. (See reference...) Figure 1 As shown, the method for performing the loyal wingman target search and lock task may include steps S101 to S104.
[0085] Step S101: Based on the environmental complexity and combat mission, construct a low-level control model and use the low-level control model for flight control;
[0086] Step S102: Based on the radar detection area, introduce the range ratio and angle ratio to construct a priori region model; wherein, the priori region model includes a priori region sub-model, a safe region sub-model, a danger region sub-model and a termination region sub-model;
[0087] Step S103: Based on the underlying control model, construct the top-level target search task model, and use the top-level target search task model to perform target search;
[0088] Step S104: Based on the underlying control model and the prior region model, construct the top-level target locking task model, and use the top-level target locking task model to lock the target.
[0089] The aforementioned method for loyal wingman target search and locking tasks achieves several key benefits. Firstly, the underlying control model controls the aircraft's throttle, rudder, elevator, and ailerons to achieve the desired heading, speed, and altitude. The prior region model utilizes prior training and de-prioritization execution methods for computation. The top-level target search task model is built upon the underlying control model, eliminating the need for retraining on flight control principles and allowing for transfer to different top-level tasks, significantly simplifying training complexity. The top-level target locking task model can automatically recover lost targets within a certain range, reducing model switching frequency and improving control system stability. Secondly, this method divides the complex loyal wingman training task into manageable components, shortening training time and providing efficient autonomous target search and target search methods, thus addressing the instability issue during algorithm model switching.
[0090] Below, we will refer to Figures 1 to 15 The steps of the above-described loyal wingman target search and locking task execution method in this example embodiment will be described in more detail.
[0091] This application is based on a six-degree-of-freedom F-16 aircraft dynamics model; such as Figure 2 The image shown is a schematic diagram of a six-degree-of-freedom aircraft model.
[0092] In step S101, based on the environmental complexity and combat mission, a tuple consisting of 13 state variables was designed to represent the state space of the loyal wingman underlying control model.
[0093] [ΔH,Δψ,ΔV,H abs ,cosφ,sinφ,sinθ,cosθ,V x V y V z ,V]
[0094] Where ΔH is the difference between the desired altitude and the wingman's current altitude, Δψ is the difference between the desired heading and the wingman's current heading, ΔV is the difference between the desired speed and the wingman's current speed, and H... abs The wingman's current altitude is given by φ, roll angle is given by θ, and pitch angle is given by V. x V represents the wingman's x-axis velocity. y V represents the y-axis velocity of the wingman. z V is the z-axis velocity of the wingman, and V is the resultant velocity of the wingman; ΔH, H abs The dimensions of are km; the dimensions of Δψ, φ, and θ are rad; the dimensions of ΔV and V are... x V y V z The dimension of V is mh.
[0095] The action space of the underlying control model consists of a quadruple of four consecutive control variables.
[0096] [C a C e C r C t ]
[0097] Among them, C a This indicates the throttle lever control amount, used to control the speed of the wingman; C e This indicates the rudder control input, used to control the wingman's yaw angle; C r This indicates the elevator control input, used to control the pitch angle of the wingman; C t This represents the aileron control variable, used to control the roll angle of the wingman. The value ranges of each control variable are shown in Table 1.
[0098] Table 1. Range of Control Quantities
[0099] <![CDATA[C a ]]> [0,1] <![CDATA[C r ]]> [-1,1] <![CDATA[C e ]]> [-1,1] <![CDATA[C t ]]> [-1,1]
[0100] Furthermore, the continuous action space is discretized, that is, each continuous action is converted into a multi-discrete action space of size 50.
[0101] The goal is for the drone to fly at the expected heading, altitude, and speed, with a limited roll angle to prevent stalling and altitude loss due to excessive roll. Therefore, the following four reward functions are defined:
[0102]
[0103] All four reward functions are Gaussian in form, restricting the reward value to the interval (0,1). Different standard deviations of the Gaussian reward functions represent different tolerances for different indicators; the smaller the standard deviation, the lower the tolerance and the higher the accuracy requirement. In order to meet the requirements of heading, altitude, speed, and roll angle simultaneously as quickly as possible during training, this application uses the geometric mean of four rewards as the overall measure of these four reward indicators.
[0104]
[0105] In addition, a penalty will be imposed when a drone descends to a dangerous altitude of 1km.
[0106] R PH =-2
[0107] Finally, the sum of the above rewards is calculated as the reward for the wingman's underlying control model.
[0108] R1 = R C +R PH
[0109] Among them, R ψ R represents the heading reward function; H Represents a high-reward function; R V R represents the speed reward function; φ Represents the rolling reward function; R C R is the overall control reward function; PH R is a high-penalty constant; R is the overall reward of the underlying control model. The range of values for the above rewards is shown in Table 2.
[0110] Table 2. Range of values for the reward function
[0111] <![CDATA[R ψ ]]> (0,1] <![CDATA[R H ]]> (0,1] <![CDATA[R V ]]> (0,1] <![CDATA[R φ ]]> (0,1] <![CDATA[R C ]]> (0,1] <![CDATA[R PH ]]> -2 <![CDATA[R1]]> (-2,1]
[0112] Based on the characteristics of the mission, we define the following 7 termination conditions for the loyal wingman underlying control model.
[0113] 1. Failed to achieve the desired heading, speed, and altitude within the specified simulation steps: 500 simulation steps;
[0114] 2. The loyal wingman descended to a dangerous altitude of 1 km;
[0115] 3. Run time exceeded: 2×10 3 When the simulation step size is 1;
[0116] 4. The loyal wingman's altitude exceeds 10. 5 km;
[0117] 5. The angular velocities (p, q, r) of the loyal wingman are greater than 10. 3 rad / s;
[0118] 6. The speed of a loyal wingman exceeds 100m / h;
[0119] 7. The acceleration of a loyal wingman exceeds 20g.
[0120] like Figure 3 The diagram shows a hierarchical control architecture. The bottom-level model is responsible for controlling the aircraft's throttle, rudder, elevator, and ailerons to control the aircraft's flight and achieve the desired heading, speed, and altitude. Figure 4 The figure shows the reward curves for the training rounds of the underlying model of this method and the PPO algorithm.
[0121] In step S102, it is assumed that the loyal wingman has no attack capability (launching missiles). Therefore, the wingman's threat to the target lies only in guiding the missiles launched by the lead aircraft. Obviously, the target poses a greater threat to the wingman. Based on this characteristic, this application considers the worst-case scenario, that is, the constructed area model only considers the wingman's own trajectory and does not consider the target's heading. (When the threat difference is not significant, the target's heading is usually also considered, such as when our aircraft tail-chases the target or head-on with the target, the area will change significantly.)
[0122] A region model is constructed based on the radar detection area (the area enclosed by the dark red fan-shaped outline), encompassing both the in-fan termination area and the out-of-fan termination area. For example... Figure 5 The diagram shown illustrates the prior region model. The prior region model constructed in this application encompasses the termination region, the loss region (prior region), the danger region, and the safe region. This region model is key to the "prior-based training, de-prior-based execution" proposed in this application.
[0123] Since the regional model is based on the radar detection area, this application introduces two definitions: range ratio and angle ratio.
[0124] 1. Define the distance ratio F D This is the ratio of the target distance within the area to the radar detection radius of our aircraft; the range of the area distance ratio is defined as [0, 1.5].
[0125] 2. Define angle F A The ratio is the ratio of the angle between the target's velocity vector and the radian of the radar detection area; the range of the area angle ratio is defined as [-1.5, 1.5].
[0126] Angle ratio F A The sign indicates that the target is to the left of the wingman's velocity vector, and the sign indicates that the target is to the right of the wingman's velocity vector.
[0127] Since this definition is a relative ratio definition, different radar parameters (different detection ranges and detection angles) will not affect the target locking model. Therefore, the range ratio and angle ratio definitions for different region models in this application are shown in Table 3.
[0128] Table 3. Distance and Angle Ratios of Models in Different Regions
[0129]
[0130]
[0131] The radar parameters are shown in Table 4.
[0132] Table 4 Radar Model Parameters
[0133] Loyal Wingman Radar Parameters [-60°,60°] 30km Lead aircraft radar parameters - -
[0134] In step S103, the target search scheme of this application adopts waypoint search, that is, several waypoints are predetermined, and the loyal wingman will fly and scan the battlefield according to the waypoints, while clearing the way for the lead aircraft behind.
[0135] The model only needs to process the coordinates of the target waypoint at the top level, convert them into the desired heading, speed, and altitude of the loyal wingman, and then calculate the difference between these coordinates and the current heading, speed, and altitude of the aircraft. This difference is used as part of the observation and control of the bottom-level model.
[0136] The loyal wingman needs to travel to the i-th waypoint O. i Its latitude and longitude are (λ i ,φ i ,h i The loyal wingman's current latitude, longitude, and altitude are (λ, φ, h), and its horizontal velocity vector is... The latitude, longitude, and altitude of the origin of the combat area are (λ) o ,φ o ,h o By converting latitude and longitude to the navigation coordinate system, the waypoint coordinates are obtained as (x... i ,y i ,z i The loyal wingman's coordinates are (x, y, z). Therefore, the height difference can be obtained in the vertical plane as follows:
[0137] Δh=z i -z
[0138] On the horizontal plane, the position vector of the target waypoint relative to the loyal wingman can be obtained as follows:
[0139]
[0140] Then the magnitude of the heading angle difference can be obtained as follows:
[0141]
[0142] However, it is still uncertain whether the sign of the angle difference, i.e. whether the wingman should yaw to the left or to the right.
[0143] This application uses the following formula to determine the sign of the angle difference.
[0144]
[0145] The speed difference Δv can be selected to accelerate (the speed difference is a positive number), decelerate (the speed difference is a negative number), or remain at a constant speed (the speed difference is 0) according to the user's wishes.
[0146] At this point, the top-level processing is complete and the differences in heading, speed, and altitude are obtained. These differences can be used as input to the bottom-level control model to complete the waypoint-based target search task.
[0147] like Figure 6 The image shown is a schematic diagram of a waypoint target search mission for a leader-chief formation. Figure 7 The image shows the actual simulation trajectory of the top-level target search task model. It can be seen that the model can follow... Figure 6 The target search task was completed using the preset waypoints.
[0148] In step S104, during the target search performed by the loyal wingman, if a target is detected, the system switches to the top-level target locking task model. The target locking model is based on a priori region model. The top-level target locking task model establishes a lost zone outside the radar detection area but within the termination zone to represent that the locked target has left the detection area; that is, when the target is in the lost zone, it means the target is lost. This lost zone is also called the priori region because during training, it is assumed that the target's state information within the lost zone is known, while during deployment, since this area is outside the detection area, the target information is unknown. Because this application introduces priori information, namely the target's range ratio and angle ratio are known, it can guide the aircraft to quickly recover the target during the training phase. However, during the evaluation phase, since the target is outside the detection area, the target's range ratio and angle ratio are unknown, but the range ratio at the moment before the target is lost can be used as a reference. and angle ratio Determine the type of loss. If:
[0149] 1. Distance loss: The distance ratio to the observed value is set to the worst-case maximum of 1.5, and the angle ratio to the observed value is set to...
[0150] 2. Angle loss: If The angle ratio to the observed value is set to -1.5 in the worst-case scenario. The angle ratio to the observed value is set to 1.5 in the worst-case scenario, and the distance ratio to the observed value is set to...
[0151] 3. Distance + Angle Loss: If The angle ratio to the observed value is set to -1.5 in the worst-case scenario. The angle ratio to the observation is set to 1.5 in the worst-case scenario, and the distance ratio to the observation is set to 1.5 in the worst-case scenario.
[0152] By setting the distance ratio and angle ratio observations in this way, the model can guide the wingman to quickly and automatically find the target even without prior information during the model deployment phase.
[0153] In addition, by introducing the concept of a "lost zone" outside the detection zone as a buffer for target search and localization, this algorithm model reduces the switching frequency of different algorithm models and improves the robustness and stability of the entire combat system.
[0154] This application modifies and supplements the state space of the underlying control model, and designs a tuple consisting of 16 state variables to represent the state space of the top-level target locking control model for loyal wingmen.
[0155] [ΔH,Δψ,ΔV x H abs ,cosφ,sinφ,sinθ,cosθ,V x V y V z ,V,R,ψ r ,F A ,F D ]
[0156] Where ΔH represents the difference between the target enemy aircraft's altitude and the wingman's current altitude; Δψ represents the difference between the target enemy aircraft's horizontal velocity direction and the wingman's current heading; ΔV x H represents the difference between the target enemy aircraft's longitudinal axis velocity and the wingman's current longitudinal axis velocity; abs Indicates the wingman's current altitude; φ represents the wingman's roll angle; θ represents the wingman's pitch angle; V x V y V z ψ represents the wingman's velocity along the x, y, and z axes; V represents the wingman's resultant velocity; R represents the relative distance between the target enemy aircraft and the wingman; r F represents the angle between the target enemy aircraft's velocity vector and its line-of-sight (LOS); A Indicates the angle ratio; F D Indicates the distance ratio. ΔH, H abs The dimensions are km; Δψ, φ, θ, ψ r The dimension of ΔV is rad; x V x V y V z The dimension of V is mh; F A F D Dimensionless.
[0157] The action space of the top-level target locking control model consists of a triplet of three discrete error quantities, which controls the wingman's heading, speed, and altitude through five control levels.
[0158] [ΔH d ,Δψ d ,ΔV d ]
[0159] Where, ΔH d Indicates the height error; Δψ d Indicates the heading error; ΔV d This represents the speed error. The value ranges of each control variable are shown in Table 5.
[0160] Table 5. Range of values for top-level target locking control parameters
[0161]
[0162] For the reward function, firstly, consider the reward for the lost area. The reward elements for the lost area are distance loss, angle loss, and distance + angle loss. When the enemy target is located in this area, it means that the target can be automatically retrieved through this model. Define guidance penalties for three cases:
[0163]
[0164] Second, considering the penalty for being in the safe zone, the strategy is to maneuver the target to stay in the safe zone for as long as possible. Therefore, the penalty for the target being in the safe zone is as follows:
[0165] P safe =0
[0166] Third, considering the danger zone penalty, a Gaussian reward was designed to encourage wingmen to maneuver away from the danger zone by maneuvering themselves to keep the target as far away as possible.
[0167]
[0168] Fourth, consider the termination zone penalty. When the target is within the termination zone (inside the fan), it means the distance is too close, and the wingman lacks countermeasures, thus ending the target lock. When the target is outside the termination zone (outside the fan), it means the distance is too far, and the target cannot be retrieved, thus ending the turn.
[0169] P terminate =-5
[0170] Finally, applying only the above four penalties can lead to the model prematurely ending the round due to cumulative penalties, resulting in the target falling into the termination zone too early. Therefore, the reward for detecting targets outside the termination zone also needs to be considered. This application considers the reward given when the target is located in the safe zone and danger zone because the target is always within the detection area:
[0171] R detect =1
[0172] Among them, P DL This refers to the guidance penalty for lost distance, the purpose of which is to shorten the distance between our aircraft and the target to reduce the penalty; P AL This refers to the guidance penalty for angle loss, the purpose of which is to reduce the angular difference between the aircraft's radar detection zone and the target; P DAL The guidance penalty representing the loss of both distance and angle is R. DL With R AL The geometric mean aims to simultaneously reduce both distance and angle differences; P safe Indicates punishment within the safe zone; P danger Indicates punishment within the danger zone; P terminate Indicates termination of penalties within the designated area; R detect =1 indicates a reward within the detection zone. The range of the above penalty and reward values is shown in Table 6.
[0173] Table 6. Range of Values for Top-Level Target Locking Reward Function
[0174] <![CDATA[P DL ]]> [-3,-2) <![CDATA[P danger ]]> (-2,0] <![CDATA[P AL ]]> [-3,-2) <![CDATA[P terminate ]]> -5 <![CDATA[P DAL ]]> [-3,-2) <![CDATA[R detect ]]> 1 <![CDATA[P safe ]]> 0
[0175] In addition to the termination conditions of the underlying control model, based on the characteristics of this mission, the round should also end when the target enemy aircraft is in the termination zone.
[0176] The top-level target locking task model is based on prior training, while the execution method is based on deprioritization.
[0177] The training phase is as follows:
[0178] 1. Initialize parameters. The action network uses orthogonal initialization, with gain = 0.01 to ensure that each action has a chance to be taken at the beginning. The value network gain is set to 1. Load the underlying control model. The initial position of the target enemy aircraft is randomly given within the lost zone, and the target enemy aircraft information within the lost zone is known. The target's initial heading is random, and its heading and speed change randomly at specific intervals.
[0179] 2. Loop 1: Set the update count to num_updates
[0180] 3. Learning rate decay lrnow = 1.0 - (update - 1.0) / num_updates * lr
[0181] 1. Loop 2: Set the number of steps in a preview to the batch size.
[0182] 2. The global step count increases based on the number of parallel environments.
[0183] 3. Store the top-level observations; is the storage complete?
[0184] 4. Input the observations into the neural network (policy network and value network) to obtain the action, log probability, and estimated state value.
[0185] 5. Store the top-level state value, action value, and probability logarithmic value.
[0186] 6. The vector environment executes one step, using the top-level action signal as the bottom-level observation signal to obtain the next observation, reward, and whether to terminate.
[0187] 7. Store reward values, update round rewards and round length.
[0188] 8. End of loop 2
[0189] 4. Use a neural network value network to estimate the state value of the next observation returned after the pre-play ends, use generalized advantage estimation to calculate the advantage, and store the sum of the advantage value and the state value as the reward value.
[0190] 5. This gives us a preview of the next batch size, along with the logarithmic probability values, action values, advantage values, reward values, and estimated state values.
[0191] 1. Loop 3: Update the number of epochs
[0192] 2. Disrupt the storage experience pool
[0193] 1. Loop 4: Loop between batch size and mini-batch size
[0194] 2. Input small batches of shuffled observations and shuffled actions into the neural network (to calculate the new batch probability logarithm, entropy, and new state value).
[0195] 3. Calculate the ratio
[0196] 4. Calculate strategy losses based on small-batch advantages, ratios, and trimming factors.
[0197] 5. Calculate the mean square error between the new state value and the small batch return value as the value loss.
[0198] 6. Calculate the mean of the entropy as the entropy loss.
[0199] 7. Calculate the total loss
[0200] 8. Parameter update and gradient clipping
[0201] 9. End loop 4
[0202] 3. End loop 3
[0203] 6. End loop 1
[0204] The deployment phase is detailed as follows:
[0205] Target enemy aircraft information is no longer known within the lost zone.
[0206] 1. Target-locking model activation
[0207] 2. If the target leaves the detection area, return the distance ratio and angle ratio from the moment it was lost.
[0208] 3. Infer the loss type and assign the worst-case distance ratio and angle ratio values to the model.
[0209] 1. Distance loss: The distance ratio to the observed value is set to the worst-case maximum of 1.5, and the angle ratio to the observed value is set to...
[0210] 2. Angle loss: If The angle ratio to the observed value is set to -1.5 in the worst-case scenario. The angle ratio to the observed value is set to 1.5 in the worst-case scenario, and the distance ratio to the observed value is set to...
[0211] 3. Distance + Angle Loss: If The angle ratio to the observed value is set to -1.5 in the worst-case scenario. The angle ratio to the observed values is set to 1.5 in the worst-case scenario, and the distance ratio to the observed values is also set to 1.5 in the worst-case scenario.
[0212] 4. Loop 1: Given the maximum seek time step
[0213] 1. Successfully retrieved the target, exited the loop.
[0214] 2. Once the maximum retrieval step size is reached, if the retrieval fails, exit the loop and switch to the target search model.
[0215] 5. Loop 1 ends
[0216] like Figure 8 The figure shows the reward curve for the training rounds of the top-level target locking task model. To verify the effectiveness of the model, this application simulates target locking under three scenarios: triangular maneuver, head-on maneuver, and escape maneuver. The side and top views of the trajectories of both sides are shown below. Figures 9 to 14 As shown. Among them, Figure 9 Side view of the wingman's target lock-on maneuver and mission execution (triangular maneuver to target); Figure 10 A top-down view of the target triangular maneuver and wingman target locking mission execution. Figure 11 A side view of the wingman maneuvering towards the target and locking onto the target during mission execution; Figure 12 A top-down view of a maneuvering towards the target while the wingman locks onto the target and executes the mission. Figure 13 Side view of the wingman's target lock mission during the target escape maneuver; Figure 14A top-down view of the wingman's target lock mission during the target escape maneuver.
[0217] Analysis shows that the trained algorithm model can delay the approach of the target / maintain the distance by climbing, reducing speed, and turning when the target is approaching, and can continuously lock onto the target.
[0218] like Figure 15 As shown in the figure, the algorithm model switching frequency is compared with and without the lost region (prior region). It can be seen that the method of this application greatly reduces the model switching frequency and enhances the stability of the control system.
[0219] The aforementioned method for loyal wingman target search and locking tasks achieves several key benefits. Firstly, the underlying control model controls the aircraft's throttle, rudder, elevator, and ailerons to achieve the desired heading, speed, and altitude. The prior region model utilizes prior training and de-prioritization execution methods for computation. The top-level target search task model is built upon the underlying control model, eliminating the need for retraining on flight control principles and allowing for transfer to different top-level tasks, significantly simplifying training complexity. The top-level target locking task model can automatically recover lost targets within a certain range, reducing model switching frequency and improving control system stability. Secondly, this method divides the complex loyal wingman training task into manageable components, shortening training time and providing efficient autonomous target search and target search methods, thus addressing the instability issue during algorithm model switching.
[0220] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of embodiments of this disclosure, "a plurality of" means two or more, unless otherwise explicitly specified.
[0221] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this disclosure. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. In addition, those skilled in the art can combine and integrate the different embodiments or examples described in this specification.
[0222] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the appended claims.
Claims
1. A method for performing a loyal wingman target search and locking mission, characterized in that, include: Step S101: Based on the environmental complexity and combat mission, construct a low-level control model and use the low-level control model for flight control; among which, The steps involved in building the underlying control model include: Based on the environmental complexity and combat mission, several state variables are designed, and a low-level control model is constructed based on these state variables; the expression of the low-level control model is as follows: In the formula, The difference between the desired altitude and the wingman's current altitude. The difference between the desired heading and the wingman's current heading. The difference between the desired speed and the wingman's current speed. The wingman's current altitude. Roll the corner for the wingman. For the wingman's pitch angle, Let x be the wingman's velocity. Let y be the wingman's velocity. Let z be the wingman's z-axis velocity. For wingman speed; The action space of the underlying control model consists of a quadruple of four consecutive control variables: In the formula, The amount controlled by the throttle lever. For rudder control inputs, For elevator control parameters, For aileron control; Discretize the continuous action space to balance computational complexity and simulation accuracy; The steps to discretize a continuous action space include: Define the reward function: In the formula, For the heading reward function, For a high reward function, For the speed reward function, For the rolling reward function; The overall control reward function is calculated based on the heading reward function, altitude reward function, speed reward function, and roll reward function, and is used as the overall metric; the overall control reward function is expressed as: Set height penalty constant According to the height penalty constant And the overall reward of the wingman's underlying control model calculated by the overall metric: In the formula, The overall reward for the wingman's underlying control model; Step S102: Based on the radar detection area, introduce the range ratio and angle ratio to construct a priori region model; wherein, the priori region model includes a priori region sub-model, a safe region sub-model, a danger region sub-model and a termination region sub-model; The range ratio is the ratio of the target's distance to the radar's detection radius within the area, and the range ratio ranges from [0, 1.5]. The angle ratio is the ratio of the angle between the target and the wingman's velocity vector to the radian of the radar's detection area, and the angle ratio ranges from [-1.5, 1.5]. The angle ratio is signed; a positive sign indicates that the target is to the left of the wingman's velocity vector, and a negative sign indicates that the target is to the right of the wingman's velocity vector. Step S103: Based on the underlying control model, construct the top-level target search task model, and use the top-level target search task model to perform target search; Step S104: Based on the underlying control model and the prior region model, construct the top-level target locking task model, and use the top-level target locking task model to perform target locking; specifically including: Based on the underlying control model and the prior region model, a top-level target locking task model is constructed: In the formula, The difference between the target aircraft's longitudinal speed and the wingman's current longitudinal speed. The relative distance between the target enemy aircraft and its wingman. Let be the angle between the target enemy aircraft's velocity vector and its line of sight. For angle ratio, It is the distance ratio; The action space of the top-level target locking control model consists of a triplet of three discrete error quantities, which controls the wingman's heading, speed, and altitude through five control levels: In the formula, This is the height error amount. This is the heading error. This is the speed error amount; Penalties for sub-models in different regions are calculated to obtain penalties for lost areas, safe areas, danger zones, and termination zones; among them, The penalty for lost areas is: In the formula, As a guide penalty for lost distance, As a penalty for lost angle guidance, A penalty for losing both distance and angle; The safe zone penalty is: In the formula, Punishment in the safe zone; The penalty for dangerous zones is: In the formula, Punishment for dangerous zones; The penalty for termination is: In the formula, Penalty for termination zone; The reward given for a target remaining within the detection area regardless of whether it is in the safe zone or the danger zone: In the formula, Rewards will be given for the exploration area; The top-level target locking task model locks the target based on the penalties for lost zones, safe zones, danger zones, termination zones, and exploration zones.
2. The method for searching and locking a loyal wingman target according to claim 1, characterized in that, Based on the characteristics of the task, define the termination conditions of the underlying control model.
3. The method for searching and locking a loyal wingman target according to claim 2, characterized in that, The steps for performing target search using a top-level target search task model include: The top-level target search task model processes the coordinates of the top-level target waypoints, converting them into the desired heading, speed, and altitude of the loyal wingman, and then subtracts these from the current heading, speed, and altitude of the main aircraft. Convert the latitude and longitude to the navigation coordinate system to obtain the waypoint coordinates. The coordinates of the loyal wingman are Then, on the vertical plane, the height difference is: On the horizontal plane, the position vector of the target waypoint relative to the loyal wingman is obtained as follows: The magnitude of the heading angle difference can then be obtained as: In the formula, The velocity vector on the horizontal plane; Determine the sign of the angle difference: The difference in heading angle is used as the input to the underlying control model to complete the waypoint-based target search task.
Citation Information
Patent Citations
Priori information-based heuristic indoor environment robot exploration method and system
CN113110482A
Sequential Bayesian geoacoustic parameter inversion method based on shallow sea double-node modal time difference of arrival
CN116577826A