Multi-aircraft reinforcement learning space-time collaborative guidance method for dynamic threat area
By adjusting the IFDS algorithm parameters through deep reinforcement learning, the obstacle handling problem of collaborative guidance of multiple aircraft in dynamic threat areas was solved, and efficient trajectory planning and coordinated strike integration were achieved.
Patent Information
- Application Number
- CN202510698456.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-09-19
AI Technical Summary
Existing technologies fail to effectively handle obstacles in complex environments during multi-vehicle collaborative guidance, especially in dynamic threat areas, and it is difficult to strike a balance between computational efficiency and robustness.
A deep reinforcement learning algorithm is used to dynamically adjust the IFDS algorithm parameters, build multi-aircraft kinematic and obstacle models, and generate high-quality obstacle avoidance trajectories.
It improves the dynamic adaptability and coordinated strike effect of multiple aircraft in complex environments, reduces computational complexity, and ensures real-time performance and computational efficiency.
Smart Images

Figure CN120669714A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of collaborative guidance, and in particular, to a multi-aircraft reinforcement learning spatiotemporal collaborative guidance method for dynamic threat zones. Background Art
[0002] With the rapid development of modern anti-missile weapons and defense systems, the performance improvements of individual aircraft are no longer sufficient to meet increasingly complex mission requirements. Multi-aircraft collaboration has become a key research direction, showing a trend towards clustering and intelligent control. Multi-aircraft utilize collaborative guidance and control technologies to improve overall performance through information exchange between them.
[0003] Aircraft coordination encompasses both temporal and spatial coordination. However, existing research focuses solely on single coordinated attack scenarios, neglecting more complex environments, such as those with multiple obstacles. Designing collaborative trajectory planning methods that balance computational efficiency and robustness for the unexpected obstacles encountered by multiple aircraft during flight is a worthy research topic. This approach builds upon existing collaborative research by expanding its focus on obstacle avoidance in complex environments.
[0004] Traditional obstacle avoidance research has largely focused on drones and robotics, with common methods including Artificial Potential Field (APF), RRT, and A*. The IFDS (Interfered Fluid Dynamical System) algorithm, with its advantages such as rapid response to environmental changes and naturally smooth paths, has become a hot topic in the field. However, parameters in IFDS can affect the quality of the generated trajectory. Therefore, adaptively adjusting parameters to generate high-quality paths is particularly important in complex environments with various obstacles. Summary of the Invention
[0005] In order to overcome at least one deficiency in the prior art, the present application provides a multi-aircraft reinforcement learning spatiotemporal collaborative guidance method for dynamic threat areas.
[0006] In a first aspect, a multi-aircraft reinforcement learning spatiotemporal collaborative guidance method for a dynamic threat zone is provided, comprising:
[0007] Construct a multi-vehicle kinematic model and determine the space-time constraints;
[0008] Construct obstacle models;
[0009] Obtain information about the aircraft's surroundings, including obstacle locations, aircraft status, and target locations;
[0010] Based on the multi-aircraft kinematic model, obstacle model and the aircraft's surrounding environment information, a deep reinforcement learning algorithm is used to obtain the parameters of the IFDS algorithm; the parameters of the IFDS algorithm include repulsion parameters, tangential parameters and rotation angle parameters;
[0011] According to the aircraft's surrounding environment information and the parameters of the IFDS algorithm, the IFDS algorithm is used to generate the obstacle avoidance trajectory of each aircraft to the target position.
[0012] In one embodiment, the multi-vehicle kinematic model is:
[0013]
[0014]
[0015] Among them, P i is the distance from the i-th aircraft to the target, R i The derivative of V i is the speed of the i-th aircraft, θ i is the ballistic inclination angle of the i-th aircraft, is the line of sight angle of the i-th aircraft, is the sight angle of the i-th aircraft, φ i is the trajectory deviation angle of the i-th aircraft; is θ i The derivative of for The derivative of for The derivative of is φ i The derivative of a iy is the component of the acceleration of the i-th aircraft in the y-axis direction, a iz is the component of the acceleration of the i-th aircraft in the z-axis direction.
[0016] In one embodiment, the spatiotemporal constraints include time constraints and space constraints:
[0017] The time constraints are:
[0018]
[0019] Among them, η i is the total error of the remaining flight time between the i-th aircraft and other aircraft, ξ i,j is the remaining flight time error between the i-th aircraft and the j-th aircraft, n is the number of other aircraft, η D is the minimum time error value expected to be achieved, is the remaining flight time of the i-th aircraft, is the remaining flight time of the jth aircraft;
[0020] The spatial constraints are:
[0021] ||p i (t)-p j (t)||≥τ,i≠j
[0022] Among them, p i (t) is the position of the i-th aircraft at time t, p j (t) is the position of the j-th aircraft at time t, and τ is the minimum safety distance.
[0023] In one embodiment, the obstacle model is:
[0024]
[0025] Among them, F(x,y,z) is the obstacle model equation, (x,y,z) is any space coordinate, (x c ,y c ,z c ) are the coordinates of the obstacle center, and a, b, c, d, e, and f are coefficients related to the shape and size of the obstacle.
[0026] In one embodiment, the state space of the deep reinforcement learning algorithm is:
[0027] s t =[C1,C2,d,C3,L std ] T
[0028] Among them, s t is the state space, C1 is the vector pointing from the aircraft to the nearest obstacle surface, C2 is the vector pointing from the aircraft to the target, C3 is the difference vector between the trajectory of different aircraft and the mean trajectory, d is the distance from the aircraft to the nearest obstacle surface, L std is the standard deviation between the final tracks.
[0029] In one embodiment, the action space of the deep reinforcement learning algorithm is:
[0030] a t =[ρ n ,σ n ,θ n ] T
[0031] Among them, a t is the action space, ρ n is the exclusion parameter, σ n is the tangential parameter, θ n is the rotation angle parameter.
[0032] In one embodiment, the reward function of the deep reinforcement learning algorithm is:
[0033] reward=c1·reward1+c2·reward2+c3·reward3
[0034] Among them, reward is the reward function, reward1 is the obstacle and threat area reward, reward2 is the track reward, reward3 is the additional terminal reward, c1, c2, c3 are weights;
[0035] Obstacle and Threat Zone rewards are:
[0036]
[0037] Among them, d sur is the distance from the aircraft to the obstacle surface, d cen is the distance from the aircraft to the center of the obstacle area, r and ε are both thresholds;
[0038] The track rewards are:
[0039]
[0040] Among them, d tar is the distance from the aircraft to the target, d0 is the initial distance from the aircraft to the target;
[0041] Additional terminal rewards are:
[0042]
[0043] Among them, λ is the weight factor, r mean is the average reward of the track, L i is the track length of the i-th aircraft, L is the sum of the track lengths of all aircraft, mean is the mean, and N is the number of aircraft.
[0044] In a second aspect, a multi-aircraft reinforcement learning spatiotemporal collaborative guidance device for dynamic threat zones is provided, comprising:
[0045] The first model building module is used to build a multi-aircraft kinematic model and determine the time and space constraints;
[0046] The second model building module is used to build an obstacle model;
[0047] An information acquisition module is used to obtain information about the aircraft's surrounding environment, including obstacle locations, aircraft status, and target locations.
[0048] The deep reinforcement learning module is used to obtain the parameters of the IFDS algorithm based on the multi-aircraft kinematic model, obstacle model and the aircraft's surrounding environment information using a deep reinforcement learning algorithm. The parameters of the IFDS algorithm include repulsion parameters, tangential parameters and rotation angle parameters.
[0049] The trajectory generation module is used to generate the obstacle avoidance trajectory of each aircraft to the target position using the IFDS algorithm based on the aircraft's surrounding environment information and the parameters of the IFDS algorithm.
[0050] In a third aspect, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned multi-aircraft reinforcement learning spatiotemporal collaborative guidance method for dynamic threat areas.
[0051] In a fourth aspect, a computer program product is provided, comprising a computer program / instruction. When the computer program / instruction is executed by a processor, the above-mentioned multi-aircraft reinforcement learning spatiotemporal collaborative guidance method for dynamic threat areas is used.
[0052] Compared with the prior art, this application has the following beneficial effects:
[0053] 1. Use deep reinforcement learning algorithms to dynamically adjust IFDS algorithm parameters, enabling multiple aircraft to achieve spatiotemporal coordinated strikes in complex and uncertain environments, and improve adaptability to dynamic environments.
[0054] 2. Integrate trajectory planning and coordinated strikes, optimize parameters through deep reinforcement learning, improve strike effectiveness, and realize the integration of trajectory planning and coordinated strikes.
[0055] 3. Use deep reinforcement learning to simplify parameter management and optimization, improve coordination accuracy and reliability, and reduce the complexity of multi-aircraft coordination.
[0056] 4. Optimize the algorithm to reduce computational complexity, ensure the real-time execution of multi-aircraft coordinated strike missions, and improve real-time performance and computational efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] The present application may be better understood by referring to the following description in conjunction with the accompanying drawings, which together with the following detailed description are incorporated into and form a part of this specification. In the drawings:
[0058] Figure 1 The flowchart of the multi-aircraft reinforcement learning spatiotemporal collaborative guidance method for dynamic threat areas is shown;
[0059] Figure 2The dynamic parameters of the IFDS algorithm in different scenarios are shown. (a) is the dynamic parameters of the IFDS algorithm in scenario 1, (b) is the dynamic parameters of the IFDS algorithm in scenario 2, and (c) is the dynamic parameters of the IFDS algorithm in scenario 3.
[0060] Figure 3 The figure shows the aircraft trajectory diagrams under different scenarios, where (a) is the aircraft trajectory diagram under scenario 1, (b) is the aircraft trajectory diagram under scenario 2, and (c) is the aircraft trajectory diagram under scenario 3;
[0061] Figure 4 The structural block diagram of the multi-aircraft reinforcement learning spatiotemporal collaborative guidance device for dynamic threat areas is shown. DETAILED DESCRIPTION
[0062] Exemplary embodiments of the present application are described below with reference to the accompanying drawings. For the sake of clarity and conciseness, not all features of actual embodiments are described in this specification. However, it should be understood that in the process of developing any such actual embodiment, many implementation-specific decisions may be made to achieve the developer's specific goals, and these decisions may vary from one implementation to another.
[0063] It is also necessary to explain here that, in order to avoid obscuring the present application due to unnecessary details, the accompanying drawings only show the device structure closely related to the solution according to the present application, while other details that are not closely related to the present application are omitted.
[0064] It should be understood that the present application is not limited to the described embodiments due to the following description with reference to the accompanying drawings. In this document, where feasible, the embodiments may be combined with each other, features between different embodiments may be replaced or borrowed, and one or more features may be omitted in one embodiment.
[0065] The present invention provides a multi-aircraft reinforcement learning spatiotemporal collaborative guidance method for dynamic threat areas. Figure 1 The flowchart of the multi-aircraft reinforcement learning spatiotemporal collaborative guidance method for dynamic threat areas is shown in Figure 2. Figure 1 , methods include:
[0066] Step S1: construct a multi-aircraft kinematic model and determine the time and space constraints.
[0067] Specifically, the multi-vehicle kinematic model is:
[0068]
[0069] Among them, R i is the distance from the i-th aircraft to the target, Ri The derivative of V i is the speed of the i-th aircraft, θ i is the ballistic inclination angle of the i-th aircraft, is the line of sight angle of the i-th aircraft, is the sight angle of the i-th aircraft, φ i is the trajectory deviation angle of the i-th aircraft; is θ i The derivative of for The derivative of for The derivative of is φ i The derivative of a iy is the component of the acceleration of the i-th aircraft in the y-axis direction, a iz is the component of the acceleration of the i-th aircraft in the z-axis direction.
[0070] Spatiotemporal constraints include time constraints and space constraints:
[0071] The time constraints are:
[0072]
[0073] Among them, η i is the total error of the remaining flight time between the i-th aircraft and other aircraft, ξ i,j is the remaining flight time error between the i-th aircraft and the j-th aircraft, n is the number of other aircraft, η D is the minimum time error value expected to be achieved, is the remaining flight time of the i-th aircraft, is the remaining flight time of the jth aircraft;
[0074] The spatial constraints are:
[0075] ||p i (t)-p j (t)||≥τ,i≠j
[0076] Among them, p i (t) is the position of the i-th aircraft at time t, p j (t) is the position of the j-th aircraft at time t, and τ is the minimum safety distance.
[0077] Step S2: construct an obstacle model.
[0078] Specifically, the obstacle model is:
[0079]
[0080] Among them, F(x,y,z) is the obstacle model equation, (x,y,z) is any space coordinate, (x c ,y c ,z c ) are the coordinates of the obstacle center, and a, b, c, d, e, and f are coefficients related to the shape and size of the obstacle.
[0081] Step S3: Acquire the aircraft's surrounding environment information, which includes obstacle locations, aircraft status, and target locations.
[0082] Step S4, based on the multi-aircraft kinematic model, obstacle model and aircraft surrounding environment information, a deep reinforcement learning algorithm is used to obtain the parameters of the IFDS algorithm; the parameters of the IFDS algorithm include repulsion parameters, tangential parameters and rotation angle parameters.
[0083] Here, we first train the deep reinforcement learning algorithm to obtain the agent weights. The specific training process includes:
[0084] 1. Initialize the policy network and value network: The policy network is used to output the action probability distribution, and the value network is used to estimate the state value.
[0085] 2. Sampling: Under the current strategy, interact with the environment and collect a set of trajectory data.
[0086] 3. Calculate the Advantage Function: The advantage function measures the pros and cons of taking a certain action in a certain state relative to the current strategy.
[0087] 4. Update the policy network: By optimizing the objective function, update the parameters of the policy network to ensure that the policy update does not deviate too far from the current policy.
[0088] 5. Update the value network: Update the parameters of the value network by minimizing the prediction error of the value network.
[0089] 6. Repeat: Repeat the above steps until the strategy converges.
[0090] After training, the environment information surrounding the aircraft is used as input to the deep reinforcement learning algorithm based on the agent weights. Here, the deep reinforcement learning algorithm runs on the basis of the multi-aircraft kinematic model and obstacle model, and finally obtains the parameters of the IFDS algorithm.
[0091] Specifically, the state space of the deep reinforcement learning algorithm is:
[0092] s t =[C1,C2,d,C3,L std ] T
[0093] Among them, s t is the state space, C1 is the vector pointing from the aircraft to the nearest obstacle surface, C2 is the vector pointing from the aircraft to the target, C3 is the difference vector between the trajectory of different aircraft and the mean trajectory, d is the distance from the aircraft to the nearest obstacle surface, L std is the standard deviation between the final tracks.
[0094] The action space of the deep reinforcement learning algorithm is:
[0095] a t =[ρ n ,σ n ,θ n ] T
[0096] Among them, a t is the action space, ρ n is the exclusion parameter, σ n is the tangential parameter, θ n is the rotation angle parameter, and the value range is set to [0.1,3].
[0097] Repulsion parameter ρ n , controls the degree of fluid repulsion to obstacles. A higher value will cause the aircraft track to be further away from obstacles, thus generating a safer obstacle avoidance track; the tangential parameter σ n , which determines the tangent degree of the fluid when it bypasses the obstacle and affects the direction of fluid flow; the rotation angle parameter θ n , controls the rotation angle of the fluid flow and affects the curvature of the track.
[0098] Specifically, the reward function of the deep reinforcement learning algorithm is:
[0099] reward=c1·reward1+c2·reward2+c3·reward3
[0100] Among them, reward is the reward function, reward1 is the obstacle and threat area reward, reward2 is the track reward, reward3 is the additional terminal reward, c1, c2, c3 are weights;
[0101] To ensure the safety of the planned trajectory, if the next planned trajectory point falls into an obstacle, a certain penalty should be given. At the same time, a safe distance from the obstacle surface should be maintained. If the spacecraft enters the threat zone, a penalty is also added. When pre-training is completed and the weights are updated, training stops if the spacecraft enters an obstacle. The rewards for obstacles and threat zones are:
[0102]
[0103] Among them, d sur is the distance from the aircraft to the obstacle surface, d cen is the distance from the aircraft to the center of the obstacle area, r and ε are both thresholds;
[0104] In order to ensure that the spacecraft arrive at the target in a coordinated manner with the shortest possible trajectory, a trajectory penalty is required to avoid unnecessary exploration behavior. The trajectory reward is:
[0105]
[0106] Among them, d tar is the distance from the aircraft to the target, d0 is the initial distance from the aircraft to the target;
[0107] When the training is finished or terminated, the spacecraft's track lengths are compared and rewards are given. The goal is to make the difference in track length as small as possible. The additional terminal rewards are:
[0108]
[0109] Among them, λ is the weight factor, r mean is the average reward of the track, L i is the track length of the i-th aircraft, L is the sum of the track lengths of all aircraft, mean is the mean, and N is the number of aircraft.
[0110] Figure 2 The dynamic parameters of the IFDS algorithm in different scenarios are shown. Among them, (a) is the dynamic parameters of the IFDS algorithm in scenario 1, (b) is the dynamic parameters of the IFDS algorithm in scenario 2, and (c) is the dynamic parameters of the IFDS algorithm in scenario 3.
[0111] Step S5: Based on the aircraft's surrounding environment information and the parameters of the IFDS algorithm, the IFDS algorithm is used to generate an obstacle avoidance track for each aircraft to the target position.
[0112] Figure 3 The aircraft trajectory diagrams under different scenarios are shown, where (a) is the aircraft trajectory diagram under scenario 1, (b) is the aircraft trajectory diagram under scenario 2, and (c) is the aircraft trajectory diagram under scenario 3.
[0113] This embodiment uses a deep reinforcement learning algorithm to dynamically adjust the IFDS algorithm parameters, enabling multiple aircraft to achieve spatiotemporal coordinated strikes in complex and uncertain environments, improving adaptability to dynamic environments. By integrating trajectory planning with coordinated strikes, the parameters are optimized through deep reinforcement learning to improve strike effectiveness and achieve the integration of trajectory planning and coordinated strikes.
[0114] Based on the same inventive concept as the multi-aircraft reinforcement learning spatiotemporal collaborative guidance method for dynamic threat zones, this embodiment also provides a corresponding multi-aircraft reinforcement learning spatiotemporal collaborative guidance device for dynamic threat zones. Figure 4 The structural block diagram of the multi-aircraft reinforcement learning spatiotemporal collaborative guidance device for dynamic threat areas is shown. The device includes:
[0115] The first model building module is used to build a multi-aircraft kinematic model and determine the time and space constraints;
[0116] The second model building module is used to build an obstacle model;
[0117] An information acquisition module is used to obtain information about the aircraft's surrounding environment, including obstacle locations, aircraft status, and target locations.
[0118] The deep reinforcement learning module is used to obtain the parameters of the IFDS algorithm based on the multi-aircraft kinematic model, obstacle model and the aircraft's surrounding environment information using a deep reinforcement learning algorithm. The parameters of the IFDS algorithm include repulsion parameters, tangential parameters and rotation angle parameters.
[0119] The trajectory generation module is used to generate the obstacle avoidance trajectory of each aircraft to the target position using the IFDS algorithm based on the aircraft's surrounding environment information and the parameters of the IFDS algorithm.
[0120] The multi-aircraft reinforcement learning spatiotemporal collaborative guidance device for dynamic threat areas in this embodiment has the same inventive concept as the multi-aircraft reinforcement learning spatiotemporal collaborative guidance method for dynamic threat areas mentioned above. Therefore, the specific implementation method of the device can be seen in the embodiment part of the optimization method of the Fourier stack imaging illumination system in the previous text, and its technical effects correspond to the technical effects of the above-mentioned method, which will not be repeated here.
[0121] An embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned multi-aircraft reinforcement learning spatiotemporal collaborative guidance method for dynamic threat areas.
[0122] An embodiment of the present application provides a computer program product, including a computer program / instruction. When the computer program / instruction is executed by a processor, it implements the above-mentioned multi-aircraft reinforcement learning spatiotemporal collaborative guidance method for dynamic threat areas.
[0123] The above descriptions are merely examples of various embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any modifications or substitutions that can be readily conceived by a person skilled in the art within the technical scope disclosed in the present application should be included within the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A multi-aircraft reinforcement learning spatiotemporal collaborative guidance method for dynamic threat areas, characterized by: include: Construct a multi-vehicle kinematic model and determine the space-time constraints; Construct obstacle models; Acquiring information about the aircraft's surrounding environment, including obstacle locations, aircraft status, and target locations; Based on the multi-aircraft kinematic model, the obstacle model and the aircraft surrounding environment information, a deep reinforcement learning algorithm is used to obtain parameters of the IFDS algorithm; the parameters of the IFDS algorithm include a repulsion parameter, a tangential parameter and a rotation angle parameter; According to the aircraft surrounding environment information and the parameters of the IFDS algorithm, the IFDS algorithm is used to generate an obstacle avoidance track for each aircraft to the target position.
2. The method according to claim 1, wherein The multi-aircraft kinematic model is: Among them, R i is the distance from the i-th aircraft to the target, R i The derivative of V i is the speed of the i-th aircraft, θ i is the ballistic inclination angle of the i-th aircraft, is the line of sight angle of the i-th aircraft, is the sight angle of the i-th aircraft, φ i is the trajectory deviation angle of the i-th aircraft; is θ i The derivative of for The derivative of for The derivative of is φ i The derivative of a iy is the component of the acceleration of the i-th aircraft in the y-axis direction, a iz is the component of the acceleration of the i-th aircraft in the z-axis direction.
3. The method according to claim 1, wherein The spatiotemporal constraints include time constraints and space constraints: The time constraints are: Among them, η i is the total error of the remaining flight time between the i-th aircraft and other aircraft, ξ u,j is the remaining flight time error between the i-th aircraft and the j-th aircraft, n is the number of other aircraft, η D is the minimum time error value expected to be achieved, is the remaining flight time of the i-th aircraft, is the remaining flight time of the jth aircraft; The spatial constraints are: ||p i (t)-p j (t)||≥τ,i≠j Among them, p i (t) is the position of the i-th aircraft at time t, p j (t) is the position of the j-th aircraft at time t, and τ is the minimum safety distance.
4. The method according to claim 1, wherein The obstacle model is: Among them, F(x,y,z) is the obstacle model equation, (x,y,z) is any space coordinate, (x c ,y c ,z c ) are the coordinates of the obstacle center, and a, b, c, d, e, and f are coefficients related to the shape and size of the obstacle.
5. The method according to claim 1, wherein The state space of the deep reinforcement learning algorithm is: s t =[C1,C2,d,C3,L std ] T Among them, s t is the state space, C1 is the vector pointing from the aircraft to the nearest obstacle surface, C2 is the vector pointing from the aircraft to the target, C3 is the difference vector between the trajectory of different aircraft and the mean trajectory, d is the distance from the aircraft to the nearest obstacle surface, L std is the standard deviation between the final tracks.
6. The method according to claim 1, wherein The action space of the deep reinforcement learning algorithm is: a t =[ρ n ,s n ,i n ] T Among them, a t is the action space, ρ n is the exclusion parameter, σ n is the tangential parameter, θ n is the rotation angle parameter.
7. The method according to claim 1, wherein The reward function of the deep reinforcement learning algorithm is: Among them, reward is the reward function, reward1 is the obstacle and threat area reward, reward2 is the track reward, reward3 is the additional terminal reward, c1, c2, c3 are weights; The Obstacle and Threat Zone rewards are: Among them, d sur is the distance from the aircraft to the obstacle surface, d cen is the distance from the aircraft to the center of the obstacle area, r and ε are both thresholds; The track rewards are: Among them, d tar is the distance from the aircraft to the target, d0 is the initial distance from the aircraft to the target; The additional terminal rewards are: Among them, λ is the weight factor, r mean is the average reward of the track, L i is the track length of the i-th aircraft, L is the sum of the track lengths of all aircraft, mean is the mean, and N is the number of aircraft.
8. A multi-aircraft reinforcement learning spatiotemporal collaborative guidance device for dynamic threat areas, characterized by: include: The first model building module is used to build a multi-aircraft kinematic model and determine the time and space constraints; The second model building module is used to build an obstacle model; An information acquisition module is used to acquire information about the aircraft's surrounding environment, including obstacle locations, aircraft status, and target locations; A deep reinforcement learning module is configured to obtain parameters of an IFDS algorithm using a deep reinforcement learning algorithm based on the multi-aircraft kinematic model, the obstacle model, and information about the aircraft's surrounding environment; the parameters of the IFDS algorithm include a repulsion parameter, a tangential parameter, and a rotation angle parameter; The track generation module is used to generate an obstacle avoidance track for each aircraft to a target position using the IFDS algorithm according to the aircraft surrounding environment information and the parameters of the IFDS algorithm.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the multi-aircraft reinforcement learning spatiotemporal collaborative guidance method for dynamic threat areas as described in any one of claims 1 to 7.
10. A computer program product, characterized in that The invention comprises a computer program / instruction, which, when executed by a processor, implements the multi-aircraft reinforcement learning spatiotemporal collaborative guidance method for dynamic threat areas as described in any one of claims 1 to 7.