A Three-Dimensional Adaptive Multi-Missile Cooperative Guidance Method Based on Deep Network Optimization

By introducing deep network optimization technology in multi-elastic coordinated guidance and adaptively adjusting the proportional guidance coefficient, the problems of complex design of multi-elastic coordinated guidance law and difficult parameter selection in the existing technology are solved, and more efficient multi-elastic coordinated guidance effect and the autonomous coordination ability of the aircraft cluster are achieved.

CN119002545BActive Publication Date: 2025-06-10HARBIN INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410919842.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-10
Publication Date
2025-06-10
Estimated Expiration
2044-07-10

AI Technical Summary

Technical Problem

The existing multi-elastic coordinated guidance law is complex, relies on precise dynamic models, and the selection of proportional guidance coefficients lacks effective methods, which affects the guidance effect.

Method used

The three-dimensional adaptive multi-elastic collaborative guidance method based on deep network optimization is adopted. By establishing a three-dimensional collaborative guidance model, designing a limited time multi-elastic collaborative guidance law, and using deep neural network to adaptively adjust the proportional guidance coefficient, the collaborative guidance remaining flight time consistency error is optimized.

Benefits of technology

The multi-elastic coordinated guidance effect is improved, the dependence on precise dynamic models is reduced, the independent selection and online update of guidance parameters are realized, and the autonomous coordination and strike capabilities of the aircraft cluster are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119002545B_ABST
    Figure CN119002545B_ABST
Patent Text Reader

Abstract

The present invention discloses a three-dimensional adaptive multi-missile cooperative guidance method based on deep network optimization, comprising the following steps: Step 1, establishment of a three-dimensional cooperative guidance model; Step 2, design of a multi-missile cooperative guidance law; Step 3, optimization of multi-missile cooperative guidance law parameters; The present invention constructs a cooperative training and testing framework for a missile swarm based on a deep network. By introducing a deep neural network and based on the interactive learning mechanism between an agent and the environment, it realizes the autonomous selection of guidance parameters within a given parameter range, effectively enhances the cooperative guidance effect of the missile swarm, and avoids problems such as strong dependence on the cooperative guidance algorithm model and difficulty in parameter selection in the prior art. By adjusting the flight time of each aircraft, simultaneous or sequential attacks are achieved, the strike probability against three-dimensional targets is increased, and an intelligent missile swarm cooperative guidance method with autonomous decision-making ability is formed, thereby making up for the defects of traditional methods and improving the autonomous cooperation and strike ability of the aircraft cluster.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of aircraft guidance and control, and specifically to a three-dimensional adaptive multi-missile cooperative guidance method based on deep network optimization. Background Technique

[0002] In the prior art, for the design of cooperative guidance laws for homogeneous / heterogeneous missile groups, it can be divided into designing cooperative guidance laws based on guidance time consistency and cooperative guidance laws based on simultaneous constraints of attack angle and guidance time. Specifically, there are the following several types: 1. Leader-follower strategy, which transforms the problem of simultaneous arrival of attack time into the problem of tracking the lead angle of the leader-follower, so that each follower missile attacks the target at the same time as the leader's attack time, realizing multi-missile time cooperative guidance; 2. Using the first derivative of the missile normal acceleration as the control quantity to control the attack angle, and the remaining time error feedback term as an additional term to achieve consistency in time, and finally realizing cooperative attack with the desired attack angle and specified guidance time based on the optimal control theory; 3. In order to achieve the desired attack time and attack angle constraints, the prior art has proposed a line-of-sight angle rate shaping method, and designed a cooperative attack guidance law based on the second-order sliding mode theory of finite-time convergence.

[0003] Most of the existing multi-missile cooperative guidance law designs are based on non-linear control theory, and online estimation of external disturbances or target maneuver information is realized by designing complex non-linear disturbance observers. This type of method has the disadvantage of relying on an accurate dynamic model, and the cooperative guidance law design process is complex, with many parameters to be selected. In practical engineering applications, the cooperative guidance method based on the improved proportional navigation law is widely used in practical engineering due to its advantages of fewer required guidance information parameters and high model robustness. However, the quality of multi-missile coordinated guidance effect depends seriously on the selection of the proportional navigation coefficient. At present, there is no effective method for selecting the proportional guidance coefficient. Summary of the Invention

[0004] The purpose of the present invention is to provide a three-dimensional adaptive multi-missile cooperative guidance method based on deep network optimization to solve the problems raised in the above background technique.

[0005] To achieve the above purpose, the present invention provides the following technical solution: A three-dimensional adaptive multi-missile cooperative guidance method based on deep network optimization, including the following steps: Step 1, establishing a three-dimensional cooperative guidance model; Step 2, designing a multi-missile cooperative guidance law; Step 3, optimizing the parameters of the multi-missile cooperative guidance law.

[0006] Among them, in the above Step 1, a missile-target relative kinematic model of the missile group relative to a fixed target is established, the influence of the remaining flight time on the cooperative guidance effect is analyzed, and a three-dimensional cooperative guidance model is constructed.

[0007] In the above step 2, based on the finite-time consistency theory, a finite-time multi-missile cooperative guidance law applicable to fixed targets is designed on the basis of the classical cooperative proportional guidance law;

[0008] In the above step 3, based on the deep network, the proportional guidance coefficient of the multi-missile cooperative guidance law obtained in step 2 is adaptively adjusted to optimize the consistency error of the remaining flight time of the cooperative guidance.

[0009] Preferably, in the above step 1, specifically:

[0010] Taking a fixed target point as the research background, a three-dimensional cooperative guidance model is constructed. Let (OXYZ) I , (OXYZ) LOS , (OXYZ) M be the ground inertial coordinate system, the line-of-sight coordinate system, and the vehicle velocity coordinate system respectively. R, θ LOS , φ LOS are the relative distance between the missile and the target, the line-of-sight inclination angle, and the line-of-sight deflection angle respectively. V m , θ m , φ m are the flight speed, the lead speed inclination angle, and the lead speed deflection angle respectively. The above variables satisfy the following dynamic equations:

[0011]

[0012] In the formula are the two acceleration components of the vehicle in the normal direction, which are the control input variables of the guidance model and satisfy the following relationship:

[0013]

[0014] The field of view angle of the vehicle relative to the line-of-sight system is defined as σ m , and there is:

[0015] cosσ m =cosθ m cosφ m (3)

[0016] A multi-vehicle system composed of n agents has an undirected connected communication topology. Let the variable A = [a ij ∈ R n×n be the adjacency matrix of the multi-vehicle system. In the matrix elements, a ij =1 indicates that member i can receive the information of member j, otherwise a ij =0. The variable L = [l ij ∈ R n×n is the Laplace matrix of the multi-vehicle system, where the matrix element lij = -a ij (i ≠ j), N i The set of neighbor nodes of the member, satisfying N i = {j | a ij = 1, 1 ≤ j ≤ n};

[0017] Under the condition of a given target point, design the cooperative guidance commands for multiple aircraft so that the multiple aircraft can reach the target point simultaneously within a given time. Let the whole flight time of the aircraft be T f,i , define the remaining flight time t of a single aircraft as follows go,i , (1 ≤ i ≤ n), then there is:

[0018] t go,i = T f,i - t(4)

[0019] Where t is the current flight time;

[0020] For a fixed target point, the remaining flight time of a single aircraft can be estimated by the following formula:

[0021]

[0022] Where R i , σ m,i , V m,i are respectively the missile-to-target distance, field of view angle and flight speed of member i, and N ~ [3, 7] is the proportional navigation coefficient;

[0023] To achieve the cooperative arrival of multiple aircraft at the desired target point, define the remaining flight time consistency error of member i as follows:

[0024]

[0025] Preferably, in the second step, specifically:

[0026] The classical cooperative proportional navigation law is as follows:

[0027]

[0028] Where N is the proportional navigation coefficient to be selected, then there is:

[0029]

[0030] Get the following relationship:

[0031]

[0032] Based on the small angle assumption, take the derivative of equation (5) to get:

[0033]

[0034] Define the average remaining flight time of multiple aircraft at time t as The remaining flight time error of a single aircraft is At the same time, let

[0035]

[0036] The multi-missile cooperative guidance law is designed as follows:

[0037]

[0038] Preferably, the stability of the multi-missile cooperative guidance law is proved by the following method:

[0039] Before proving the stability of the multi-missile cooperative guidance law, the following lemma is first given: For the nonlinear system If there exists a positive definite radially unbounded function V(x): x→R that satisfies the following inequality:

[0040]

[0041] where a, b > 0, then the function V(x) converges to a small neighborhood of the origin in finite time, and the convergence time T f satisfies:

[0042]

[0043] Based on the designed multi-missile cooperative guidance law in Equation (12), the following theorem is given: For a multi-aircraft system with undirected full connection, for a three-dimensional fixed target point, using the three-dimensional guidance model shown in Equation (1) and the cooperative guidance law shown in Equation (12), then the multi-aircraft will reach the desired target point in finite time;

[0044] Proof: Design the reference Lyapunov function Take the derivative of it to get:

[0045]

[0046] From Equation (11) Δ j According to the definition, the above equation can be further written as:

[0047]

[0048] where σ min,j is the minimum field of view angle of member j;

[0049] Let Then according to the lemma, the function V converges in finite time T f :

[0050]

[0051] If it converges to the small neighborhood of the 0 point internally, the theorem is proved.

[0052] Preferably, in the third step, specifically: model multiple aircraft as multiple agent systems, and use the remaining flight time t of member i go,i and the total remaining aircraft time consistency error at the current moment as the input of the deep neural network, and use the current proportional navigation coefficient N of member i i as the network output. By designing certain reward and punishment rules and network training and learning, adaptively adjust the amplitude of the proportional guidance coefficient to complete the optimization of the remaining flight time consistency error of cooperative guidance;

[0053] Based on Equation (12), the multi-missile cooperative guidance law optimized by the deep network is modified to:

[0054]

[0055] In the formula is the proportional guidance coefficient after neural network optimization.

[0056] Preferably, the deep network uses the DDPG algorithm as the training framework for the agent. In the DDPG algorithm, the Q value and V value of the agent are approximated by two deep neural networks, where the network Q value is expressed as:

[0057] y t = r t + γQ(s t+1 , π(s t+1 |φ π ′)|θ Q ′) (19)

[0058] where θ′ Q and φ π ′ are the hyperparameters of the target Q value network and V value network;

[0059] In the DDPG algorithm, the following formula is used to optimize the Q value network parameters:

[0060]

[0061] where N is the sample size;

[0062] In the DDPG algorithm, the following formula is used to optimize the action network parameters:

[0063]

[0064] The target network is updated using the following soft update method:

[0065]

[0066] where τ is the soft update factor.

[0067] Preferably, the action network Actor of the deep neural network is modeled as a multi-input multi-output MIMO system. Suppose the multi-agent system contains n members, then the neural network input space is designed as:

[0068] Input = [t go1 (t), t go,2 (t),..., t go,n (t)] T ∈ R n (23)

[0069] The network output is designed as:

[0070] Output = [N 1 (t), N 2 (t),..., N n (t)] T ∈ R n . (24)

[0071] Preferably, the comprehensive reward function of the deep neural network consists of the following three parts:

[0072] 1) Distance reward:

[0073]

[0074] where R(t - 1) and R(t) are the relative distances between the missile and the target at times t - 1 and t respectively, is the magnitude of the velocity vector w 1 is the reward weight coefficient, with a value of 0.1;

[0075] 2) Remaining flight time error function reward:

[0076]

[0077] In the formula: t go,i is the current remaining flight time of member i, is the average remaining flight time of the current multi-aircraft system, w 2 is the reward weight coefficient, with a value of 0.05;

[0078] 3) Process constraint reward:

[0079]

[0080] In the formula: t max , Rmax are the maximum training step and the maximum allowable miss distance. The training end reward function is triggered when the aircraft successfully reaches the target point or other termination conditions occur.

[0081] Preferably, the other termination conditions include training step timeout and aircraft state singularity.

[0082] Preferably, the comprehensive reward function is:

[0083]

[0084] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention constructs a collaborative training and testing framework for a missile swarm based on a deep network. By introducing a deep neural network and based on the interaction learning mechanism between the agent and the environment, it realizes the autonomous selection of guidance parameters within a given parameter range, effectively enhances the collaborative guidance effect of the missile swarm, and avoids problems such as strong dependence on the collaborative guidance algorithm model and difficulty in parameter selection in the prior art. By adjusting the flight time of each aircraft, it realizes simultaneous or sequential attacks, improves the strike probability against three-dimensional targets, and forms an intelligent collaborative guidance method for a missile swarm with autonomous decision-making ability, thereby making up for the defects of traditional methods and improving the autonomous collaboration and strike ability of the aircraft cluster. BRIEF DESCRIPTION OF THE DRAWINGS

[0085] Figure 1 is a schematic diagram of a three-dimensional collaborative guidance model;

[0086] Figure 2 is a graph of the change of the reward function during the training process of multiple aircraft;

[0087] Figure 3 is a comparison diagram of the three-dimensional flight trajectory before and after optimization;

[0088] Figure 4 is a comparison diagram of the overload of multiple aircraft before and after optimization;

[0089] Figure 5 is a comparison diagram of the Tgo error of multiple aircraft before and after optimization;

[0090] Figure 6 is a comparison diagram of the Tgo of multiple aircraft before and after optimization;

[0091] Figure 7 is a comparison diagram of the cumulative error of Tgo before and after optimization;

[0092] Figure 8 is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0093] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0094] Please refer to Figure 1-8 , a technical solution provided by the present invention:

[0095] Embodiment:

[0096] A three-dimensional adaptive multi-missile cooperative guidance method based on deep network optimization, comprising the following steps: Step 1, establishment of a three-dimensional cooperative guidance model; Step 2, design of a multi-missile cooperative guidance law; Step 3, optimization of multi-missile cooperative guidance law parameters.

[0097] Among them, in the above Step 1, a missile-target relative kinematic model of the missile group relative to a fixed target is established, the influence of the remaining flight time on the cooperative guidance effect is analyzed, and a three-dimensional cooperative guidance model is constructed. Specifically:

[0098] Taking the fixed target point as the research background, a three-dimensional cooperative guidance model is constructed. Let (OXYZ) I , (OXYZ) LOS , (OXYZ) M be the ground inertial coordinate system, the line-of-sight coordinate system and the vehicle velocity coordinate system respectively, R, θ LOS , φ LOS be the relative distance between the missile and the target, the line-of-sight inclination angle and the line-of-sight deflection angle respectively, V m , θ m , φ m be the flight speed, the lead speed inclination angle and the lead speed deflection angle respectively. The above variables satisfy the following dynamic equations:

[0099]

[0100] In the formula are the two acceleration components of the vehicle in the normal direction, which are the control input variables of the guidance model and satisfy the following relationship:

[0101]

[0102] The field of view angle of the vehicle relative to the line-of-sight system is defined as σ m , and there is:

[0103] cosσ m =cosθ m cosφ m (3)

[0104] A multi - vehicle system composed of n agents has an undirected connected communication topology. Let the variable A = [a ij ∈ R n×n be the adjacency matrix of the multi - vehicle system. In the matrix elements, a ij = 1 indicates that member i can receive information from member j; otherwise, a ij = 0. The variable L = [l ij ∈ R n×n is the Laplace matrix of the multi - vehicle system, where the matrix element l ij = -a ij (i ≠ j), N i is the set of neighbor nodes of member i, satisfying N i = {j|a ij = 1, 1 ≤ j ≤ n};

[0105] Under the condition of a given target point, design the cooperative guidance command for the multi - vehicle so that the multi - vehicle can reach the target point simultaneously within a given time. Let the whole - journey flight time of the vehicle be T f,i , and define the remaining flight time t go,i of a single vehicle as follows, (1 ≤ i ≤ n), then there is:

[0106] t go,i = T f,i - t (4)

[0107] where t is the current flight time;

[0108] For a fixed target point, the remaining flight time of a single vehicle can be estimated by the following formula:

[0109]

[0110] where R i , σ m,i , V m,i are the missile - target distance, field - of - view angle, and flight speed of member i respectively, and N ~ [3, 7]

[0111] is the proportional navigation coefficient;

[0112] To achieve the cooperative arrival of the multi - vehicle at the desired target point, define the remaining flight time consistency error of member i as follows:

[0113]

[0114] Among them, in the above step two, based on the finite - time consistency theory, on the basis of the classical cooperative proportional navigation law, design a finite - time multi - missile cooperative guidance law applicable to a fixed target, specifically:

[0115] The classical cooperative proportional navigation law is as follows:

[0116]

[0117] Where N is the proportional navigation coefficient to be selected, then there is:

[0118]

[0119] The following relationship is obtained:

[0120]

[0121] Based on the small-angle assumption, taking the derivative of Equation (5) gives:

[0122]

[0123] Define the average remaining flight time of multiple aircraft at time t as The remaining flight time error of a single aircraft is At the same time, let

[0124]

[0125] The multi-missile cooperative guidance law is designed as follows:

[0126]

[0127] The stability of the multi-missile cooperative guidance law is proved by the following method:

[0128] Before proving the stability of the multi-missile cooperative guidance law, the following lemma is first given: For the nonlinear system If there exists a positive definite radially unbounded function V(x): x → R, satisfying the following inequality:

[0129]

[0130] Where a, b > 0, then the function V(x) converges to a small neighborhood of the origin in finite time, and the convergence time T f Satisfies:

[0131]

[0132] Based on the designed multi-missile cooperative guidance law in Equation (12), the following theorem is given: For a multi-aircraft system with undirected full connection, for a three-dimensional fixed target point, using the three-dimensional guidance model shown in Equation (1) and the cooperative guidance law shown in Equation (12), then the multi-aircraft will reach the desired target point in finite time;

[0133] Proof: Design the reference Lyapunov function Taking the derivative of it gives:

[0134]

[0135] From the definition of Equation (11) Δ j , the above equation can be further written as:

[0136]

[0137] where σ min,j is the minimum field of view angle of member j;

[0138] Let Then, according to the lemma, the function V converges to a small neighborhood of the 0 point within a finite time T f :

[0139]

[0140] and the theorem is proved;

[0141] Among them, in the above step three, based on the deep network, the proportional navigation coefficient adaptive adjustment of the multi-missile cooperative guidance law obtained in step two is realized, so as to optimize the consistency error of the remaining flight time of the cooperative guidance. Specifically: the multi-aircraft is modeled as multiple intelligent agent systems, and the remaining flight time t go,i of member i and the current total remaining flight vehicle time consistency error are used as the inputs of the deep neural network, and the current proportional navigation coefficient N i of member i is used as the network output. By designing certain reward and punishment rules and network training and learning, the amplitude of the proportional navigation coefficient is adaptively adjusted to complete the optimization of the consistency error of the remaining flight time of the cooperative guidance; based on Equation (12), the multi-missile cooperative guidance law optimized by the deep network is modified to:

[0142]

[0143] where is the proportional navigation coefficient after neural network optimization;

[0144] The deep network uses the DDPG algorithm as the training framework for the intelligent agent. In the DDPG algorithm, the Q value and V value of the intelligent agent are approximated by two deep neural networks. Among them, the network Q value is expressed as:

[0145] y t = r t + γQ(s t+1 , π(s t+1 |φ π ′)|θ Q ′) (19)

[0146] where θ′Q and φ π ′ are the hyperparameters of the target Q-value network and the V-value network;

[0147] In the DDPG algorithm, the following formula is used to optimize the Q-value network parameters:

[0148]

[0149] where N is the sample size;

[0150] In the DDPG algorithm, the following formula is used to optimize the action network parameters:

[0151]

[0152] The target network is updated in the following soft update manner:

[0153]

[0154] where τ is the soft update factor;

[0155] The action network Actor of the deep neural network is modeled as a multi-input multi-output MIMO system. Suppose the multi-agent system contains n members, then the neural network input space is designed as:

[0156] Input = [t go1 (t), t go,2 (t),..., t go,n (t)] T ∈R n (23)

[0157] The network output is designed as:

[0158] Output = [N 1 (t), N 2 (t),..., N n (t)] T ∈R n (24)

[0159] The comprehensive reward function of the deep neural network consists of the following three parts:

[0160] 1) Distance reward:

[0161]

[0162] where R(t - 1) and R(t) are the relative distances between the projectile and the target at times t - 1 and t respectively, is the magnitude of the velocity vector w 1 is the reward weight coefficient, with a value of 0.1;

[0163] 2) Remaining flight time error function reward:

[0164]

[0165] In the formula: t go,i is the current remaining flight time of member i, is the average remaining flight time of the current multi - vehicle system, w 2 is the reward weight coefficient, with a value of 0.05;

[0166] 3) Process constraint reward:

[0167]

[0168] In the formula: t max , R max are the maximum training step and the allowed maximum miss distance. The training end reward function is triggered when the aircraft successfully reaches the target point or other termination conditions occur. Other termination conditions include training step timeout and aircraft state singularity;

[0169] The comprehensive reward function is:

[0170]

[0171] Experimental example:

[0172] To verify the effectiveness of the method proposed in the embodiment, a training scenario is set. Taking the case where there are 4 aircraft members in the system as an example, the members are connected through an undirected communication topology, and the target point position is Target = [15, 15, 15] T km, and the initial state parameters of the 4 aircraft are:

[0173]

[0174] The parameter settings of the multi - missile cooperative guidance law (12) are:

[0175] c 1 = c 2 = 2, α 1 = 0.8 (30)

[0176] The proportional navigation coefficient of member i benchmark is N 0,i = 3, i = 1, 2, 3, 4, and the expression of the proportional navigation coefficient optimized by the deep network is:

[0177] N i (t) = N 0,i + ΔN i (t), ΔN i (t) ~ [0, 4] (31)

[0178] where ΔN i (t) is the output of the time-varying deep network, and its value ranges from 0 to 4;

[0179] The parameters of the DDPG algorithm (Deep Deterministic Policy Gradient) are set as follows:

[0180] Action network structure: 4 - 256 - 256 - 4; Evaluation network structure: 8 - 256 - 256 - 1; Learning rate of the action network: 3×10 -4 ; Learning rate of the evaluation network: 3×10 -4 ; Discount rate γ: 0.99; Soft update rate τ: 0.05; Sample quantity batchsize: 4*256; Size of the experience pool: 1×10 7 ; Training episodes: 200; By setting the neural network parameters and simulation scenarios, offline training of the multi - vehicle system can be achieved. In the present invention, 200 training episodes are set, and the converged policy network is saved after 200 training episodes for generating online guidance commands for the aircraft. Under the above - mentioned simulation scenarios, comparative analysis of the simulation results before and after optimization is carried out.

[0181] Based on the above, the advantages of the present invention are as follows. When the invention is used, by constructing a three - dimensional cooperative guidance model and a multi - missile cooperative guidance law, and by adopting the theory of deep reinforcement learning to realize the autonomous selection of guidance law coefficients, the consistency error of the remaining flight time of multi - vehicles is optimized. Deep reinforcement learning has the ability to explore the environment offline and give strategies online. Through a large number of offline simulations, a set of guidance parameter selection strategies adapted to the task environment is obtained. In specific applications, by only inputting information such as the current flight state of the missile group and the remaining flight time, the autonomous generation and online update of cooperative guidance commands can be realized. Within the range of the aircraft's range, by randomly initializing the initial state of the missile group and the position of the target point, the cooperative guidance of the missile group can be realized, and the multi - missile cooperative overload commands designed based on the finite - time consistency theory can enable the missile group to reach the expected target simultaneously within a finite time.

[0182] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above - mentioned exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms. Therefore, in any aspect, the embodiments should be regarded as exemplary and non - restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be regarded as limiting the claimed rights.

Claims

1. A three-dimensional adaptive multi-buzzer collaborative guidance method based on deep network optimization, including the following steps: Step 1: Establishment of a three-dimensional collaborative guidance model; Step 2: Design of a multi-missile collaborative guidance law; Step 3: Optimization of multi-missile collaborative guidance law parameters; It is characterized by: In the above step 1, a relative kinematic model of the missile group relative to the fixed target is established, the influence of the remaining flight time on the cooperative guidance effect is analyzed, and a three-dimensional cooperative guidance model is constructed; In the above step 2, based on the finite time consistency theory and on the basis of the classical cooperative proportional guidance law, a finite time multi-missile cooperative guidance law applicable to fixed targets is designed; In the above step 3, the proportional guidance coefficient of the multi-missile cooperative guidance law obtained in step 2 is adaptively adjusted based on the deep network, so as to optimize the consistency error of the cooperative guidance remaining flight time; In the third step, the specific one is to model the multi-aircraft into multiple agent systems, and the remaining flight time t of member i go,i The consistency error of the remaining aircraft time and the current time As input to the deep neural network, the current proportional navigation coefficient N of member i is N i As the network output, by designing certain reward and punishment rules and network training and learning, the amplitude of the proportional guidance coefficient is adaptively adjusted to optimize the consistency error of the remaining flight time of the cooperative guidance. Based on formula (12), the multi-missile cooperative guidance law optimized by the deep network is modified as follows: In the formula It is the proportional guidance coefficient after neural network optimization.

2. A three-dimensional adaptive multi-ejection collaborative guidance method based on deep network optimization according to claim 1, characterized in that: In the step 1, specifically: Taking the fixed target point as the research background, a three-dimensional collaborative guidance model is constructed. I , (OXYZ) LOS , (OXYZ) M They are the ground inertial coordinate system, the line of sight coordinate system and the aircraft velocity coordinate system. They are the relative distance between the projectile and the target, the inclination angle of sight and the deflection angle of sight. are respectively the flight speed, the leading speed inclination angle and the leading speed deviation angle. The above variables satisfy the following dynamic equation: In the formula The two acceleration components of the aircraft in the normal direction are the control input variables of the guidance model and satisfy the following relationship: The field of view angle of the aircraft relative to the line of sight is defined as σ m , and there are: A multi-aircraft system consisting of n intelligent agents has an undirected communication topology. Let variable A = [a ij ]∈R n×n is the adjacency matrix of the multi-aircraft system, and the matrix elements a ij =1 means member i can receive information from member j, otherwise a ij =0, variable L = [l ij ]∈R n×n is the Laplace matrix of the multi-aircraft system, where the matrix elements l ij =-a ij , i≠j, N i The set of neighbor nodes of a member satisfies N i ={j|a ij =1,1≤j≤n}; Under the given target point condition, design multi-aircraft collaborative guidance instructions so that multiple aircraft can reach the target point at the same time within a given time. Suppose the total flight time of the aircraft is T f,i , the remaining flight time t of a single aircraft is defined as follows go ,i,1≤i≤n, then: t go,i =T f,i -t (4) Where t is the current flight time; For a fixed target point, the remaining flight time of a single aircraft can be estimated by the following formula: Where R i ,σ m,i ,V m,i are the missile-target distance, field of view angle and flight speed of member i respectively, and N~[3,7] is the proportional guidance coefficient; In order to achieve the coordinated arrival of multiple aircraft at the desired target point, the remaining flight time consistency error of member i is defined as follows:

3. According to the three-dimensional adaptive multi-missile coordinated guidance method based on deep network optimization according to claim 1, it is characterized in that: In the step 2, specifically: The classic cooperative proportional guidance law is as follows: Where N is the proportional guidance coefficient to be selected, then: The following relationship is obtained: Based on the small angle assumption, the derivative of equation (5) is: Define the average remaining flight time of multiple aircraft at time t as The remaining flight time error of a single aircraft is At the same time, The multi-missile coordinated guidance law is designed as follows:

4. According to claim 3, a three-dimensional adaptive multi-missile coordinated guidance method based on deep network optimization is characterized in that: The stability of the multi-missile coordinated guidance law is proved by the following method: Before proving the stability of the multi-missile coordinated guidance law, we first give the following lemma: For a nonlinear system If there exists a positive definite radially unbounded function V(x):x→R that satisfies the following inequality: Where a, b>0, then the function V(x) converges to a small neighborhood of the origin in a finite time, and the convergence time T f satisfy: Based on the designed multi-missile cooperative guidance law of formula (12), the following theorem is given: For a multi-aircraft system with undirected full connection, for a three-dimensional fixed target point, using the three-dimensional guidance model shown in formula (1) and the cooperative guidance law shown in formula (12), the multi-aircraft can reach the desired target point within a limited time; Proof: Design reference Lyapunov function Taking its derivative we get: From formula (11)Δ j The definition of , the above formula can be further written as: Where σ mi n,j is the minimum field of view angle of member j; make Then, by the lemma, we can get that the function V has a finite time T f : If it converges to a small neighborhood of 0, the theorem is proved.

5. According to claim 1, a three-dimensional adaptive multi-missile coordinated guidance method based on deep network optimization is characterized in that: The deep network adopts the DDPG algorithm as the training framework of the agent. In the DDPG algorithm, the Q value and V value of the agent are approximated by two deep neural networks, where the network Q value is expressed as: y t =r t +γQ(s t+1 ,π(s t+1 |φ′ π )|θ′ Q ) (19) where θ′ Q and φ′ π are the hyperparameters of the target Q-value network and V-value network; In the DDPG algorithm, the following formula is used to optimize the Q value network parameters: Where N is the sample size; In the DDPG algorithm, the following formula is used to optimize the action network parameters: The target network is updated using the following soft update method: Where τ is the soft update factor.

6. The three-dimensional adaptive multi-missile coordinated guidance method based on deep network optimization according to claim 5 is characterized in that: The action network Actor of the deep neural network is modeled as a multi-input multi-output MIMO system. Assuming that the multi-agent system contains n members, the neural network input space is designed as: Input=[t go1 (t),t go,2 (t),...,t go,n (t)] T ∈R n (23) The network output is designed to be: Output=[N1(t),N2(t),...,N n (t)] T ∈R n (24)。 7. The three-dimensional adaptive multi-missile coordinated guidance method based on deep network optimization according to claim 5 is characterized in that: The comprehensive reward function of the deep neural network consists of the following three parts: 1) Distance Reward: Among them, R(t-1) and R(t) are the relative distances between the projectile and the target at time t-1 and t respectively. is the velocity vector magnitude w1 is the reward weight coefficient, which is 0.1; 2) Remaining flight time error function reward: Where: t go,i is the current remaining flight time of member i, is the average remaining flight time of the current multi-aircraft system, w2 is the reward weight coefficient, and its value is 0.05; 3) Process Constraint Rewards: Where: t max , R max For the maximum training step size and the maximum allowable off-target amount, the training end reward function is triggered when the aircraft successfully reaches the target point or other termination conditions appear.

8. The three-dimensional adaptive multi-missile coordinated guidance method based on deep network optimization according to claim 7 is characterized in that: The other termination conditions include training step timeout and vehicle state singularity.

9. The three-dimensional adaptive multi-missile coordinated guidance method based on deep network optimization according to claim 7 is characterized in that: The comprehensive reward function is:

Citation Information

Patent Citations

  • Three-dimensional multi-missile cooperative guidance method and system with finite time convergence

    CN106843265A

  • Variable proportionality coefficient multi-missile cooperative guidance method and system based on reinforcement learning

    CN117989923A