A decentralized collaborative guidance method based on strategy fusion and differential game
Through a decentralized collaborative guidance method based on strategy fusion and differential game, combined with adaptive dynamic programming and improved artificial potential field method, the collision avoidance and strike problems of missile clusters in high-intensity confrontation environments are solved, and efficient collaborative guidance of missile clusters is achieved.
Patent Information
- Application Number
- CN202411630006.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-15
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2044-11-15
AI Technical Summary
Existing guidance methods are difficult to meet the requirements of missile cluster combat in high-intensity confrontation environments in complex nonlinear systems, especially when the target can be actively identified and avoided. The probability of multi-missile cluster collision is high, and the existing ADP technology is insufficiently applied, failing to effectively solve the problem of missile-to-missile collision avoidance.
A decentralized collaborative guidance method based on strategy fusion and differential game is adopted. By integrating consensus and distance cost functions, a differential game guidance law is designed. The optimal control strategy is solved online using adaptive dynamic programming technology. Dynamic weight strategy fusion and improved artificial potential field method are introduced to solve the problem of inter-projectile collision.
It realizes intelligent collision avoidance and efficient strike of missile clusters in complex environments, simplifies the neural network structure, reduces the amount of calculation, improves the real-time and scalability of the system, and solves the problems of unknown target maneuverability and limited control input.
Smart Images

Figure CN119535974B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of aircraft guidance, and in particular to a decentralized collaborative guidance method based on strategy fusion and differential game. Background Art
[0002] As modern warfare becomes increasingly non-contact, precise, information-based, and systematic, the development of precision-guided weapons, the primary means of striking targets on the modern battlefield, and related technologies and equipment, has garnered significant attention worldwide. These technologies and equipment are flourishing at an unprecedented pace, playing an irreplaceable role. Among these, airborne guided bombs are the most numerous type of high-precision airborne weapon worldwide. Their key characteristics are their simple structure, ease of use, long range, high accuracy, low cost, and high cost-effectiveness. In-depth research into the development and application of airborne guided bomb swarms is an effective means of transforming the monolithic nature of future combat operations and enhancing the overall combat effectiveness of systems. It is a prerequisite and guarantee for seizing the initiative on future battlefields and winning future wars.
[0003] The performance of traditional single missiles is no longer sufficient to meet the strike requirements of today's complex air combat systems. In the 1970s, the United States proposed the concept of swarm-based coordinated operations, which holds significant military significance. Coordinated guidance of multiple missile clusters can significantly enhance the attack and defense capabilities of missiles under high-level confrontation conditions, enabling penetration of enemy missile interception systems and saturation attacks on targets, ensuring the successful destruction of key battlefield targets and improving combat effectiveness. Existing guidance methods, whether based on proportional guidance, optimal coordinated guidance, adaptive coordinated guidance, or sliding-mode coordinated guidance, all rely on the consistency of residual time to ensure simultaneous attacks by multiple missiles on stationary or slightly maneuvering targets. Consequently, these coordinated guidance design methods require pre-acquired maneuver information about the target, without considering the dynamic game between the missile and the target. Consequently, the resulting guidance laws are insufficient for combating high-speed, highly maneuverable intelligent targets.
[0004] Therefore, in order to address the uncertain dynamic characteristics between missile groups and between missiles and targets in the process of multi-missile coordinated guidance under non-ideal confrontation conditions, it is necessary to study a new coordinated guidance method with low requirements for inter-missile communication and excellent reliability, anti-destruction and robustness to meet the requirements of missile cluster combat in a high-intensity confrontation environment.
[0005] In addition, many researchers at home and abroad have attempted to linearize the system to solve the HJI equation in optimal control, but this is often difficult to achieve in practice. For complex nonlinear systems, the design of differential game cooperative guidance laws inevitably requires solving coupled nonlinear partial differential equations. In recent years, adaptive dynamic programming (ADP) technology has attracted considerable attention as an effective tool for solving optimal control problems for nonlinear systems. Its basic principle is to approximate a performance indicator function using a function approximation structure (such as a neural network). The parameters of the function approximation structure are then updated according to the Bellman optimality principle to obtain the optimal performance indicator function and the optimal control strategy. However, to date, the application of ADP technology in the aerospace field is limited, especially in the design of nonlinear differential game guidance laws. Furthermore, existing research has not considered the problem of collision avoidance between missiles. When targets can actively identify incoming missiles and evade them, the collision probability of multiple missiles in a swarm increases, significantly reducing their strike effectiveness. Therefore, it is necessary to break through the existing cooperative guidance design framework for complex nonlinear cooperative guidance systems and conduct research on the theory and methods of cooperative collision avoidance guidance based on differential games. Summary of the Invention
[0006] This paper aims to propose a decentralized collaborative guidance method based on strategy fusion and differential games. This method addresses the challenges of a multi-to-one missile pursuit and evasion system, such as limited control, target identification and evasion, high cluster computational complexity, and the high probability of inter-missile collisions. The paper first proposes a new decentralized scheme that represents the cost function by integrating consensus and distance. Considering control input constraints and unknown target maneuvers, a differential game guidance law is designed, combined with adaptive dynamic programming techniques, to solve the optimal control strategy for both the pursuer and the evader online. Next, a dynamic weighted strategy fusion technique is introduced to address the dilemma of multiple evasion strategies caused by paired pursuit and evasion control strategies. Finally, to address the collision issues that arise when coordinating the relative distance between missiles and targets, and the high probability of inter-missile collisions, an improved artificial potential field method is proposed. This method introduces a dynamic adjustment factor and solves the local minimum problem, achieving collision avoidance among the intelligent agents in the entire swarm system.
[0007] In order to achieve the above technical objectives, the technical solution adopted by the present invention is:
[0008] A decentralized collaborative guidance method based on strategy fusion and differential game theory, the missile distributed collaborative guidance method comprising the following steps:
[0009] S1: Establish a relative motion model of the projectile and the target, select the relative distance between the projectile and the target as the state variable, and obtain the nonlinear multi-agent system as follows:
[0010]
[0011] in, is the state variable of the i-th missile, is x i The derivative of x i,1 is the relative distance between the projectile and the target, x i,2 is the rate of change of the relative distance between the projectile and the target, n is the dimension of the system state variable, α i is the track angle of the i-th missile, θ i is the sight angle of the i-th missile, β is the target track angle, u i and υ T are the normal control inputs perpendicular to the velocity direction of the i-th missile and target, respectively, used to generate the lateral acceleration of the missile and target in the autopilot, is the derivative of the sight angle;
[0012] S2: For a single missile and target, we select the consistent formation error and miss distance as performance indicators, construct the performance indicator function of both parties in the game, and obtain the optimal control strategy based on the necessary conditions of the Nash-Pontryagin maximum-minimum principle. The time-varying HJI equation is then obtained as follows:
[0013]
[0014] Among them, J i * (x i ) is x i The optimal cost function, J i * (x i ) About x i The partial differential of Q(x i ) is a positive definite function, R s is a symmetric positive definite matrix, γ represents the restriction on the missile control input;
[0015] S3, using ADP technology to approximate the optimal cost function online, the network weight update law is obtained as:
[0016]
[0017] in, is the derivative of the ideal weight estimate, is the estimated value of the ideal weight, e c is the error function, is the partial derivative of the activation function σ(x) with respect to x, η is the learning rate, F1 and F2 are adjustment parameters greater than 0, J s is a smooth, radially unbounded differentiable function that satisfies It is a stabilization item;
[0018] S4, for the target, proposes a strategy fusion scheme, and combines the target escape strategy v1,v1,…,v N The optimal escape strategy obtained by vector fusion is:
[0019]
[0020] Among them, the correction weight coefficient ψ i (t) is the interest value of the target at time t for missile i, satisfying the relationship 0≤Ψ i (t)≤1 and N is the number of missiles;
[0021] S5: Use the artificial potential field method to deal with the collision problem, and use the relative distance between the missile and the target as the coordination variable to obtain the command acceleration u of the i-th missile i "for:
[0022] u″ i (t) = u i (t)+u′ i (t)
[0023] Among them, u i (t) is the command acceleration for the i-th missile to achieve coordination, u' i (t) is the command acceleration of missile i to avoid collision.
[0024] Furthermore, in step S1, the process of obtaining the nonlinear multi-agent system includes the following steps:
[0025] S11: Assume that the missile and the target are both point masses, and the velocity of the i-th missile is V i , the target speed is V T , is the derivative of the sight angle of the i-th missile, r i and are the relative distance between missile and target and the derivative of the relative distance between missile and target respectively, then the relative kinematic equation between the i-th missile and the target is:
[0026]
[0027] S12: Select state variables The state equation of the guidance system is:
[0028]
[0029] The feasible region of the system is N is the number of missiles, is x i1 The derivative of is x i2 The derivative of
[0030] S13: The state equation of the nonlinear system in step S12 is further described as a nonlinear multi-agent system: in u i and v T are the normal control inputs perpendicular to the velocity directions of the i-th missile and the target, respectively.
[0031] Furthermore, in step S2, the process of obtaining the time-varying HJI equation includes the following steps:
[0032] S21: For missiles, construct a continuous-time nonlinear differential game system:
[0033]
[0034] Among them, x i ∈R n is the state variable of the i-th missile, u i ∈R m and v T ∈R Z are the control inputs of both parties in the game, f(x i )∈R n , g(x i )∈R n×m and k(x i )∈R n×Z It is a locally Lipschitz continuous nonlinear function with f(0)=0, n, m, Z are the dimensions of the variables;
[0035] Three missiles form a triangle formation to pursue, and the distance between the missiles is r ij It can be expressed as:
[0036]
[0037] The missile's local consistency error δ i It can be expressed as:
[0038]
[0039] Where h(t) is the expected distance between missiles, a ij is the weight value of the communication between missiles, N i ={j:(i,j)∈E}, where E represents the communication set;
[0040] Taking the consistency formation error and miss amount as performance indicators, the performance indicator function J is selected i :
[0041]
[0042] Among them, λ1 and λ2 represent weight values, satisfying λ1+λ2=1, Q 1i (x i )=δ i T Q1δ i , Q 2i (x i )=x i T Q2x i , Q1 and Q2 are both positive definite functions, R v is a symmetric positive definite matrix, γ represents the restriction on our missile control input;
[0043] S22: Find the Nash equilibrium solution (u,v T ), so that the performance indicators meet:
[0044] J(u * ,v T )≤J(u * ,v T * )≤J(u,v T * )
[0045] Among them, u * ,v T * It represents the optimal control input of our missile and target;
[0046] The Hamiltonian function is defined as:
[0047]
[0048] in That is J i (x i ) for variable x i Find partial derivatives;
[0049] In the infinite time domain, the derivative of the optimal performance index with respect to time is 0, that is,
[0050] S23, according to the necessary conditions of the Nash-Pontryagin maximum-minimum principle and Get the optimal control strategy:
[0051]
[0052] in, u i * (x i ,t),v T * (x i ,t) represents the optimal control input perpendicular to the missile, i.e., the acceleration command;
[0053] Substituting the control strategy into the Hamiltonian function in step S22, the time-varying HJI equation is obtained:
[0054]
[0055] Furthermore, in step S3, the process of online approximating the optimal cost function using the ADP technology includes the following steps:
[0056] S31, using neural network to approximate the optimal cost function J online i * (x i ):
[0057]
[0058] Among them, W c ∈R L is the ideal weight vector of the neural network, σ(x i )∈R L is the activation function, ε(x i ) is the approximation error of the neural network;
[0059] Taking the partial derivative of the optimal cost function, we get:
[0060]
[0061] in, Denotes the activation function and the approximation error x i The partial derivative of
[0062] Substitute the partial derivative of the optimal cost function into the optimal control strategy shown in step S23:
[0063]
[0064] The time-varying HJI equation is obtained:
[0065]
[0066] The residual error caused by the neural network approximation is sorted out:
[0067]
[0068] S32, estimate the ideal weights based on the output of the neural network:
[0069]
[0070] in, and are the estimated values of the optimal cost function and ideal weights respectively;
[0071] right Taking the partial derivatives we get:
[0072]
[0073] Substituting the optimal control strategy into the optimal control strategy, the estimate of the optimal control strategy is:
[0074]
[0075] in, FA represents an estimate of the optimal control input perpendicular to the missile, i.e., the acceleration command;
[0076] Substituting the estimate of the optimal control strategy into the time-varying HJI equation in step S31, we obtain:
[0077]
[0078] S33, design neural network weight update law to optimize actual weights Let it infinitely approach the ideal weight W c , define the error function as:
[0079]
[0080] in, is the Hamiltonian function whose control input is the estimated optimal policy, is the Hamiltonian function whose control input is the optimal policy;
[0081] The loss function is defined as:
[0082]
[0083] The weight coefficient of the network is corrected using the gradient descent method:
[0084]
[0085] Among them, η>0 is the learning rate of the neural network;
[0086] Substituting into the error function we get:
[0087]
[0088] S34, assume that there exists a continuously differentiable radially unbounded Lyapunov function J s (x):
[0089]
[0090] in, Indicates J s The partial derivative of (x) with respect to x, there exists a positive definite function Λ(x) that satisfies and So that the following inequality holds:
[0091]
[0092] Assume that the closed-loop system is bounded under the differential game control strategy and the bounding function is an expression with respect to the variable x:
[0093]
[0094] Where c is a positive constant and c>0;
[0095] By adding two terms to formula (*), the network weight update law is proposed as follows:
[0096]
[0097] in, is a stabilization term, defined as:
[0098]
[0099] When the system is stable, at this time The stabilization item does not work. When the system is unstable, The stabilization item works to maintain system stability.
[0100] Furthermore, in step S4, a strategy fusion solution is proposed for the target, and the process of obtaining the target's optimal escape strategy includes the following steps:
[0101] S41, consider the strategy fusion scheme as the target escape strategy v1,v1,...,v N The resultant vector is expressed as:
[0102]
[0103] Among them, the modified weight coefficient ψ i (t) represents the missile's interest in the target, satisfying 0≤ψ i (t)≤1 and N is the number of missiles;
[0104] S42, define ψ i 1 (t) is the distance interest function S r and the speed interest function S V The sum is expressed as:
[0105] ψ i 1 (t) = S r +S v
[0106] in, r i is the relative distance between the projectile and the target, r c is the escape radius of the target, V i is the speed of the i-th missile, V T The speed of the target;
[0107] S43, get the weight coefficient ψ i (t) specific representation:
[0108]
[0109] Furthermore, in step S5, the collision problem is processed by the artificial potential field method, and the relative distance between the missile and the target is used as the coordination variable. The process of obtaining the command acceleration of the i-th missile includes the following steps:
[0110] S51, assuming that the artificial potential field function is represented by U(x,y), and that U(x,y) is differentiable for every position in space, the force acting on the missile is the control force F(x,y) in the direction of the negative gradient of the artificial potential field function, expressed as:
[0111]
[0112] Among them, (x, y) represents the coordinates of the missile in the two-dimensional plane, Represents the gradient of U(x,y) at (x,y), and its direction is the direction of the maximum rate of change of the potential field at the coordinate (x,y);
[0113] The artificial potential field function between missiles i and j is U ij , which satisfies:
[0114]
[0115] Among them, r ij represents the distance between missiles, n represents the number of missiles involved in the coordinated operation, It means that missile i is controlled by the artificial potential field of missile j, It means that missile j is controlled by the artificial potential field of missile i, U iIt represents the sum of the artificial potential field control forces exerted on missile i by the rest of the missiles;
[0116] S52, artificial potential field function U determined by the relative distance between projectiles ij (r ij )for:
[0117]
[0118] Among them, the coefficient K r Adjust the potential field strength, r min is the minimum safety radius between missile i and missile j, r s is the anti-collision threshold, when r ij >r s When there is no artificial potential field between the bullets, when r ij →r min When U ij →∞;
[0119] Combined with the artificial potential field control force, the control force F exerted on missile i in the potential field of missile j is ij (x i ,y i )for:
[0120]
[0121] Among them, (x i ,y i ) represents the coordinates of missile i in the two-dimensional plane.
[0122] The control force on missile i and F i (x i ,y i )for:
[0123]
[0124] S53, control force F i The command acceleration u′ caused by i for:
[0125]
[0126] Among them, m i is the mass of the missile, F ia is the control force F i Normal projection of velocity;
[0127] Combined with the cooperative guidance law, the command acceleration of the missile is obtained as:
[0128] u″ i =u i +u′ i
[0129] Among them, u i is the normal control input perpendicular to the velocity direction of the i-th missile. min <r ij ≤r s When the missile is at u i and u i 'When flying under the simultaneous action of the two commands, when the two commands are equal in size but opposite in direction at a certain moment, the missile will fall into a local minimum state;
[0130] S54, design a dynamic adjustment factor k based on the distance between bullets d To prevent the missile from falling into the local minimum:
[0131]
[0132] in, r dom is a random number, satisfying 0≤r dom ≤1, is the gain factor, and at the same time, -1<α d <1;
[0133] After the improvement, the control force on missile i is:
[0134]
[0135] When in r min <r ij ≤r s When the missile is at u i and u i 'Flying under the action of i +u′ i =0, k d Start dynamic adjustment of u′ i The value of r ij ≤r min When the potential field force is infinite, the command u i It doesn't work. The missile is mainly used for collision avoidance.
[0136] Compared with the prior art, the present invention has the following beneficial effects:
[0137] First, the decentralized collaborative guidance method based on strategy fusion and differential game of the present invention takes into account the scenario where our control is limited, the target maneuver is unknown, and the incoming air-guided bombs (missiles) can be actively identified and avoided. A differential game guidance law based on ADP is designed to predict the target maneuver.
[0138] Second, the decentralized collaborative guidance method based on strategy fusion and differential game theory uses a neural network to approximate the solution of the HJI equation online. Unlike the traditional evaluation network-execution network structure, this paper uses a single evaluation network to approximate the optimal cost function online, simplifying the approximate structure of the neural network.
[0139] Third, the decentralized collaborative guidance method based on strategy fusion and differential games addresses the challenges of limited control input and unknown target maneuvers. This method uses differential game theory to enable missiles to attack targets in predetermined formations. By combining differential game theory with consistent formation theory, a cost function is proposed to achieve decentralization. This method eliminates the need to estimate the remaining time of moving targets and offers advantages such as a simple structure, good real-time performance, low communication overhead, and excellent scalability.
[0140] Fourth, the decentralized collaborative guidance method based on strategy fusion and differential game of the present invention uses an improved artificial potential field method to introduce a dynamic adjustment factor to solve the collision problem caused by coordinating the relative distance between the missile and the target, thereby solving the local minimum problem and realizing collision avoidance among the intelligent agents of the entire cluster system. BRIEF DESCRIPTION OF THE DRAWINGS
[0141] Figure 1 This is an overall flow chart of a decentralized collaborative guidance method based on strategy fusion and differential game.
[0142] Figure 2 This is the system structure block diagram based on the artificial potential field method.
[0143] Figure 3 Schematic diagram of the relative motion model between projectile and target.
[0144] Figure 4 This is a diagram of the formation of the cluster system. DETAILED DESCRIPTION
[0145] The embodiments of the present invention are described in further detail below with reference to the accompanying drawings.
[0146] See also Figure 1 and Figure 2 The present invention discloses a decentralized collaborative guidance method based on strategy fusion and differential game, comprising the following steps:
[0147] Step 1: Establish a relative motion model between the projectile and the target, select the relative distance between the projectile and the target as the state variable, and obtain a nonlinear multi-agent system;
[0148] Consider the problem of coordinated pursuit of a maneuvering target by N missiles in a two-dimensional plane. The missiles seek the optimal control strategy to pursue the maneuvering target T, while the target T tries to seek the optimal control strategy to avoid the attack. The N missiles achieve coordination through formation, such as Figure 3 、 Figure 4 shown.
[0149] The research of this method is based on the final guidance stage, when the missile speed remains unchanged and the direction of motion is changed by changing the magnitude of the normal acceleration. Assume that the missile and the target are both point masses, and the speed of the i-th missile is V i , the track angle is α i , the sight angle is θ i The target's velocity is V T , the track angle is β. u i and υ T are the normal control inputs perpendicular to the missile and target velocity directions, respectively, used to generate the lateral acceleration of the missile and target in the autopilot, r i and is the relative distance between the missile and the target and its derivative. Then the relative kinematic equation between the i-th missile and the target is:
[0150]
[0151] Selecting state variables Therefore, the state equation of the guidance system is:
[0152]
[0153] Among them, |α i -θ i When |=0 and |β-θ|=0, the system is uncontrollable, which is the point that causes the system to be unstable. i →0, the state equation tends to infinity and the guidance process is destroyed. Therefore, the feasible region of the system is But in reality, in the final stage of terminal guidance, the missile continues to move only by inertia. i If the radius is smaller than the capture radius, it is considered a successful capture. Therefore, it does not affect the pursuit accuracy and escape effect of both parties.
[0154] Formula (2) can be further described as a nonlinear multi-agent system in the following form:
[0155]
[0156] in u i and v T is the control input of the system.
[0157] Step 2: For the missile and target, select the consistent formation error and miss distance as performance indicators and construct the performance indicator function for both parties in the game. According to the necessary conditions of the Nash-Pontryagin maximum-minimum principle, the optimal control strategy is obtained, and then the time-varying HJI equation is obtained.
[0158] For missiles, a continuous-time nonlinear differential game system is constructed:
[0159]
[0160] Among them, x i ∈R n is the state variable, u i ∈R m and v T ∈R Z is the control input of both parties in the game, f(x i )∈R n , g(x i )∈R n×m and k(x i )∈R n×Z is a locally Lipschitz continuous nonlinear function, n, m, z are the dimensions of the variables, and f(0) = 0. Obviously, the system (4) is controllable, and x = 0 is the equilibrium point of the system.
[0161] Three missiles form a triangle formation for pursuit. The distance between the missiles can be expressed as:
[0162]
[0163] The local consistency error of the missile can be expressed as:
[0164]
[0165] Where h(t) is the expected distance between missiles, a ij is the weight value of the communication between missiles, N i ={j:(i,j)∈E}, where E represents the communication set.
[0166] Taking the consistent formation error and miss distance as performance indicators, the performance indicator function is selected as:
[0167]
[0168] Among them, λ1 and λ2 represent weight values, satisfying λ1+λ2=1, Q 1i (x i )=δ i T Q1δ i , Q 2i (x i )=x i T Q2x i , Q1 and Q2 are both positive definite functions, R v is a symmetric positive definite matrix, γ represents the restriction on our missile control input;
[0169] Since our control strategy u hopes to minimize the amount of miss, the enemy control strategy v T In order to maximize the miss distance, a zero-sum differential game is formed. In the end, both the chasing and the escaping parties need to find a Nash equilibrium solution (u, v T ),make:
[0170] J(u * ,v T )≤J(u * ,v T * )≤J(u,v T * ) Formula (8)
[0171] The Hamiltonian function is defined as:
[0172]
[0173] in That is J i (x i ) for variable x i Find the partial derivative; in the infinite time domain, the derivative of the optimal performance index with respect to time is 0, that is, Combined with the boundary conditions, we can get:
[0174]
[0175] According to the necessary conditions of the Nash-Pontryagin maximum-minimum principle and The optimal control strategy is obtained as:
[0176]
[0177] in,
[0178] Substituting the control strategy (11) into (9), we get the time-varying HJI equation:
[0179]
[0180] in By solving the HJI equation, we obtain that since the nonlinear partial differential equation is difficult to solve, the ADP technology is used to approximate the optimal cost function online.
[0181] Step 3: Since nonlinear partial differential equations are difficult to solve, the optimal cost function is approximated online using the ADP technique. To minimize the square of the error and keep the signal bounded during closed-loop learning, a simplified network weight update law is proposed.
[0182] Assumption 1 Considering the affine nonlinear system (4) and the given performance index function (7), under the action of the differential game control strategy (11), it is assumed that there exists a continuously differentiable radially unbounded Lyapunov function J s (x), so that the inequality Established. Among them, Indicates J s (x) with respect to x. Now we can determine that there exists a positive definite function Λ(x) that satisfies and So that the following inequality holds:
[0183]
[0184] In addition, it is assumed that the closed-loop system is bounded under the differential game control strategy and the bounding function is an expression with respect to the variable x:
[0185]
[0186] Among them, c>0 is a positive constant.
[0187] Using neural network to approximate the optimal cost function J online * (x):
[0188]
[0189] Among them, W c ∈R L is the ideal weight vector of the neural network, σ(x i )∈R L is the activation function, ε(x i ) is the approximation error of the neural network.
[0190] Taking partial derivative of formula (15), we get:
[0191]
[0192] in, represents the partial derivative of the activation function with respect to x.
[0193] Substituting equation (16) into equation (11), the optimal control strategy is:
[0194]
[0195] Substituting equation (16) into equation (12), we get the HJI equation:
[0196]
[0197] The residual error caused by the neural network approximation is sorted out:
[0198]
[0199] Since the ideal weights are unknown, the output of CNN will be used to estimate the ideal weights. The specific expression is as follows:
[0200]
[0201] in, and Represents an estimate of the optimal cost function and ideal weights.
[0202] Taking partial derivative of formula (20) we get:
[0203]
[0204] Substituting into formula (11), we can get the estimate of the optimal control strategy:
[0205]
[0206] Substituting equation (22) into equation (12), we get the HJI equation:
[0207]
[0208] It is necessary to design a CNN weight update law to optimize the actual weights Let it infinitely approach the ideal weight W c Since the ideal weights are unknown, the error function is defined as:
[0209]
[0210] Define the loss function:
[0211]
[0212] Using the gradient descent method to correct the network weight coefficient, we have
[0213]
[0214] Among them, η>0 represents the learning rate of CNN.
[0215] Substituting equation (24) into equation (26) yields
[0216]
[0217] Due to the actual weight The stability of nonlinear systems cannot be guaranteed, so a simplified network weight update law is proposed, which can not only minimize the square of the error but also ensure that the signal is bounded during the closed-loop system learning process. Two terms are added to the basis of formula (27) to reduce the approximation error of the neural network and ensure the stability of the system, and we get
[0218]
[0219] in, F1 and F2 are adjustment parameters greater than 0; This has been given in hypothesis 1; is defined as follows:
[0220]
[0221] When the system (4) is stable, at this time The stabilization item does not work. When the system is unstable, The stabilization item works to maintain system stability.
[0222] Defining CNN weight estimation error According to (18), we can get:
[0223]
[0224] Combining equations (24), (28) and (30), the CNN weight estimation error dynamics can be expressed as:
[0225]
[0226] Step 4: Propose a strategy fusion solution for the target. The process of obtaining the optimal escape strategy for the target includes the following steps:
[0227] The strategy fusion scheme can be regarded as the target escape strategy v1,v1,...,v N The resultant vector is expressed as:
[0228]
[0229] Where N is the number of missiles, and the corrected weight coefficient ψ i (t) represents the missile's interest in the target, satisfying 0≤ψ i (t)≤1 and:
[0230]
[0231] Define ψ i 1 (t) is the distance interest function S r and the speed interest function S VThe sum is expressed as:
[0232] ψ i 1 (t) = S r +S v Formula (34)
[0233] in, r c is the target’s escape radius;
[0234] Get the weight coefficient ψ i The specific expression of (t) is:
[0235]
[0236] Step 5: Use the artificial potential field method to deal with the problem of missile collisions when the relative distance between the missile and the target is used as the coordination variable in the distributed coordination process. Obtain the command acceleration of the missile;
[0237] Assumption 2: The artificial potential field function is represented by U(x,y), which must be differentiable for every position in space.
[0238] The artificial potential field method determines the direction and speed of the missile by searching for the direction in which the potential field value decreases. Therefore, the force acting on the missile is the control force in the direction of the negative gradient of the artificial potential field function, which is expressed as:
[0239]
[0240] in, Represents the gradient of U(x,y) at (x,y), which is a vector whose direction is the direction of the maximum rate of change of the potential field at position (x,y).
[0241] The artificial potential field function between missile i and missile j in the system is U ij , which has the following properties:
[0242] <1> The potential field function has symmetry, namely:
[0243]
[0244] <2> During the flight, missile i may be in the potential field of multiple missiles around it, so the total potential field U i It can be expressed as:
[0245]
[0246] Where r ij represents the distance between missiles, and n represents the number of missiles participating in the coordinated operation.
[0247] like Figure 2 This is the system structure block diagram based on the artificial potential field method.
[0248] The artificial potential field function determined by the relative distance between projectiles is defined as:
[0249]
[0250] Among them, the coefficient K r Adjust the potential field strength, r min is the minimum safety radius between missile i and missile j, r s is the anti-collision threshold, when there is no artificial potential field between the bullets. ij →r min When U ij →∞.
[0251] Therefore, combined with formula (36), the control force on missile i in the potential field of missile j is:
[0252]
[0253] Furthermore, according to formula (38), the control force on missile i is:
[0254]
[0255] Control force F i The resulting command acceleration is:
[0256]
[0257] Among them, m i is the mass of the missile, F ia is the control force F i In the normal projection of the velocity, the normal control input changes the velocity heading angle of the missile, thereby changing its direction of movement without changing the magnitude of the velocity.
[0258] Combining the consistency control algorithm and artificial potential field collision avoidance algorithm in the previous section, the command acceleration of the missile can be obtained as:
[0259] u i ”=u i +u i Formula (43)
[0260] When in r min <r ij ≤r s When the missile is at u i and u i 'When flying under the simultaneous action of two commands, when the two commands are equal in size but opposite in direction at a certain moment, the missile will fall into a local minimum state.
[0261] Design a dynamic adjustment factor k based on the distance between bullets d To prevent the missile from falling into the local minimum:
[0262]
[0263] in, r dom is a random number, satisfying 0≤r dom ≤1, is the gain factor; at the same time, -1<α d <1.
[0264] After the improvement, the control force on missile i is:
[0265]
[0266] When in r min <r ij ≤r s When the missile is at u i and u i 'Flying under the action of i +u i When '=0, k d Start dynamic adjustment u i ' value, when r ij ≤r min When the potential field force is infinite, the command u i It doesn't work. The missile is mainly used for collision avoidance.
[0267] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code. The scheme in the embodiment of the present application can be implemented in various computer languages, for example, object-oriented programming language Java and literal translation scripting language JavaScript, etc.
[0268] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0269] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0270] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions for executing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0271] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.
[0272] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.
Claims
1. A decentralized collaborative guidance method based on strategy fusion and differential game, characterized by: The method comprises the following steps: S1: Establish a relative motion model of the projectile and the target, select the relative distance between the projectile and the target as the state variable, and obtain the nonlinear multi-agent system as follows: in, is the state variable of the i-th missile, is x i The derivative of x i,1 is the relative distance between the projectile and the target, x i,2 is the rate of change of the relative distance between the projectile and the target, n is the dimension of the system state variable, α i is the track angle of the i-th missile, θ i is the sight angle of the i-th missile, β is the target track angle, u i and v T are the normal control inputs perpendicular to the velocity direction of the i-th missile and target, respectively, used to generate the lateral acceleration of the missile and target in the autopilot, is the derivative of the sight angle; S2: For a single missile and target, we select the consistent formation error and miss distance as performance indicators, construct the performance indicator function of both parties in the game, and obtain the optimal control strategy based on the necessary conditions of the Nash-Pontryagin maximum-minimum principle. The time-varying HJI equation is then obtained as follows: Among them, J i * (x i ) is x i The optimal cost function, J i * (x i ) About x i The partial differential of Q(x i ) is a positive definite function, R s is a symmetric positive definite matrix, γ represents the restriction on the missile control input; S3, using ADP technology to approximate the optimal cost function online, the network weight update law is obtained as: in, is the derivative of the ideal weight estimate, is the estimated value of the ideal weight, e c is the error function, ▽σ(x) is the partial derivative of the activation function σ(x) with respect to x, η is the learning rate, F1 and F2 are adjustment parameters greater than 0, J s is a smooth, radially unbounded differentiable function that satisfies It is a stabilization item; S4: Propose a strategy fusion solution for the target, and combine the target escape strategy v1,v1,…,v N The optimal escape strategy obtained by vector fusion is: Among them, the correction weight coefficient ψ i (t) is the interest value of the target at time t for missile i, satisfying the relationship 0≤Ψ i (t)≤1 and N is the number of missiles; The step S4 comprises: S41, consider the strategy fusion scheme as the target escape strategy v1,v1,...,v N The resultant vector is expressed as: Among them, the modified weight coefficient ψ i (t) represents the missile's interest in the target, satisfying 0≤ψ i (t)≤1 and N is the number of missiles; S42, define ψ i 1 (t) is the distance interest function S r and the speed interest function S V The sum is expressed as: ψ i 1 (t) = S r +S v Formula (6) in, r i is the relative distance between the projectile and the target, r c is the escape radius of the target, V i is the speed of the i-th missile, V T The speed of the target; S43, get the weight coefficient ψ i (t) specific representation: S5: Use the artificial potential field method to deal with the collision problem, and use the relative distance between the missile and the target as the coordination variable to obtain the command acceleration u of the i-th missile i "for: u i ”(t) = u i (t) + u i '(t) Equation (8) Among them, u i (t) is the command acceleration for the i-th missile to achieve coordination, u i '(t) is the command acceleration of missile i to avoid collision.
2. The decentralized collaborative guidance method based on strategy fusion and differential game as claimed in claim 1, characterized in that: The step S1 comprises: S11: Assume that the missile and the target are both point masses, and the velocity of the i-th missile is V i , the target speed is V T , is the derivative of the sight angle of the i-th missile, r i and are the relative distance between missile and target and the derivative of the relative distance between missile and target respectively, then the relative kinematic equation between the i-th missile and the target is: S12: Select state variables The state equation of the guidance system is: The feasible region of the system is N is the number of missiles, is x i1 The derivative of is x i2 The derivative of S13: The state equation of the nonlinear system in step S12 is further described as a nonlinear multi-agent system: in u i and v T are the normal control inputs perpendicular to the velocity directions of the i-th missile and the target, respectively.
3. The decentralized collaborative guidance method based on strategy fusion and differential game as claimed in claim 1, characterized in that: The step S2 comprises: S21: For missiles, construct a continuous-time nonlinear differential game system: Among them, x i ∈R n is the state variable of the i-th missile, u i ∈R m and v T ∈R Z are the control inputs of both parties in the game, f(x i )∈R n , g(x i )∈R n×m and k(x i )∈R n×Z It is a locally Lipschitz continuous nonlinear function with f(0)=0, n, m, Z are the dimensions of the variables; Three missiles form a triangle formation to pursue, and the distance between the missiles is r ij It can be expressed as: The missile's local consistency error δ i It can be expressed as: Where h(t) is the expected distance between missiles, a ij is the weight value of the communication between missiles, N i ={j:(i,j)∈E}, where E represents the communication set; Taking the consistency formation error and miss amount as performance indicators, the performance indicator function J is selected i : Among them, λ1 and λ2 represent weight values, satisfying λ1+λ2=1, Q 1i (x i )=δ i T Q1δ i , Q 2i (x i )=x i T Q2x i , Q1 and Q2 are both positive definite functions, R v is a symmetric positive definite matrix, γ represents the restriction on our missile control input; S22: Find the Nash equilibrium solution (u,v T ), so that the performance indicators meet: J(u * , v T ) ≤ J(u * , v T * ) ≤ J(u, v T * ) Equation (15) Among them, u * ,v T * It represents the optimal control input of our missile and target; The Hamiltonian function is defined as: in That is J i (x i ) for variable x i Find partial derivatives; In the infinite time domain, the derivative of the optimal performance index with respect to time is 0, that is, S23, according to the necessary conditions of the Nash-Pontryagin maximum-minimum principle and Get the optimal control strategy: in, u i * (x i ,t),v T * (x i ,t) represents the optimal control input perpendicular to the missile, i.e., the acceleration command; Substituting the control strategy into the Hamiltonian function in step S22, the time-varying HJI equation is obtained:
4. The decentralized collaborative guidance method based on strategy fusion and differential game as claimed in claim 1, characterized in that: The step S3 comprises: S31, using neural network to approximate the optimal cost function J online i * (x i ): Among them, W c ∈R L is the ideal weight vector of the neural network, σ(x i )∈R L is the activation function, ε(x i ) is the approximation error of the neural network; Taking the partial derivative of the optimal cost function, we get: ▽J i * (x i )=▽(σ(x i )) T W c +▽ε(x i ) Formula (20) Among them, ▽σ(x i ),▽ε(x i ) represents the activation function and the approximation error x i The partial derivative of Substitute the partial derivative of the optimal cost function into the optimal control strategy shown in equation (14) in step S23: The time-varying HJI equation is obtained: The residual error caused by the neural network approximation is sorted out: S32, estimate the ideal weights based on the output of the neural network: in, and are the estimated values of the optimal cost function and ideal weights respectively; Taking the partial derivative of formula (21) we get: Substituting into equation (14), we can get the estimate of the optimal control strategy: in, FA represents an estimate of the optimal control input perpendicular to the missile, i.e., the acceleration command; Substituting the estimate of the optimal control strategy into the time-varying HJI equation in step S31, we obtain: S33, design neural network weight update law to optimize actual weights Let it infinitely approach the ideal weight W c , define the error function as: in, The Hamiltonian function whose control input is the estimated optimal policy, H(x,u * ,v T * ,▽J i * (x i )) is the Hamiltonian function whose control input is the optimal strategy; The loss function is defined as: The weight coefficient of the network is corrected using the gradient descent method: Among them, η>0 is the learning rate of the neural network; Substituting into the error function we get: S34, assume that there exists a continuously differentiable radially unbounded Lyapunov function J s (x): Among them, ▽J s (x) represents J s The partial derivative of (x) with respect to x, there exists a positive definite function Λ(x) that satisfies and So that the following inequality holds: (▽J s (x)) T (f(x)+g(x)u * +k(x)v T * ) < -(▽J s (x)) T Λ(x)▽J s (x) Equation (33) Assume that the closed-loop system is bounded under the differential game control strategy and the bounding function is an expression with respect to the variable x: Where c is a positive constant and c>0; By adding two terms to formula (28), the network weight update law is proposed as follows: in, is a stabilization term, defined as: J(x) is the cost function. When the system is stable, at this time The stabilization item does not work. When the system is unstable, The stabilization item works to maintain system stability.
5. The decentralized collaborative guidance method based on strategy fusion and differential game as claimed in claim 1, characterized in that: The step S5 comprises: S51, assuming that the artificial potential field function is represented by U(x,y), and that U(x,y) is differentiable for every position in space, the force acting on the missile is the control force F(x,y) in the direction of the negative gradient of the artificial potential field function, expressed as: F(x,y)=-▽U(x,y) Formula (37) Where (x, y) represents the coordinates of the missile in the two-dimensional plane, ▽U(x, y) represents the gradient of U(x, y) at (x, y), and its direction is the direction of the maximum rate of change of the potential field at the coordinate (x, y); The artificial potential field function between missiles i and j is U ij , which satisfies: Among them, r ij represents the distance between missiles, n represents the number of missiles involved in the coordinated operation, ▽U ij (r ij ) indicates that missile i is subject to the artificial potential field control force of missile j, ▽U ji (r ij ) indicates that missile j is controlled by the artificial potential field of missile i, U i It represents the sum of the artificial potential field control forces exerted on missile i by the rest of the missiles; S52, artificial potential field function U determined by the relative distance between projectiles ij (r ij )for: Among them, the coefficient K r Adjust the potential field strength, r min is the minimum safety radius between missile i and missile j, r s is the anti-collision threshold, when r ij >r s When there is no artificial potential field between the bullets, when r ij →r min When U ij →∞; Combined with the artificial potential field control force, the control force F exerted on missile i in the potential field of missile j is ij (x i ,y i )for: Among them, (x i ,y i ) represents the coordinates of missile i in the two-dimensional plane; The control force on missile i and F i (x i ,y i )for: S53, control force F i The command acceleration u caused by i 'for: Among them, m i is the mass of the missile, F ia is the control force F i Normal projection of velocity; Combined with the cooperative guidance law, the command acceleration of the missile is obtained as: u i ” = u i + u i ' Equation (43) Among them, u i is the normal control input perpendicular to the velocity direction of the i-th missile. min <r ij ≤r s When the missile is at u i and u i 'When flying under the simultaneous action of the two commands, when the two commands are equal in size but opposite in direction at a certain moment, the missile will fall into a local minimum state; S54, design a dynamic adjustment factor k based on the distance between bullets d To prevent the missile from falling into the local minimum: in, r dom is a random number, satisfying 0≤r dom ≤1, is the gain factor, and at the same time, -1<α d <1; After the improvement, the control force on missile i is: When in r min <r ij ≤r s When the missile is at u i and u i 'Flying under the action of i +u i When '=0, k d Start dynamic adjustment u i ' value, when r ij ≤r min When the potential field force is infinite, the command u i It doesn't work. The missile is mainly used for collision avoidance.
Citation Information
Patent Citations
Missile non-escapable area rapid calculation method based on K-sparse self-encoding SVM (Support Vector Machine)
CN116011315A
Policy decision-making method in multi-agent collaboration and confrontation
CN116167415A