Cooperative control method of hypersonic vehicle based on adaptive dynamic programming
By combining adaptive dynamic programming and single-evaluation neural networks, the optimal tracking problem in the cooperative control of hypersonic vehicle swarms was solved, achieving efficient cooperative control in complex environments.
Patent Information
- Application Number
- CN202411964823.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2044-12-30
AI Technical Summary
Existing technologies cannot effectively solve the problem of cooperative optimal tracking and control of hypersonic vehicles in multi-vehicle cooperative flight scenarios, especially under model uncertainty and environmental distortion, it is difficult to achieve efficient cooperative control.
An adaptive dynamic programming algorithm is used to establish a 6-DOF normalized model of a hypersonic vehicle, and a single-evaluation neural network online controller is constructed. By solving the optimal tracking control law online, the cooperative tracking control of a hypersonic vehicle cluster is realized.
Considering flight constraints, cooperative tracking control of hypersonic vehicle clusters was achieved, improving control accuracy and efficiency in complex environments.
Smart Images

Figure CN119847200B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of cooperative control technology for hypersonic vehicles, and specifically relates to a cooperative control method for hypersonic vehicles based on adaptive dynamic programming. Background Technology
[0002] Hypersonic vehicles, due to their extremely high speed and agile maneuverability, hold significant importance for the scientific field. They are not only suitable for atmospheric cruise but also capable of reentering orbit after atmospheric reentry, effectively combining the advantages of aircraft and spacecraft. However, the motion models of hypersonic vehicles exhibit faster time differences, stronger coupling, and greater uncertainties than existing aircraft, bringing unprecedented difficulties and challenges to control system design. Therefore, existing design theories and methods for aircraft control systems cannot be directly transferred to hypersonic vehicles. It is necessary to research new cooperative control theories and methods without sacrificing the versatility of hypersonic vehicles.
[0003] Adaptive dynamic programming (ADMP) is an emerging optimal control strategy that effectively solves the curse of dimensionality problem inherent in existing optimal control algorithms and enables adaptive online control of targets under more severe model uncertainties and environmental distortions. ADMP is of great significance for solving the speed or attitude control problems of hypersonic vehicles, but it has received insufficient attention in the cooperative optimal tracking control problem of multi-vehicle cooperative flight scenarios. Therefore, how to apply ADMP-based online control algorithms to the cooperative tracking control of hypersonic vehicles is a pressing issue that needs to be addressed. Summary of the Invention
[0004] To address the aforementioned technical problems, this invention provides a cooperative control method for hypersonic vehicles based on adaptive dynamic programming. It uses an adaptive dynamic programming algorithm with a single-evaluation neural network architecture to calculate the optimal cooperative control input for the vehicle, and is applicable to cooperative tracking control based on adaptive dynamic programming in the field of cooperative control of hypersonic vehicles.
[0005] The technical solution adopted in this invention is: a cooperative control method for hypersonic vehicles based on adaptive dynamic programming, the specific steps of which are as follows:
[0006] S1. Establish a 6-DOF normalized model for hypersonic vehicles and clarify the cooperative tracking and control problem of hypersonic vehicle swarms.
[0007] S2. Based on step S1, use the adaptive dynamic programming algorithm to obtain the optimal tracking control form under the constraint of control input;
[0008] S3. Based on step S2, construct a single-evaluation neural network online controller to solve the optimal tracking control law online;
[0009] S4. The optimal tracking control law obtained in step S3 is applied to the hypersonic vehicle cluster to achieve cooperative tracking control.
[0010] Furthermore, step S1 is specifically as follows:
[0011] S11. Establish a 6-DOF normalized model for hypersonic vehicles;
[0012] The equation of motion for a 6-DOF hypersonic vehicle without power takeoff and landing is as follows:
[0013]
[0014] Where r represents the radial distance from the Earth's center to the hypersonic vehicle, t represents time; θ and φ represent longitude and latitude, respectively; V represents the Earth's relative velocity; ψ represents the velocity heading angle relative to local north; γ represents the track angle; m and σ represent the hypersonic vehicle's mass and roll angle, respectively; D and L represent aerodynamic drag and lift, respectively. The Earth's angular rate is denoted as ω. e The local gravitational acceleration is denoted as g.
[0015] The dynamic equations of a 6-DOF particle are normalized, and the normalized variable expressions are defined as follows:
[0016]
[0017] Among them, R e Let g represent the Earth's radius, and g0 represent gravitational acceleration.
[0018] Substituting the normalized variables into the dynamic equations, we obtain the normalized dynamic equations for the hypersonic vehicle, as shown below:
[0019]
[0020] S12. Clarify the cooperative tracking and control issues of hypersonic vehicle swarms;
[0021] The hypersonic vehicle cluster is defined as having one leader and N followers, where the leader is denoted as vehicle 0 and the followers are denoted as vehicles 1, ..., N.
[0022] Then, the communication network between the hypersonic vehicle clusters is modeled as a directed graph. From a set of vertices An edge set and a weighted adjacency matrix Composition; a is defined if and only if the i-th follower hypersonic vehicle is able to receive information from follower hypersonic vehicle j. ij >0; otherwise, a ij =0; for a ii =0; node v i The neighbor set is represented as And the diagram The edge matrix between the leader hypersonic vehicle and the follower hypersonic vehicles This indicates that when the i-th follower hypersonic vehicle has a direct communication link with the leader hypersonic vehicle 0, b i =1, otherwise b i =0.
[0023] The normalized leader hypersonic vehicle cooperative control model, considering the distortion, is expressed with respect to time t as follows:
[0024]
[0025] Where X0(t) = [r0, θ0, φ0, V0, γ0, ψ0] represents the system state of the leader hypersonic vehicle, u0(t) = [α0, σ0] represents the control input to the leader hypersonic vehicle, α0 represents the angle of attack of the leader hypersonic vehicle, σ0 represents the roll angle of the leader hypersonic vehicle, and d0(t) represents the disturbance experienced by the leader hypersonic vehicle. Let f be the derivative of the system state of the hypersonic vehicle leader. The input-output mapping of the dynamic equation in equation (3) is denoted as f.
[0026] Meanwhile, a normalized cooperative control model for the follower hypersonic vehicles in the hypersonic vehicle swarm is defined, with the following expression:
[0027]
[0028] Among them, X i (t)=[r i ,θ i ,φ i V i ,γ i ,ψ i ] represents the system state of the i-th follower hypersonic vehicle, u i (t)=[α i ,σ i ] represents the control input for the i-th follower hypersonic vehicle, α i Let σ represent the angle of attack of the i-th follower hypersonic vehicle. iLet d represent the tilt angle of the i-th follower hypersonic vehicle. i (t) represents the disturbance experienced by the i-th follower aircraft. Let represent the derivative of the system state of the i-th follower hypersonic vehicle.
[0029] Define the normalized distributed cooperative control error e of the i-th follower hypersonic vehicle. i (t), the expression is as follows:
[0030]
[0031] Then, the derivative of the normalized distributed cooperative control error of the i-th follower hypersonic vehicle is obtained. The expression is as follows:
[0032]
[0033] Among them, e ij e i0 Let represent the errors of the i-th follower hypersonic vehicle with respect to its neighbors and leader, respectively, and represent the cooperative tracking control law of the i-th follower hypersonic vehicle. and Unified as u ei , This represents the gradient of the error of the i-th follower hypersonic vehicle relative to its neighbors and leader, respectively. Let represent the gradient of the i-th follower hypersonic vehicle's control input to its neighbors and leader, respectively.
[0034] Finally, if the conditions are met It is then assumed that the hypersonic vehicle cluster achieves cooperative tracking and control.
[0035] Furthermore, step S2 is specifically as follows:
[0036] The objective of the hypersonic vehicle swarm cooperative control is to control the follower hypersonic vehicles to track the desired trajectory of the leader vehicle from the initial state at the lowest cost; that is, the objective is to determine the tracking control law u. ei This makes the cost L(e) i Minimize (t) as follows:
[0037]
[0038] Where τ represents the integral variable with respect to time, J(e i (τ),u ei (e i (τ))) represents the utility function consisting of the state-dependent term Q and the feedback control-dependent term R. Q is defined as Q(e i) = e i T Q0e i Q0 is a positive definite diagonal matrix representing the weight of each state element, and T represents the transpose operation.
[0039] Set a saturation constraint for the optimal tracking control law, and define R such that the feedback remains within the limit, as shown in the following expression:
[0040]
[0041] Where v represents the control input u ei The integral variable is R0, which is a positive definite matrix, and λ1 and λ2 represent α. i and σ i Amplitude limitations.
[0042] The Hamiltonian function is then obtained through the HJB equation, and its expression is as follows:
[0043]
[0044] in, Let represent the augmented form of the gradient of the i-th hypersonic follower vehicle with respect to the error. Let represent the augmented form of the gradient of the i-th follower hypersonic vehicle with respect to the control input. L(e) i The partial derivative of with respect to ei. Then the optimal cost function L(e) is... i ) * The expression is as follows:
[0045]
[0046] Where Ψ(Ω) represents u ei The permissible control region. According to the Bellman optimality principle, V(e i ) * The solution to the HJB equation is expressed as follows:
[0047]
[0048] in, This represents the optimal tracking control law, which is obtained by solving... The partial differential equation is obtained as follows:
[0049]
[0050] in, L(e) i ) * For e i The partial derivatives of .
[0051] Furthermore, step S3 is specifically as follows:
[0052] First, a single neural network is used to analyze L. i (e i The approximation is performed using the following method: the number of neurons in the hidden layer is denoted by l, the weight matrix between the input layer and the hidden layer is denoted by Y, and the weight matrix between the hidden layer and the output layer is denoted by W. The single-evaluation network is selected as a three-layer feedforward neural network, and the output expression of the three-layer neural network is as follows:
[0053]
[0054] Where Ξ represents the activation function and b represents the activation function threshold.
[0055] Then, based on the higher-order Weierstrass approximation theorem, the optimal cost function L obtained by fitting a single-evaluation neural network is... ci (e i ) * Its gradient is represented by an infinite-dimensional linearly independent set of basis functions, and L is constructed using an evaluation neural network with finite-dimensional activation functions. ci (e i ) * The expression is as follows:
[0056]
[0057] in, This represents the weight vector of the neural network. And L ci (e i ) * Regarding e i The expressions for the first-order partial derivative and the second-order partial derivative are as follows:
[0058]
[0059] Then L ci (e i ) * , The expressions for the construction error are as follows:
[0060]
[0061] Based on the derivation of equations (13)-(17), then u ei The approximate expression is as follows:
[0062]
[0063] And L i (e i ) *For error e i If it is continuous, then a linearized neural network can be obtained as the evaluation network for adaptive dynamic programming, that is, the observation form of equations (15) and (16) is obtained, and the expression is as follows:
[0064]
[0065] in, Let represent the optimal weights of the single-evaluation neural network for the i-th follower hypersonic vehicle. Right now Approximation, ∈ c This represents the approximate error that converges to zero after appropriate training under continuous excitation. Therefore, the practical expression for the approximate cost function is as follows:
[0066]
[0067] Next, the gradient descent method is used to update the weights of the single-evaluation network. First, the error function of the single-evaluation network is defined as follows:
[0068]
[0069] The goal of the weight update is then to minimize the loss function, expressed as follows:
[0070]
[0071] Finally, using the gradient descent principle, the online update rule expression for the network weights is designed as follows:
[0072]
[0073] Among them, K W This represents the positive learning rate.
[0074] The beneficial effects of this invention are as follows: First, a 6-DOF normalized model of a hypersonic vehicle is established, and the cooperative tracking control problem of a hypersonic vehicle swarm is defined. An adaptive dynamic programming algorithm is used to obtain the optimal tracking control form under constrained control input. Based on this, a single-evaluation neural network online controller is constructed to solve the optimal tracking control law online. The obtained optimal tracking control law is then applied to the hypersonic vehicle swarm to achieve cooperative tracking control. This invention's method, targeting a hypersonic vehicle system described by a 6-DOF model, considers achieving cooperative tracking control of a hypersonic vehicle swarm under flight constraints and employs an adaptive dynamic programming algorithm for online control, making it applicable to the field of cooperative control of hypersonic vehicle swarms. Attached Figure Description
[0075] Figure 1This is a flowchart of a hypersonic vehicle cooperative control method based on adaptive dynamic programming according to the present invention.
[0076] Figure 2 This is a schematic diagram of the communication topology of a multi-robot system hypersonic aircraft in an embodiment of the present invention.
[0077] Figure 3 This is a diagram illustrating the cooperative tracking behavior exhibited after the method of the present invention is deployed on the hypersonic vehicle in an embodiment of the present invention. Detailed Implementation
[0078] The method of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0079] like Figure 1 The flowchart of a hypersonic vehicle cooperative control method based on adaptive dynamic programming according to the present invention is shown below. The specific steps are as follows:
[0080] S1. Establish a 6-DOF normalized model for hypersonic vehicles and clarify the cooperative tracking and control problem of hypersonic vehicle swarms.
[0081] S2. Based on step S1, use the adaptive dynamic programming algorithm to obtain the optimal tracking control form under the constraint of control input;
[0082] S3. Based on step S2, construct a single-evaluation neural network online controller to solve the optimal tracking control law online;
[0083] S4. The optimal tracking control law obtained in step S3 is applied to the hypersonic vehicle cluster to achieve cooperative tracking control.
[0084] In this embodiment, step S1 is specifically as follows:
[0085] S11. Establish a 6-DOF normalized model for hypersonic vehicles;
[0086] The equation of motion for a 6-DOF hypersonic vehicle without power takeoff and landing is as follows:
[0087]
[0088] Where r represents the radial distance from the Earth's center to the hypersonic vehicle, t represents time; θ and φ represent longitude and latitude, respectively; V represents the Earth's relative velocity; ψ represents the velocity heading angle relative to local north; γ represents the track angle; m and σ represent the hypersonic vehicle's mass and roll angle, respectively; D and L represent aerodynamic drag and lift, which are related to aerodynamic coefficients, angle of attack, and velocity. The Earth's angular rate is denoted as ω. e The local gravitational acceleration is denoted as g.
[0089] To make the mathematical model of cooperative control of hypersonic vehicles applicable to adaptive dynamic programming, the dynamic equations of the 6-DOF particle are normalized. The expressions for the normalized variables are defined as follows:
[0090]
[0091] Among them, R e Let g represent the Earth's radius, and g0 represent gravitational acceleration.
[0092] Substituting the normalized variables into the dynamic equations, we obtain the normalized dynamic equations (6-DOF normalized model) for the hypersonic vehicle as follows:
[0093]
[0094] S12. Clarify the cooperative tracking and control issues of hypersonic vehicle swarms;
[0095] The hypersonic vehicle cluster is defined as having one leader and N followers, where the leader is denoted as vehicle 0, and the followers are denoted as vehicles 1, ..., N. The communication topology diagram of the multi-robot system hypersonic vehicles in this embodiment is shown below. Figure 2 As shown.
[0096] Then, the communication network between the hypersonic vehicle clusters is modeled as a directed graph. From a set of vertices An edge set and a weighted adjacency matrix Composition; a is defined if and only if the i-th follower hypersonic vehicle is able to receive information from follower hypersonic vehicle j. ij >0; otherwise, a ij =0; for a ii =0; node v i The neighbor set is represented as And the diagram The edge matrix between the leader hypersonic vehicle and the follower hypersonic vehicles This indicates that when the i-th follower hypersonic vehicle has a direct communication link with the leader hypersonic vehicle 0, b i =1, otherwise b i =0.
[0097] The normalized leader hypersonic vehicle cooperative control model, considering the distortion, is expressed with respect to time t as follows:
[0098]
[0099] Where X0(t) = [r0, θ0, φ0, V0, γ0, ψ0] represents the system state of the leader hypersonic vehicle, u0(t) = [α0, σ0] represents the control input to the leader hypersonic vehicle, α0 represents the angle of attack of the leader hypersonic vehicle, σ0 represents the roll angle of the leader hypersonic vehicle, and d0(t) represents the disturbance experienced by the leader hypersonic vehicle. Let f be the derivative of the system state of the hypersonic vehicle leader. For ease of representation, the input-output mapping of the dynamic equation in equation (3) is denoted as f.
[0100] Meanwhile, a normalized cooperative control model for the follower hypersonic vehicles in the hypersonic vehicle swarm is defined, with the following expression:
[0101]
[0102] Among them, X i (t)=[r i ,θ i ,φ i V i ,γ i ,ψ i ] represents the system state of the i-th follower hypersonic vehicle, u i (t)=[α i ,σ i ] represents the control input for the i-th follower hypersonic vehicle, α i Let σ represent the angle of attack of the i-th follower hypersonic vehicle. i Let d represent the tilt angle of the i-th follower hypersonic vehicle. i (t) represents the disturbance experienced by the i-th follower aircraft. Let represent the derivative of the system state of the i-th follower hypersonic vehicle.
[0103] Define the normalized distributed cooperative control error e of the i-th follower hypersonic vehicle. i (t), the expression is as follows:
[0104]
[0105] Then, the derivative of the normalized distributed cooperative control error of the i-th follower hypersonic vehicle is obtained. The expression is as follows:
[0106]
[0107] Among them, e ij e i0Let represent the errors of the i-th follower hypersonic vehicle with respect to its neighbors and leader, respectively, and represent the cooperative tracking control law of the i-th follower hypersonic vehicle. and Unified as u ei , This represents the gradient of the error of the i-th follower hypersonic vehicle relative to its neighbors and leader, respectively. Let represent the gradient of the i-th follower hypersonic vehicle's control input to its neighbors and leader, respectively.
[0108] Finally, if the conditions are met It is then assumed that the hypersonic vehicle cluster achieves cooperative tracking and control.
[0109] In this embodiment, step S2 is specifically as follows:
[0110] The objective of the hypersonic vehicle swarm cooperative control is to control the follower hypersonic vehicles to track the desired trajectory of the leader vehicle from the initial state at the lowest cost; that is, the objective is to determine the tracking control law u. ei This makes the cost L(e) i Minimize (t) as follows:
[0111]
[0112] Where τ represents the integral variable with respect to time, J(e i (τ),u ei (e i (τ))) represents the utility function consisting of the state-dependent term Q and the feedback control-dependent term R. Q is defined as Q(e i ) = e i T Q0e i Q0 is a positive definite diagonal matrix representing the weight of each state element, and T represents the transpose operation.
[0113] Set a saturation constraint for the optimal tracking control law, and define R such that the feedback remains within the limit, as shown in the following expression:
[0114]
[0115] Where v represents the control input u ei The integral variable is R0, which is a positive definite matrix, and λ1 and λ2 represent α. i and σ i Amplitude limitations.
[0116] The Hamiltonian function is then obtained through the HJB equation, and its expression is as follows:
[0117]
[0118] in, Let represent the augmented form of the gradient of the i-th hypersonic follower vehicle with respect to the error. Let represent the augmented form of the gradient of the i-th follower hypersonic vehicle with respect to the control input. L(e) i The partial derivative of with respect to ei. Then the optimal cost function L(e) is... i ) * The expression is as follows:
[0119]
[0120] Where Ψ(Ω) represents u ei The permissible control region. According to the Bellman optimality principle, V(e i ) * The solution to the HJB equation is expressed as follows:
[0121]
[0122] in, This represents the optimal tracking control law, which is obtained by solving... The partial differential equation is obtained as follows:
[0123]
[0124] in, L(e) i ) * For e i The partial derivatives of .
[0125] In this embodiment, step S3 is specifically as follows:
[0126] First, a single neural network is used to analyze L. i (e i The approximation is performed using the following method: the number of neurons in the hidden layer is denoted by l, the weight matrix between the input layer and the hidden layer is denoted by Y, and the weight matrix between the hidden layer and the output layer is denoted by W. The single-evaluation network is selected as a three-layer feedforward neural network, and the output expression of the three-layer neural network is as follows:
[0127]
[0128] Where Ξ represents the activation function and b represents the activation function threshold.
[0129] Then, based on the higher-order Weierstrass approximation theorem, the optimal cost function L obtained by fitting a single-evaluation neural network is... ci (e i ) *Its gradient is represented by an infinite-dimensional linearly independent set of basis functions, and L is constructed using an evaluation neural network with finite-dimensional activation functions. ci (e i ) * The expression is as follows:
[0130]
[0131] in, This represents the weight vector of the neural network. And L ci (e i ) * Regarding e i The expressions for the first-order partial derivative and the second-order partial derivative are as follows:
[0132]
[0133] Then L ci (e i ) * , The expressions for the construction error are as follows:
[0134]
[0135] Based on the derivation of equations (13)-(17), then u ei The approximate expression is as follows:
[0136]
[0137] Because of L i (e i ) * For error e i If it is continuous, then a linearized neural network can be obtained as the evaluation network for adaptive dynamic programming, that is, the observation form of equations (15) and (16) is obtained, and the expression is as follows:
[0138]
[0139] in, Let represent the optimal weights of the single-evaluation neural network for the i-th follower hypersonic vehicle. Right now Approximation, ∈ c This represents the approximate error that converges to zero after appropriate training under continuous excitation. Therefore, the practical expression for the approximate cost function is as follows:
[0140]
[0141] Next, the gradient descent method is used to update the weights of the single-evaluation network. First, the error function of the single-evaluation network is defined as follows:
[0142]
[0143] The goal of the weight update is then to minimize the loss function, expressed as follows:
[0144]
[0145] Finally, using the gradient descent principle, the online update rule expression for the network weights is designed as follows:
[0146]
[0147] Among them, K W This represents the positive learning rate.
[0148] In this embodiment, step S4 inputs the optimal cooperative tracking control law calculated in step S3 into the hypersonic vehicle cluster, which can achieve cooperative tracking control of hypersonic vehicles under the condition of considering flight constraints. The final effect is as follows: Figure 3 As shown.
[0149] In summary, the method of this invention is designed for hypersonic vehicle systems described by a 6-DOF model. It considers the cooperative tracking control of hypersonic vehicle clusters under flight constraints and employs an adaptive dynamic programming algorithm to achieve online control. It is applicable to the field of cooperative control of hypersonic vehicle clusters.
[0150] Those skilled in the art will recognize that the embodiments described herein are intended to help the reader understand the principles of the invention, and should be understood that the scope of protection of the invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific modifications and combinations based on the technical teachings disclosed in this invention without departing from the spirit of the invention, and these modifications and combinations are still within the scope of protection of this invention.
Claims
1. A cooperative control method for hypersonic vehicles based on adaptive dynamic programming, the specific steps of which are as follows: S1. Establish a 6-DOF normalized model for hypersonic vehicles and clarify the cooperative tracking and control problem of hypersonic vehicle swarms. S2. Based on step S1, use the adaptive dynamic programming algorithm to obtain the optimal tracking control form under the constraint of control input; S3. Based on step S2, construct a single-evaluation neural network online controller to solve the optimal tracking control law online; S4. The optimal tracking control law obtained in step S3 is applied to the hypersonic vehicle cluster to achieve cooperative tracking control. Step S2 is as follows: The objective of the hypersonic vehicle swarm cooperative control is to control the follower hypersonic vehicles to track the desired trajectory of the leader vehicle from the initial state at the lowest cost; that is, the objective is to determine the tracking control law u. ei This makes the cost L(e) i Minimize (t) as follows: in, t represents time, e i (t) represents the normalized distributed cooperative control error of the i-th follower hypersonic vehicle, τ represents the integral variable over time, and J(e i (τ),u ei (e i (τ) represents the utility function consisting of the state-dependent term Q and the feedback control-dependent term R; Q is defined as... Furthermore, Q0 is a positive definite diagonal matrix representing the weight of each state element, and T represents the transpose operation; Set a saturation constraint for the optimal tracking control law, and define R such that the feedback remains within the limit, as shown in the following expression: Where v represents the control input u ei The integral variable is R0, which is a positive definite matrix, and λ1 and λ2 represent α. i and σ i Amplitude limitation; The Hamiltonian function is then obtained through the HJB equation, and its expression is as follows: in, Let represent the augmented form of the gradient of the i-th hypersonic follower vehicle with respect to the error. Let represent the augmented form of the gradient of the i-th follower hypersonic vehicle with respect to the control input. L(e) i ) for e i The partial derivatives; then the optimal cost function L(e i ) * The expression is as follows: Where Ψ(Ω) represents u ei The allowable control region; according to the Bellman optimality principle V(e i ) * The solution to the HJB equation is expressed as follows: in, This represents the optimal tracking control law, which is obtained by solving... The partial differential equation is obtained as follows: in, L(e) i ) * For e i The partial derivatives of .
2. The hypersonic vehicle cooperative control method based on adaptive dynamic programming according to claim 1, characterized in that, The specific steps of S1 are as follows: S11. Establish a 6-DOF normalized model for hypersonic vehicles; The equation of motion for a 6-DOF hypersonic vehicle without power takeoff and landing is as follows: Where r represents the radial distance from the Earth's center to the hypersonic vehicle, t represents time; θ and φ represent longitude and latitude, respectively; V represents the Earth's relative velocity; ψ represents the velocity heading angle relative to local north, γ represents the track angle; m and σ represent the hypersonic vehicle's mass and roll angle, respectively; D and L represent aerodynamic drag and lift, respectively; and the Earth's angular rate is denoted as ω. e The local gravitational acceleration is denoted as g; The dynamic equations of a 6-DOF particle are normalized, and the normalized variable expressions are defined as follows: Among them, R e g represents the Earth's radius, and g0 represents the acceleration due to gravity. Substituting the normalized variables into the dynamic equations, we obtain the normalized dynamic equations for the hypersonic vehicle, as shown below: S12. Clarify the cooperative tracking and control issues of hypersonic vehicle swarms; The hypersonic vehicle cluster is defined as having one leader and N followers, where the leader is denoted as vehicle 0 and the followers are denoted as vehicle 1, ..., N. Then, the communication network between the hypersonic vehicle clusters is modeled as a directed graph. From a set of vertices An edge set and a weighted adjacency matrix Composition; a is defined if and only if the i-th follower hypersonic vehicle is able to receive information from follower hypersonic vehicle j. ij >0; otherwise, a ij =0; for node v i The neighbor set is represented as And the diagram The edge matrix between the leader hypersonic vehicle and the follower hypersonic vehicles This indicates that when the i-th follower hypersonic vehicle has a direct communication link with the leader hypersonic vehicle 0, b i =1, otherwise b i =0; The normalized leader hypersonic vehicle cooperative control model, considering the distortion, is expressed with respect to time t as follows: Where X0(t) = [r0, θ0, φ0, V0, γ0, ψ0] represents the system state of the leader hypersonic vehicle, u0(t) = [α0, σ0] represents the control input to the leader hypersonic vehicle, α0 represents the angle of attack of the leader hypersonic vehicle, σ0 represents the roll angle of the leader hypersonic vehicle, and d0(t) represents the disturbance experienced by the leader hypersonic vehicle. The derivative of the system state of the leader hypersonic vehicle is represented by f; and the input-output mapping of the dynamic equation in equation (9) is denoted as f; Meanwhile, a normalized cooperative control model for the follower hypersonic vehicles in the hypersonic vehicle swarm is defined, with the following expression: Among them, X i (t)=[r i ,θ i ,φ i V i ,γ i ,ψ i ] represents the system state of the i-th follower hypersonic vehicle, u i (t)=[α i ,σ i ] represents the control input for the i-th follower hypersonic vehicle, α i Let σ represent the angle of attack of the i-th follower hypersonic vehicle. i Let d represent the tilt angle of the i-th follower hypersonic vehicle. i (t) represents the disturbance experienced by the i-th follower aircraft. The derivative of the system state of the i-th follower hypersonic vehicle; Define the normalized distributed cooperative control error e of the i-th follower hypersonic vehicle. i (t), the expression is as follows: Then, the derivative of the normalized distributed cooperative control error of the i-th follower hypersonic vehicle is obtained. The expression is as follows: Among them, e ij e i0 Let represent the errors of the i-th follower hypersonic vehicle with respect to its neighbors and leader, respectively, and represent the cooperative tracking control law of the i-th follower hypersonic vehicle. and Unified as u ei , This represents the gradient of the error of the i-th follower hypersonic vehicle relative to its neighbors and leader, respectively. Let represent the gradient of the i-th follower hypersonic vehicle's control input to its neighbors and leader, respectively; Finally, if the conditions are met It is then assumed that the hypersonic vehicle cluster achieves cooperative tracking and control.
3. The hypersonic vehicle cooperative control method based on adaptive dynamic programming according to claim 1, characterized in that, Step S3 is as follows: First, a single neural network is used to analyze L. i (e i The approximation is performed using l as the number of hidden layer neurons, Y as the weight matrix between the input layer and the hidden layer, and W as the weight matrix between the hidden layer and the output layer. The single evaluation network is selected as a three-layer feedforward neural network, and the output expression of the three-layer neural network is as follows: Where Ξ represents the activation function, and b represents the activation function threshold; Then, based on the higher-order Weierstrass approximation theorem, the optimal cost function L obtained by fitting a single-evaluation neural network is... ci (e i ) * Its gradient is represented by an infinite-dimensional linearly independent set of basis functions, and L is constructed using an evaluation neural network with finite-dimensional activation functions. ci (e i ) * The expression is as follows: in, L represents the weight vector of the neural network; and L ci (e i ) * Regarding e i The expressions for the first-order partial derivative and the second-order partial derivative are as follows: Then L ci (e i ) * , The expressions for the construction error are as follows: Based on the derivation of equations (6), (14)-(17), then u ei The approximate expression is as follows: And L i (e i ) * For error e i If it is continuous, then a linearized neural network can be obtained as the evaluation network for adaptive dynamic programming, that is, the observation form of equations (15) and (16) is obtained, and the expression is as follows: in, Let represent the optimal weights of a single evaluation neural network for the i-th follower hypersonic vehicle. Right now The approximation of ρ c Let represent the approximate error that converges to zero after appropriate training under continuous excitation; then the practical expression for the approximate cost function is as follows: Next, the gradient descent method is used to update the weights of the single-evaluation network. First, the error function of the single-evaluation network is defined as follows: The goal of the weight update is then to minimize the loss function, expressed as follows: Finally, using the gradient descent principle, the online update rule expression for the network weights is designed as follows: Among them, K W This represents the positive learning rate.
Citation Information
Patent Citations
Hypersonic aircraft reentry guidance method based on deep neural network
CN115951585A
Adaptive variable-structure game guidance method for variable-sweepback hypersonic gliding aircraft
CN118795783A