A reinforcement learning driven hypersonic variable configuration vehicle reentry trajectory optimization method

By using the DDPG reinforcement learning algorithm and the Kriging surrogate model, the problems of poor reentry performance and long computation time of hypersonic variant aircraft were solved, achieving efficient trajectory optimization and variant decision-making, thereby improving flight performance and battlefield survivability.

CN119885812BActive Publication Date: 2025-12-12BEIJING INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411657156.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-19
Publication Date
2025-12-12
Estimated Expiration
2044-11-19

Smart Images

  • Figure CN119885812B_ABST
    Figure CN119885812B_ABST
Patent Text Reader

Abstract

The application relates to a method for optimizing reentry trajectories of a hypersonic variable aircraft driven by reinforcement learning, and belongs to the fields of hypersonic aircrafts and trajectory optimization. The application is realized by the following method: a parameterized model of a hypersonic variable aircraft including a waverider fuselage and two variable wings is established based on a category shape transformation method (CST); a Kriging surrogate model is established by considering the influences of factors such as flight height H, flight Mach number Ma, attack angle alpha, inner wing segment sweepback angle chi1 and outer wing segment sweepback angle chi2, so as to obtain the aerodynamic performance of the hypersonic variable aircraft; and a reentry trajectory optimization model driven by a DDPG (Deep Deterministic Policy Gradient) algorithm is trained by combining expert experience and the constraints of dynamics, initial and final states, heat flow, overload and dynamic pressure during the flight process, according to the flight characteristics of the hypersonic variable aircraft during the reentry flight process. Through flight simulation prediction of different hypersonic variable aircrafts, variable decision-making is realized in real time, and the reentry trajectory optimization of the hypersonic variable aircraft is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to a method for optimizing reentry trajectory of a hypersonic variable aircraft driven by reinforcement learning, and belongs to the field of hypersonic aircraft and trajectory optimization. BACKGROUND

[0002] Due to the inherent high flight speed and high flight altitude of a hypersonic aircraft, the hypersonic aircraft has strong penetration ability, rapid response ability and excellent concealment ability, and can perform specific combat tasks such as competing for high-altitude space, attacking long-range targets and monitoring wide-area targets. Due to the diversity of tasks that can be performed, the hypersonic aircraft is a major development direction of the world in the field of aerospace in the 21st century. Due to the characteristics of the hypersonic aircraft in flight tasks, such as "crossing air domain" and "wide speed domain", although the design method of the design of a layout scheme, the overall scheme and the aerodynamic shape is very mature, it is not suitable. As a new hot spot in the field of aerospace research, the variable technology can adaptively adjust the shape according to the complex environment and sudden conditions that cannot be predicted in advance to improve the aerodynamic characteristics of the aircraft, optimize the flight performance and improve the overall performance of the whole flight envelope.

[0003] The reentry phase is an extremely important flight process of the hypersonic aircraft in the whole flight process. From the point of separation of the arrow aircraft, the aircraft is in a non-powered flight until the point of diving and pressing into the terminal guidance phase. The reentry phase accounts for the largest proportion in the whole trajectory, has the widest domain and the most complex constraints, and is the largest flight phase affecting the overall trajectory performance. However, the state quantity of the hypersonic variable aircraft with the variable technology is complex, and the influence mechanism is difficult to analyze, so that the traditional optimization method is difficult to give a suitable variable decision and difficult to realize good reentry trajectory optimization. In addition, the complex environment with sharp changes and the performance requirements limit the implementation of the variable scheme, and it is not suitable for large and rapid deformation. Therefore, designing a decision maker for determining the deformation of the aircraft according to the flight altitude, speed, attack angle and current aerodynamic shape and other state quantities can deeply tap the design potential of the hypersonic aircraft, improve the flight performance of the aircraft in the reentry trajectory, and enhance the battlefield survival ability and task completion ability of the aircraft, which is of great significance.

[0004] With the development of artificial intelligence and neural networks, reinforcement learning can be iteratively trained by the reward obtained by the agent performing actions, and the decision-making that adapts to complex environments is derived. In decision-making problems, reinforcement learning can well solve decision-making problems such as autonomous mobile robot obstacle avoidance and mechanical arm autonomous action planning. The trajectory optimization problem similar to these problems is also a state transition problem, which decides the action through the state of the body to reach the next state, and the change of the state is used to train the decision-making model. Therefore, the deep neural network decision algorithm (Deep Deterministic Policy Gradient, DDPG) based on the combination of reinforcement learning and neural network is used to give the variable decision, and the trajectory optimization of the hypersonic variable aircraft in the reentry section has a good development prospect. SUMMARY

[0005] In order to solve the problems of poor hypersonic variable aircraft reentry flight performance and long calculation time of traditional optimization algorithm, the technical problem to be solved by the hypersonic variable aircraft reentry trajectory optimization method driven by reinforcement learning disclosed in the application is to realize parameterized modeling of the hypersonic variable aircraft, realize rapid construction of different variable configurations, construct a Kriging surrogate model according to the aerodynamic simulation sampling points of the geometric model, realize aerodynamic performance prediction under different flight configurations and flight environments, train the optimization model based on the DDPG reinforcement learning algorithm, construct a decision maker with maximum lift-drag ratio as the main target, and finally obtain a reentry optimization trajectory that maintains high lift-drag ratio under most states and meets the constraints

[0006] The purpose of the application is realized by the following technical scheme:

[0007] The hypersonic variable aircraft reentry trajectory optimization method driven by reinforcement learning disclosed in the application is based on the Class / Shape Transformation (CST) to establish a parameterized model of a hypersonic variable aircraft including a waverider fuselage and two variable wings. Considering the influence of factors such as flight altitude H, flight Mach number Ma, attack angle a, inner wing segment sweep angle x1 and outer wing segment sweep angle x2, a Kriging surrogate model is established to obtain the aerodynamic performance of the hypersonic variable aircraft. Based on the DDPG reinforcement learning algorithm, according to the flight characteristics of the hypersonic variable aircraft during the reentry flight process, combined with expert experience and the constraints of dynamics, initial and final state, heat flow, overload and dynamic pressure during the flight process, the DDPG algorithm driven reentry trajectory optimization model training is realized. Through flight simulation prediction of different hypersonic variable aircrafts in actual application, variable decision is given in real time, and the reentry trajectory optimization of the hypersonic variable aircraft is realized.

[0008] The application discloses a reinforcement learning driven trajectory optimization method for a hypersonic variable aircraft reentry phase, and comprises the following steps.

[0009] Step A: establishing a parameterized model of the hypersonic variable aircraft.

[0010] The implementation method of step A is as follows.

[0011] Step A-1: constructing a reference shape and a variable form.According to the reentry flight environment characteristics of the hypersonic variable aircraft, a similar waverider configuration is adopted, that is, a waverider fuselage combined with a hypersonic wing shape.During flight, the atmosphere is rarefied, and a large-scale deformation variable post-sweep angle is adopted to improve the variable performance.

[0012] Step A-2: parameterizing and describing the geometric model.Based on the standard three-dimensional CST description equation, combined with the characteristics of the waverider fuselage shape, combined with the symmetric relationship of the leading edge line located on the horizontal plane and the longitudinal plane, the Z coordinate expression of the upper surface of the initial shape of the fuselage is obtained, that is, the parameterized model of the hypersonic variable aircraft.

[0013]

[0014] In the formula, H' u is the spanwise section upper surface control parameter, is a category function, M' u is the body axis section upper surface control parameter, N u is the upper surface thickness control parameter, η is the X, Y normalized coordinate, and the subscript l is the lower surface related control parameter.

[0015] Step B: establishing a six-degree-of-freedom trajectory model of the hypersonic variable aircraft.

[0016] The implementation method of step B is as follows.

[0017] Step B-1: in the reentry phase flight environment of the hypersonic variable aircraft, the sound speed and the atmospheric density change need to be considered, so the aerodynamic model of the hypersonic variable aircraft is jointly influenced by the flight height H, the flight Mach number Ma, the attack angle alpha, the inner wing section sweep angle chi1 and the outer wing section sweep angle chi2; according to the flight environment and the aircraft structure characteristics, a certain number of sampling points are generated in the variable range by using a Latin hypercube design method:

[0018] [H, Ma, alpha, chi1, chi2] = lhsdesign (N) (2)

[0019] In the formula, N is the number of sampling points.

[0020] A series of automatic processes of drawing a grid by the obtained sampling point information, importing aerodynamic simulation software to solve aerodynamic force, and collecting aerodynamic force data are performed to obtain a plurality of discrete lift-drag force coefficient aerodynamic coefficient data; the lift-drag force coefficient aerodynamic coefficient data are lift coefficient C L and drag coefficient C D .

[0021] In the step B of solving the aerodynamic force data, grid independence is checked, aerodynamic force data of the determined similar working condition aircraft are solved, and the correctness of the collected aerodynamic force data is verified by an experimental data comparison method to obtain the aerodynamic force data with correctness.

[0022] Step B-2: Kriging surrogate models are respectively constructed for the lift coefficient C L and the drag coefficient C D .

[0023]

[0024] In the formula, C and C are the fitting functions of the Kriging surrogate models of the lift coefficient and the drag coefficient respectively.

[0025] Leave one out cross validation (LOOCV) is performed on the established Kringing surrogate model, and a complex correlation coefficient is given.

[0026]

[0027] In the formula, C Li and C Di are the lift coefficients and the drag coefficients obtained by the surrogate model at each sampling point; if any complex correlation coefficient is less than 0.8, the number of sampling points is increased in step B-1; otherwise, step B-3 is performed.

[0028] Step B-3: According to the flight environment characteristics experienced by the hypersonic variable aircraft in the reentry segment flight, by describing the coordinate system and the atmospheric model, the state quantities are selected as the rotating earth relative velocity v, the flight path angle γ, the heading angle ψ, the longitude λ, the latitude φ, and the altitude H, and a six-degree-of-freedom trajectory model of the hypersonic variable aircraft is constructed.

[0029]

[0030] In the formula, R e is the earth radius, g is the gravity acceleration, and a nominal value is adopted; σ is the roll angle, and the value is 0; L is the lift acceleration, D is the drag acceleration, and needs to be calculated; C v2 is the relative velocity v direction component of the dependent acceleration term, C γ1 and Cγ2 γx and γy are the components of the flight path angle γ in the x and y directions, respectively, C ψ1 and C ψ2 γx and γy are the components of the flight path angle γ in the x and y directions, respectively, C

[0031]

[0032] where ω e is the earth rotation angular velocity, and the nominal value is used.

[0033] The lift acceleration L and the drag acceleration D are given by using the aerodynamic data obtained from equation (3).

[0034]

[0035] where q = 0.5 ρv 2 2 is the dynamic pressure, ρ is the atmospheric density given by the standard atmospheric model, and S is the reference area of the wing.

[0036] Step C: determining the pointing objects of the state quantity, the action quantity, the constraints on the body, and the initial and final conditions in the reinforcement learning algorithm; determining the form of the reward function and the experience pool composed of expert experience and exploration experience in the reinforcement learning algorithm; determining the network architecture setting and initial quantity assignment in the reinforcement learning algorithm; training the reentry phase trajectory optimization model driven by the reinforcement learning algorithm to obtain the trained reentry phase trajectory optimization model driven by the reinforcement learning algorithm, and performing trajectory optimization according to the trained reentry phase trajectory optimization model driven by the reinforcement learning algorithm.

[0037] The implementation method of step C is as follows:

[0038] Step C-1: setting the inner wing segment sweepback angle χ1, the outer wing segment sweepback angle χ2, the height H, the flight path angle γ, the Mach number Ma, and the attack angle α as the required basic element state quantity parameters s in the DDPG reinforcement learning algorithm.

[0039] s = [ χ1, χ2, H, γ, Ma, α ] (8)

[0040] Step C-2: setting the inner wing segment sweepback angle change amount Δχ1 and the outer wing segment sweepback angle change amount Δχ2 as the required basic element action quantity a in the DDPG reinforcement learning algorithm.

[0041] a = [ Δχ1, Δχ2 ] (9)

[0042] The state quantity parameter attack angle α changes with the speed in stages as shown in equation (10).

[0043]

[0044] where α max is the maximum attack angle, and αmax K is the maximum lift-drag ratio angle of attack, V1 is the maximum angle of attack initial speed, V2 is the maximum lift-drag ratio corresponding speed, and the remaining parameters are as follows:

[0045]

[0046] The maximum heat flow constraint, the overload constraint and the dynamic pressure constraint related constraint conditions are considered; based on the standard resistance acceleration profile, the reentry boundary constraint model is given.

[0047]

[0048] In the formula, q max is the maximum heat flow constraint, n′ max is the overload constraint, q′ max is the dynamic pressure constraint, D q,max is the resistance acceleration under the maximum heat flow constraint, D n′,max is the resistance acceleration under the overload constraint, and D q′,max is the resistance acceleration under the dynamic pressure constraint.

[0049] Considering the flight characteristics, flight environment characteristics and body itself restrictions of the hypersonic morphing aircraft during the reentry flight process, the body lift-off parameters lift-off speed v0 and lift-off height H0 are given; the initial configuration parameters are the inner wing segment sweepback angle χ 10 , the outer wing segment sweepback angle χ 20 ; the initial attitude parameters of the body are the track inclination angle γ0 and the heading angle ψ0; the initial position parameters of the body are the longitude λ0 and the latitude φ0; the remaining parameters are the mass m of the aircraft, the wing reference area S, the earth radius R e , the rotation speed ω e and the earth surface gravity acceleration g.

[0050] The reentry end condition is adopted to terminate the training, and the termination speed v s and the termination height H s are set.

[0051] Step C-2: Establish the reward function R required for reinforcement learning training, and construct the experience pool combined with expert experience according to the reward function R.

[0052] First, considering the objective function of the maximum lift-drag ratio, the smoothing requirement of the deformation amount and the related constraint requirement, the reward function R is established.

[0053] R=λ1C L / D -λ2(Δχ1+Δχ2)-λ3ΔD q,n′,q′ (13)

[0054] Wherein, λ1, λ2, λ3 are preset weight coefficients, C L / D is the lift-drag ratio, and ΔD​q,n′,q′ The heat flow, overload, and dynamic pressure constraints exceed the amount.

[0055] During the training process, the training experience pool is divided into exploration experience E explore and expert experience E expert , and the parameter a e is introduced as the ratio of the total sampling of expert experience sampling

[0056]

[0057] where M is the number of samples in the corresponding experience pool.

[0058] The exploration experience E explore in the experience pool is the experience formed by the interaction of the agent with the environment after the neural network outputs the action in the training round, and each round of training is stored until the upper limit of the set experience pool capacity is reached. The expert experience E expert in the experience pool is the experience of the inner and outer wing segments of the current state quantity under the condition of achieving the maximum lift-drag ratio.

[0059] Step C-3: Give the reinforcement learning network architecture setting.

[0060] An action network Actor is established, which is composed of two initial hyperparameters and networks Actor_e and Actor_t with the same structure. The former gives the output action quantity a when the input state quantity s is input, and the latter outputs the next action quantity a' when the next state quantity s' is input. The update of Actor_e uses normal update, and the update of Actor_t is in the form of soft update. The network architecture includes five layers of input layer, three layers of hidden layer, and output layer. Except that sigmoid is used as the activation function between the last two layers, relu is used as the activation function between the remaining layers, and the connection mode is full connection between neurons. The number of input neurons is set to the dimension of the state quantity, and the number of output neurons is set to the dimension of the action quantity.

[0061] An evaluation network Critic is established, which is composed of two initial hyperparameters and networks Critic_e and Critic_t with the same structure, which give the output value function when the input state quantity s and action quantity a are input. The update of Critic_e uses normal update, and the update of Critic_t is in the form of soft update. The network architecture includes five layers of input layer, three layers of hidden layer, and output layer. Except that sigmoid is used as the activation function between the last two layers, relu is used as the activation function between the remaining layers, and the connection mode is full connection between neurons. The number of input neurons is set to the total dimension of the action quantity and the state quantity, and the number of output neurons is set to the dimension of the value function.

[0062] The normal update adopts the weight training form of gradient descent method, and the weight training is as follows:

[0063]

[0064] Wherein, L is a damage function, θ is a network hyperparameter, E is an expectation function, R is a reward function, γ e is a discount factor, and Q is a value function.

[0065] The normal update form is as follows:

[0066]

[0067] A soft update coefficient τ is set, the original network hyperparameter θ and the updated hyperparameter θ' are set, and the update form is as follows:

[0068] θ'←τθ+(1-τ)θ',τ≤1 (17)

[0069] After the network architecture is set, the state initial value s0 and the position state initial value X0 are determined according to the initial data given in step C-1:

[0070]

[0071] Wherein, Ma0 and α0 are derived and calculated from v0.

[0072] Step C-4: performing reentry segment trajectory optimizer training under the complete DDPG algorithm architecture;

[0073] The network parameter training steps are as follows:

[0074] At time t, there are action value a t-Δt , state value s t and position state value X t :

[0075]

[0076] Wherein, the subscript t represents the value of the variable at time t.

[0077] Substitute variables H t , α t , Ma t , χ 1t , χ 2t into formula (3) to obtain aerodynamic coefficients C Lt , C Dt at the current time; substitute the aerodynamic coefficients C Lt , C Dt at the current time into formula (7) to obtain lift acceleration L t and drag acceleration D t at the current time; substitute Lt D t With X t Substituting into equation (5), the next position state value X is obtained using a Runge-Kutta fourth-order integrator. t+Δt Δt is the set specific time step; the position status value X t+Δt Chinese v t+Δt Substitute into equation (10) to obtain α t+Δt Substitute into the standard model to obtain Ma t+Δt ; will s t Get a from the network Actor t a t As shown in the following formula:

[0078] a t =[Δχ 1t ,Δχ 2t (20)

[0079] There is a relation:

[0080]

[0081] a t Substituting into equation (21), the next time step χ is obtained using a Runge-Kutta fourth-order integrator. 1t+Δt ,χ 2t+Δt Integrate location status value X t+Δt The parameter H has been obtained. t+Δt ,γ t+Δt Given s t+Δt ; Set action value a t With state value s t The input is fed into the network Critic, which provides the value function and updates the network Actor, completing the training process at time t. At t=0, the following equation holds:

[0082]

[0083] After inputting initial values ​​to begin training, the network parameter training steps are repeated until the training termination condition is met, i.e.:

[0084]

[0085] Once the training conditions are met, this training process is recorded as one training iteration, and the number of training iterations is i = i + 1; an upper limit i for training iterations is set. max , when i=i max Training ends at a certain point, resulting in a well-trained reinforcement learning algorithm-driven reentry trajectory optimization model. Trajectory optimization is then performed based on this model.

[0086] Step D: Considering the hypersonic variable aircraft reentry segment flight under different end conditions, the specific characteristic parameters of the aircraft are determined, the trajectory optimization model driven by the trained reinforcement learning algorithm in step C is used for trajectory optimization, the accuracy and efficiency of the reentry segment trajectory optimization are improved, and high-precision variable decision control of the hypersonic variable aircraft in the reentry segment flight process is realized. After that, compared with the simulation analysis under the same conditions of other methods, the flight benefits of the established optimization model are highlighted.

[0087] The implementation method of step D is as follows:

[0088] Step D-1: The trajectories under three different simulation conditions are optimized. The hypersonic variable aircraft has small differences in trajectory in the initial descent segment due to small differences in aerodynamic force, and the initial point is the vehicle separation point, which is transported to the near space at an altitude of about 60 km by the boost rocket. The simulation termination condition is that the hypersonic variable aircraft reaches the dive-down point determined by the task characteristics, and the aircraft is ready to dive down and enter the terminal guidance segment to reach the target task point. The simulation initial conditions are the same as the training initial conditions, and the trajectories under different end conditions are optimized to give the flight benefits of the hypersonic aircraft in the reentry segment flight process.

[0089] Step D-2: Further explore the benefits obtained by the established DDPG algorithm driven reentry segment trajectory optimization model for the hypersonic variable aircraft in the reentry segment flight process, and compare the flight simulation based on various methods. For the optimization method, SLSQP, linear approximation constraint optimization (COBYLA) and trust region constraint algorithm (trust-constr) are used to calculate the corresponding angle of the maximum lift-drag ratio under the established lift coefficient proxy model and drag coefficient proxy model, and the state quantity at the next time is calculated by substituting the motion equation, and the comparison of the optimization simulation results is given. For the direct control method, the pseudospectral method which performs well in solving trajectory optimization problems is used. The LG points are the roots of N-order Legendre polynomials, and under the same initial, terminal conditions and flight state parameters as the previous simulation, the flight results of the whole segment simulation are compared and given.

[0090] Beneficial effects:

[0091] 1. In view of the research status of hypersonic variable aircraft analysis model shortage, the application discloses a reinforcement learning driven hypersonic variable aircraft reentry trajectory optimization method, a hypersonic variable aircraft parameterized model based on the CST method is established, a wing deformation configuration with variable sweepback as the main deformation form is determined, and aerodynamic data calculation is performed with the aid of engineering software ICEM and high-precision aerodynamic simulation software SU2. On this basis, a reentry six-degree-of-freedom trajectory model is given, the overall modeling of the hypersonic variable aircraft is improved, and the modeling accuracy of the hypersonic variable aircraft is improved.

[0092] 2. The reinforcement learning driven hypersonic variable aircraft reentry trajectory optimization method disclosed by the application replaces the high-precision aerodynamic analysis model of the hypersonic variable aircraft with a proxy model, realizes rapid prediction of the aerodynamic parameters of the aircraft during flight, can reduce the analysis times of the high-time-consuming aerodynamic analysis model required for optimization, and improves the reentry trajectory optimization efficiency.

[0093] 3. The reinforcement learning driven hypersonic variable aircraft reentry trajectory optimization method disclosed by the application faces the reentry trajectory optimization requirements of the hypersonic variable aircraft, combines the applicable scenarios and specific characteristics of various algorithms in reinforcement learning, and adopts the DDPG algorithm which can be applied to continuous states and continuous actions. The established DDPG algorithm driven hypersonic variable aircraft reentry trajectory optimization model is trained in combination with expert experience, realizes millisecond-level variable decision control of the hypersonic variable aircraft during the reentry flight process, and has universality, efficiency and reliability. BRIEF DESCRIPTION OF DRAWINGS

[0094] Figure 1 is a schematic diagram of the overall layout of the hypersonic variable aircraft, wherein, Figure 1 (a) is a side view of the initial shape of the fuselage, Figure 1 (b) is a rear view of the initial shape of the fuselage, Figure 1 (c) is a plan view of the initial shape of the fuselage, Figure 1 (d) is a plan view of the half mold of the whole machine.

[0095] Figure 2 is a schematic diagram of the airfoil transformation of the hypersonic variable aircraft;

[0096] Figure 3 is a flowchart of aerodynamic data collection;

[0097] Figure 4 is a reentry trajectory optimization flowchart constructed by DDPG;

[0098] Figure 5 is a comparison chart of various parameters before and after trajectory optimization under the termination speed v1=3000m / s and the termination height H1=30km;

[0099] Figure 6 The parameter comparison chart before and after trajectory optimization under the termination condition of termination speed v1=3500m / s and termination height H1=35km;

[0100] Figure 7 The parameter comparison chart before and after trajectory optimization under the termination condition of termination speed v1=2500m / s and termination height H1=25km;

[0101] Figure 8 The lift-drag ratio comparison chart under the same termination condition;

[0102] Figure 9 The parameter comparison chart before and after SLSQP optimization;

[0103] Figure 10 The parameter comparison chart before and after COBYLA optimization;

[0104] Figure 11 The parameter comparison chart before and after trust-constr optimization. DETAILED DESCRIPTION

[0105] In order to better illustrate the purposes and advantages of the present application, the present application will be described in detail below with reference to the accompanying drawings and in combination with embodiments.

[0106] As shown in Figure 4 , the present embodiment discloses a reinforcement learning driven hypersonic variable aircraft reentry trajectory optimization method, and the specific implementation steps are as follows:

[0107] Step A: Establish a parameterized model of a hypersonic variable aircraft. During reentry flight, different flight configurations need to be adopted according to different flight environments to obtain the most suitable aerodynamic shape. On the basis of the reference shape, determine the variable form that can be realized, and give different variable configurations. In order to accurately draw the shape of the hypersonic variable aircraft and quickly draw after changing the geometric parameters, the aircraft shape needs to be parameterized.

[0108] The implementation method of step A is as follows:

[0109] Step A-1: Construct the reference shape and variable form. According to the characteristics of the reentry flight environment of the hypersonic variable aircraft, a similar waverider configuration is adopted, that is, a waverider fuselage combined with a hypersonic wing shape as shown in Figure 1 . Due to the thin atmosphere during flight, a large-scale deformation of the rear sweep angle is adopted to improve the performance of the variable body. For the rear sweep angle deformation, a shear type rigid-flexible coupling skin and a sliding rail type skeleton structure are adopted to achieve the effect as shown in Figure 2 .

[0110] Step A-2: Parametric description of the geometric model. In order to accurately draw the hypersonic variable aircraft shape and quickly draw after changing the geometric parameters, it is necessary to parametrically describe the aircraft shape. Based on the standard three-dimensional CST description equation combined with the characteristics of the wave-like body fuselage shape, combined with the symmetry relationship of the leading edge line in the horizontal plane and the longitudinal plane, the Z coordinate expression of the upper and lower surfaces of the fuselage initial shape is obtained.

[0111]

[0112] In the formula, the spanwise section upper surface control parameter H' u = 0.446, is a category function, the body axis section upper surface control parameter M' u = 0.752, the upper surface thickness control parameter N' u = 0.796, the body axis section lower surface control parameter M' l = 0.973, the upper surface thickness control parameter N' l = 1.948, η is the X, Y normalized coordinates.

[0113] Step B: Establishing a six-degree-of-freedom trajectory model of a hypersonic variable aircraft.

[0114] Step B implementation method is as follows:

[0115] Step B-1: In the reentry flight environment of a hypersonic variable aircraft, the speed of sound and the change of atmospheric density need to be considered, so the aerodynamic model of the hypersonic variable aircraft is influenced by the flight height h ∈ [30, 60] km, the flight Mach number Ma ∈ [8, 25], the attack angle α ∈ [10°, 25°], the inner wing segment sweep angle χ1 ∈ [18°, 30°], and the outer wing segment sweep angle χ2 ∈ [0°, χ1]. According to the flight environment and the structural characteristics of the aircraft, a certain number of sampling points are generated within the variable range by using the Latin hypercube experimental design method:

[0116] [H, Ma, α, χ1, χ2] = lhsdesign(N) (25)

[0117] Wherein, the initial sampling point number N is set to 400.

[0118] Through the obtained sampling point information, a series of automatic processes such as drawing grids, importing aerodynamic simulation software to solve aerodynamic forces, and collecting aerodynamic data as shown in Figure 3 are obtained. The lift-drag coefficient aerodynamic coefficient data is the lift coefficient C L and the drag coefficient C D .

[0119] In the process of solving aerodynamic force data, the correctness of the collected aerodynamic force data is verified by grid independence test, solving aerodynamic force data of similar working condition aircraft and comparing with experimental data.

[0120] Step B-2: Since the established variant aerodynamic model has high dimension of design variables, and the polynomial expression form between the lift and drag coefficients and each variable cannot be determined, Kriging surrogate models are constructed for the lift coefficient C L and the drag coefficient C D respectively.

[0121]

[0122] wherein C and C are the Kriging surrogate model fitting functions of the lift and drag coefficients respectively.

[0123] In order to determine the accuracy of the aerodynamic performance prediction of the surrogate model, leave-one-out validation is performed on the established Kringing surrogate model, and the complex correlation coefficients are given:

[0124]

[0125] wherein C Li and C Di are the lift and drag coefficients obtained by the surrogate model at each sampling point; if any complex correlation coefficient is less than 0.8, return to step B-1 to increase the number of sampling points; otherwise, proceed to step B-3.

[0126] The complex correlation coefficient R 2 of the drag coefficient surrogate model is 0.891, and the complex correlation coefficient R 2 of the lift coefficient surrogate model is 0.919, both of which are above 0.8, and it is considered that the surrogate model can provide relatively accurate aerodynamic force prediction and realize complete data support.

[0127] Step B-3: According to the flight environment characteristics experienced by the hypersonic variable aircraft during the reentry phase, by describing the coordinate system and the atmospheric model, the state variables are selected as the rotating earth relative velocity v, the flight path angle γ, the heading angle ψ, the longitude λ, the latitude φ, and the altitude H, and the six-degree-of-freedom trajectory model of the hypersonic variable aircraft is constructed.

[0128]

[0129] wherein the earth radius R e is 6378km, the gravitational acceleration g is 9.81m / s 2 ; the roll angle σ is 0; L is the lift acceleration, D is the drag acceleration, which needs to be calculated; C v2 is the relative velocity v direction component of the dependent acceleration term, and Cγ1 and C γ2 are the direction components of the flight path angle γ of the Coriolis acceleration term and the coupled acceleration term, respectively, C ψ1 and C ψ2 are the direction components of the heading angle ψ of the Coriolis acceleration term and the coupled acceleration term, respectively, and the expressions are:

[0130]

[0131] where ω e is the Earth rotation angular velocity of 7.29e-5 rad / s

[0132] The lift acceleration L and the drag acceleration D are given by the aerodynamic data obtained from equation (3).

[0133]

[0134] where q = 0.5 ρv 2 2 is the dynamic pressure, ρ is the atmospheric density given by the standard atmospheric model, and S is the reference area of the wing.

[0135]

[0136] where ρ0 = 1.225 kg / m 3 2, h s = (RT0) / g0 is the atmospheric height dependent constant with a value of 7254.2 m.

[0137] Step C: Determine the pointing objects of the state quantity, action quantity, constraints on the body, and initial and final conditions in the reinforcement learning algorithm; determine the form of the reward function in the reinforcement learning algorithm and the experience pool composed of expert experience and exploration experience; determine the network architecture setting and initial quantity assignment in the reinforcement learning algorithm; finally give the training steps of the reentry trajectory optimization model driven by the reinforcement learning algorithm, complete the training, and the process is shown in Figure 5 .

[0138] The implementation method of step C is as follows:

[0139] Step C-1: Set the inner wing sweep angle χ1, the outer wing sweep angle χ2, the altitude H, the flight path angle γ, the Mach number Ma, and the angle of attack α as the required basic element state quantity parameters s in the DDPG reinforcement learning algorithm.

[0140] s = [χ1, χ2, H, γ, Ma, α] (32)

[0141] Set the inner wing sweep angle change amount Δχ1 and the outer wing sweep angle change amount Δχ2 as the required basic element action quantity a in the DDPG reinforcement learning algorithm.

[0142] a = [Δχ1, Δχ2] (33)

[0143] The angle of attack a, which is a key parameter for the aerodynamic effect of the hypersonic morphing aircraft, is not included in the training of the reinforcement learning DDPG algorithm. The angle of attack changes with the speed in stages as follows:

[0144]

[0145] where the maximum angle of attack a max is 20°, the angle of attack a max(K) under the maximum lift-drag ratio is 12°, the initial speed V1 of the maximum angle of attack is 5500 m / s, the speed V2 corresponding to the maximum lift-drag ratio is 5000 m / s, and the remaining parameters are as follows:

[0146]

[0147] The maximum heat flow constraint, the overload constraint, and the dynamic pressure constraint related constraint conditions are considered. Based on the standard resistance acceleration profile, the reentry boundary constraint model is given.

[0148]

[0149] where the maximum heat flow constraint q max is 1.5×10 7 W / m 2 , the overload constraint n′ max is 2, the dynamic pressure constraint q′ max is 90 KPa, D q,max is the resistance acceleration under the maximum heat flow constraint, D n′,max is the resistance acceleration under the overload constraint, and D q′,max is the resistance acceleration under the dynamic pressure constraint.

[0150] Considering the flight characteristics, flight environment characteristics, and body restrictions of the hypersonic morphing aircraft during the reentry flight process, the body lift-off parameters are given as follows: the lift-off speed v0 is 6200 m / s, the lift-off height H0 is 60 km; the initial configuration parameters of the body are the inner wing segment sweepback angle χ 10 = 18° and the outer wing segment sweepback angle χ 20 = 0°; the initial attitude parameters of the body are the track inclination angle γ0 = 0° and the heading angle ψ0 = 0°; the body position parameters are the longitude λ0 = 0° and the latitude φ0 = 0°; the remaining parameters are the aircraft mass m = 5000 kg, the wing reference area S = 5.47 m 2 , the earth radius R e , the rotation speed ω e , and the earth surface gravity acceleration g.

[0151] The reentry end condition is adopted to terminate the training, and the termination speed v s = 3000 m / s and the termination height H s = 30 km are set.

[0152] Step C-2: Establish the reward function R required for reinforcement learning training, and build an experience pool combined with expert experience.

[0153] First, considering the target function of maximum lift-drag ratio, the smoothness requirement of deformation amount and the related constraint requirements, the reward function R is established.

[0154] R = λ1C L / D - λ2(Δχ1+Δχ2)- λ3ΔD q,n′,q′ (37)

[0155] Wherein, the preset weight coefficients λ1=1, λ2=0.2, λ3=10, C L / D is the lift-drag ratio, ΔD q,n′,q′ is the heat flow, overload, dynamic pressure constraint exceeding amount.

[0156] For the important component of experience pool setting in the reinforcement learning DDPG algorithm, considering the characteristics of large state space span and complex variables in the reentry flight process of hypersonic variable aircraft, and the long continuous action makes the action space also very large, simple exploration learning is difficult to enable the agent to effectively recognize the whole process. The experience pool combined with expert experience is established for training.

[0157] In the training process, the experience pool sampled by training is divided into exploration experience E explore and expert experience E expert , and the ratio parameter α e of the total sampling occupied by expert experience sampling is introduced 0.2

[0158]

[0159] Wherein, the total number M = M(E expert )+M(E explore ) sampled in the corresponding experience pool is 1024.

[0160] The exploration experience E explore in the experience pool is the experience formed by the output action of the agent in the training round and the interaction with the environment, which is stored in each training round until the upper limit of the set experience pool capacity is reached. And the expert experience E expert in the experience pool is the experience of the two control quantities of the inner and outer wing segment sweep angle that realizes the maximum lift-drag ratio under the current state quantity.

[0161] In the round of collecting expert experience, the action is still output by the neural network and executed, and at the same time, the sweep angle corresponding to the maximum lift-drag ratio is found by SLSQP under the aerodynamic proxy model at each time length, and the expert experience of this round is stored.

[0162] Step C-3: Give the reinforcement learning network architecture setting.

[0163] The action network Actor is established by two initial hyperparameters and networks Actor_e and Actor_t with the same structure, the former gives the output action a when the input state s, and the latter outputs the next action a' when the input next state s'. The update of Actor_e adopts normal update, and the update of Actor_t is in the form of soft update. The network architecture includes five layers of input layer, three layers of hidden layer and output layer, except that sigmoid is used as the activation function between the last two layers, and relu is used as the activation function between the other layers, and the connection mode is full connection between neurons. The number of input neurons is set to 6, and the number of output neurons is set to 2.

[0164] The evaluation network Critic is established by two initial hyperparameters and networks Critic_e and Critic_t with the same structure, which give the output value function when the input state s and action a. The update of Critic_e adopts normal update, and the update of Critic_t is in the form of soft update. The network architecture includes five layers of input layer, three layers of hidden layer and output layer, except that sigmoid is used as the activation function between the last two layers, and relu is used as the activation function between the other layers, and the connection mode is full connection between neurons. The number of input neurons is set to 8, and the number of output neurons is set to 1.

[0165] The normal update adopts the weight training form of gradient descent method, and the weight training is as follows:

[0166]

[0167] Wherein, L is the damage function, θ is the network hyperparameter, E is the expectation function, R is the reward function, γ e is the discount factor, and Q is the value function.

[0168] The normal update form is as follows:

[0169]

[0170] The soft update coefficient τ is set to 0.2, the original network hyperparameter θ and the updated hyperparameter θ', and the update form is as follows:

[0171] θ'←τθ+(1-τ)θ',τ≤1 (41)

[0172] After the network architecture setting is completed, the initial state value s0 and the initial position state value X0 are determined according to the initial data given in step C-1:

[0173]

[0174] where Ma0 and a0 are calculated from v0.

[0175] Step C-4: reentry trajectory optimizer training under the complete DDPG algorithm architecture is performed;

[0176] The network parameter training steps are as follows:

[0177] At time t, there is an action value a t-Δt , a state value s t and a position state value X t :

[0178]

[0179] where subscript t represents the value of the variable at time t.

[0180] Substitute variables H t , a t , Ma t , x 1t , x 2t into equation (3) to obtain the aerodynamic coefficients C Lt , C Dt at the current time; substitute the aerodynamic coefficients C Lt , C Dt at the current time into equation (7) to obtain the lift acceleration L t and the drag acceleration D t at the current time; substitute L t , D t and X t into equation (5) to obtain the position state value X t+Δt at the next time using a Runge-Kutta 4th order integrator, and Δt is a specific time step set; substitute v t+Δt in the position state value X t+Δt into equation (10) to obtain a t+Δt , and into the standard model to obtain Ma t+Δt ; substitute s t into the network Actor to obtain a t , and a t is as follows:

[0181] a t = [Δx 1t , Δx 2t ] (44)

[0182] There is a relationship:

[0183]

[0184] Substitute equation (21) into equation (20), and use the 4th order Runge-Kutta integrator to obtain the next time χ t ,χ 1t+Δt ,χ 2t+Δt ,χ t+Δt ,χ t+Δt ,χ t+Δt ,χ t+Δt ; the action value a t and the state value s t are input into the network Critic to obtain the value function, and the network Actor is updated to complete the complete training process at time t; at t = 0, the following equation is obtained:

[0185]

[0186] After the initial value is input to start training, the network parameter training step is repeatedly performed until the training end condition is met, i.e.:

[0187]

[0188] After the training condition is met, the training process is recorded as one training, and the training number i is i + 1; the upper limit of the training iteration i max is set to 5000, and the training is ended when i = i max , and an optimized trajectory of a hypersonic variable aircraft reentry driven by reinforcement learning is output, thereby forming a method for optimizing the trajectory of a hypersonic variable aircraft reentry driven by reinforcement learning.

[0189] Step D: considering the hypersonic variable aircraft reentry under different end conditions, the specific characteristic parameters of the aircraft are determined, the trajectory optimization model driven by the DDPG algorithm trained in step C is used for trajectory optimization, and the variable decision is given in real time. After that, the flight benefits of the established optimization model are highlighted by comparing the simulation analysis of the remaining methods under the same conditions.

[0190] The implementation method of step D is as follows:

[0191] Step D-1: The trajectory is optimized for three different simulation conditions. The hypersonic variable aircraft has a small difference in trajectory due to the small difference in aerodynamic force in the initial descent segment, and the initial point is the vehicle separation point, which is transported by the booster rocket to the near space at an altitude of about 60 km. The simulation termination condition is that the hypersonic variable aircraft reaches the dive down pressure point determined by the task characteristics, and the aircraft is ready for dive down and enters the terminal guidance segment to reach the target task point. The simulation initial conditions are the same as the training initial conditions, and the end termination conditions are optimized for different trajectories, which gives the flight benefits of the hypersonic aircraft in the reentry segment flight process. The termination conditions are set to termination speed v1 = 2500, 3000, 3500 m / s, and the corresponding termination height H1 = 25, 30, 35 km.

[0192] Step D-2: In order to further explore the benefits obtained by the hypersonic variable aircraft in the reentry segment flight process by the established DDPG algorithm driven reentry segment trajectory optimization model, flight simulation based on various optimization methods is performed for comparison. For the optimization method, SLSQP, COBYLA and trust-constr are used to calculate the angle corresponding to the maximum lift-drag ratio under the established lift coefficient proxy model and drag coefficient proxy model, and the state quantity at the next time is calculated by substituting the angle into the motion equation, and the comparison of the optimization simulation results is given. For the direct control method, the pseudospectral method which performs well in solving trajectory optimization problems is used. The LG points are obtained by using the roots of N-order Legendre polynomials, and the flight results of the whole segment simulation are given under the same initial and termination conditions and flight state parameters as the previous simulation.

[0193] In order to better reflect the effectiveness and engineering practicability of the present application, the following takes the reentry segment trajectory optimization problem of the hypersonic variable aircraft as an example, and the present application is further described in combination with the drawings and tables.

[0194] In this case, the overall layout of the hypersonic variable aircraft is shown in Figure 1 wherein, Figure 1 (a) is a side view of the initial shape of the fuselage, Figure 1 (b) is a rear view of the initial shape of the fuselage, Figure 1 (c) is a top view of the initial shape of the fuselage, Figure 1 (d) is a top view of the whole machine half mold. The comparison of various parameters before and after trajectory optimization under different termination conditions is shown in Figures 5 to 7 wherein, Figures 5 to 7 (a) is an altitude-range graph, Figures 5 to 7 (b) is an altitude-velocity graph, Figures 5 to 7 (c) is a velocity-time graph, Figure 5 (d) is a body bank angle-time graph. The comparison of various parameters before and after trajectory optimization under different optimization algorithms is shown in Figures 9 to 11 wherein, Figures 9 to 11(a) is an altitude-range map, Figures 9 to 11 (b) is an altitude-velocity map, Figures 9 to 11 (c) is a velocity-time map, Figures 9 to 11 (d) is a body bank angle-time map. The range and lift-drag ratio before and after optimization of different termination conditions are shown in Table 1, the range and lift-drag ratio of different algorithms under the same termination condition are shown in Table 2, and the range and lift-drag ratio of the Gauss pseudospectral method configuration and other configurations are shown in Table 3.

[0195] Table 1 Comparison of range and lift-drag ratio of hypersonic variable aircraft before and after optimization

[0196]

[0197] Table 2 Comparison of range and lift-drag ratio of hypersonic variable aircraft before and after optimization of different algorithms

[0198]

[0199] Table 3 Comparison of range and lift-drag ratio of Gauss pseudospectral method configuration and other configurations

[0200]

[0201] The above optimization results show that the present application can obtain a trajectory of a hypersonic variable aircraft in reentry flight with a longer range at a smaller calculation cost, and can realize millisecond-level variable decision control.

[0202] The above specific description further describes the purpose, technical solutions and beneficial effects of the present application. It should be understood that the above description is only a specific embodiment of the present application, which is used to explain the present application and does not limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. within the spirit and principles of the present application should be included in the protection scope of the present application.

Claims

1. A method for hypersonic variable configuration vehicle re-entry trajectory optimization driven by reinforcement learning, characterized in that: Comprising the following steps: Step A: Establishing a hypersonic variable aircraft parameterized model; The method for realizing step A is as follows: Step A-1: Constructing a reference shape and a variable form; according to the reentry flight environment characteristics of the hypersonic variable aircraft, a configuration similar to a waverider body is adopted, that is, a waverider body fuselage combined with a hypersonic wing shape; due to the rarefied atmosphere during flight, a large-scale deformation of a variable back-swept angle is adopted to improve the variable performance; Step A-2: Parameterized description of the geometric model; based on the standard three-dimensional CST description equation combined with the shape characteristics of the waverider body fuselage, combined with the symmetric relationship of the leading edge line located in the horizontal plane and the longitudinal plane, the Z coordinate expression of the upper and lower surfaces of the initial shape of the fuselage is obtained, that is, the parameterized model of the hypersonic variable aircraft; where H u ′ is the spanwise section upper surface control parameter, M is the class function, u ′ is the body axis section upper surface control parameter, N u ′ is the upper surface thickness control parameter, η is the X, Y normalized coordinates, and subscript l is the lower surface related control parameter. Step B: Establishing a six-degree-of-freedom trajectory model of the hypersonic variable aircraft; The method for realizing step B is as follows: Step B-1: In the reentry flight environment of the hypersonic variable aircraft, the speed and atmospheric density changes need to be considered, so the aerodynamic model of the hypersonic variable aircraft is jointly influenced by the flight height H, the flight Mach number Ma, the attack angle a, the inner wing segment back-swept angle χ1, and the outer wing segment back-swept angle χ2; according to the flight environment and the structural characteristics of the aircraft, the Latin hypercube experimental design method is used to generate a predetermined number of sampling points within the variable range: [H,Ma,α,χ1,χ2]=lhsdesign(N) (2) In the formula, N is the number of sampling points; A series of automatic processes of drawing a grid through the obtained sampling point information, importing aerodynamic simulation software to solve aerodynamic force, and collecting aerodynamic force data are performed to obtain a plurality of groups of discrete lift-drag coefficient aerodynamic coefficient data; the lift-drag coefficient aerodynamic coefficient data are lift coefficient C L and drag coefficient C D ; Step B-2: Construct Kriging surrogate model for lift coefficient C L , drag coefficient C D , respectively; wherein Kriging surrogate model fitting functions for the lift and drag coefficients, respectively; The Kringing surrogate model established is subjected to leave-one-out validation LOOCV, and the complex correlation coefficient is given; wherein C Li , C Di are the lift and drag coefficients obtained by the proxy model at each sampling point; if any of the complex correlation coefficients is less than 0.8, return to step B-1 to increase the number of sampling points; otherwise, proceed to step B-3; Step B-3: According to the flight environment characteristics experienced by the hypersonic variable aircraft during the reentry flight, through the description of the coordinate system and the description of the atmospheric model, the state quantity is selected as the rotating earth relative velocity v, the track inclination angle γ, the heading angle ψ, the longitude λ, the latitude φ, and the altitude H, and the six-degree-of-freedom trajectory model of the hypersonic variable aircraft is constructed; where R e is the earth radius, g is the acceleration of gravity, which is taken as a nominal value; σ is the angle of bank, which is taken as 0; L is the lift acceleration, D is the drag acceleration, which needs to be calculated; C v2 is the relative velocity v direction component of the dependent acceleration term, C γ1 and C γ2 are the track inclination γ direction components of the Coriolis acceleration term and the dependent acceleration term, respectively, C ψ1 and C ψ2 are the heading ψ direction components of the Coriolis acceleration term and the dependent acceleration term, respectively, and the expressions are: where ω e is the angular velocity of the earth rotation, taken as the nominal value; The lift acceleration L and the drag acceleration D are given by using the aerodynamic data obtained in formula (3); where the dynamic pressure q = 0.5p v 2 , p is the atmospheric density given by the standard atmospheric model, and S is the wing reference area. Step C: Determining the pointing objects of the state quantity, the action quantity, the constraints on the body, and the initial and final conditions in the reinforcement learning algorithm; determining the form of the reward function in the reinforcement learning algorithm and the experience pool composed of expert experience and exploration experience; determining the network architecture setting and initial quantity assignment in the reinforcement learning algorithm; training the reentry trajectory optimization model driven by the reinforcement learning algorithm to obtain the trained reentry trajectory optimization model driven by the reinforcement learning algorithm, and performing trajectory optimization according to the trained reentry trajectory optimization model driven by the reinforcement learning algorithm; The method for realizing step C is as follows: Step C-1: Setting the inner wing segment back-swept angle χ1, the outer wing segment back-swept angle, the height H, the track inclination angle γ, the Mach number Ma, and the attack angle a as the required basic element state quantity parameters s in the DDPG reinforcement learning algorithm; s=[χ1,χ2,H,γ,Ma,α] (8) Setting the inner wing segment back-swept angle change amount Δχ1 and the outer wing segment back-swept angle change amount Δχ2 as the required basic element action quantity a in the DDPG reinforcement learning algorithm; a=[Δχ1,Δχ2] (9) The state quantity parameter attack angle a changes with the speed in stages as shown in formula (10): where α max is the maximum angle of attack, α max(K) is the angle of attack at maximum lift-drag ratio, V1 is the initial speed at maximum angle of attack, V2 is the speed corresponding to maximum lift-drag ratio, and the remaining parameters are as follows: According to the maximum heat flow constraint, overload constraint and dynamic pressure constraint related constraints, based on the standard resistance acceleration profile, the reentry boundary constraint model is given; where q max is the maximum heat flux constraint, n′ max is the g-load constraint, q′ max is the dynamic pressure constraint, D q,max is the drag acceleration under the maximum heat flux constraint, D n′,max is the drag acceleration under the g-load constraint, D q′,max is the drag acceleration under the dynamic pressure constraint; According to the flight characteristics, flight environment characteristics and the body itself limitation of the hypersonic variable aircraft in the reentry phase flight process, the body lift-off parameters lift-off speed v0, lift-off height H0 are given; the initial configuration parameters inner wing segment back sweep angle χ 10 , outer wing segment back sweep angle χ 20 ; the body initial attitude parameters track angle γ0, heading angle ψ0; the body initial position parameters longitude λ0, latitude φ0; the rest parameters aircraft mass m, wing reference area S, earth radius R e , rotation speed ω e and earth surface gravity acceleration g; The training is terminated by taking the reentry segment end condition, setting a termination velocity v s , a termination altitude H s ; Step C-2: Establish the reward function R required for reinforcement learning training, and build an experience pool combining expert experience according to the reward function R; Considering the target function of maximum lift-drag ratio, the smoothing requirement of deformation amount and the related constraint requirement, the reward function R is established; R λ1C L / D λ2(χ1 Δχ2)λ3ΔD q,n′,q′ (13) Wherein, λ1, λ2, λ3 are preset weight coefficients, C L / D is the lift-drag ratio, ΔD q,n′,q′ is the heat flow, overload, dynamic pressure constraint exceeds the amount; During training, the training experience pool is divided into exploration experience E explore and expert experience E expert , and a parameter a e is introduced as the ratio of the total sampling of expert experience sampling Wherein, M is the number of sampling in the corresponding experience pool; Exploration experience E in the experience pool explore For the experience formed by the agent outputting actions from the neural network and interacting with the environment in the training round, each round of training is stored until the upper limit of the set experience pool capacity is reached and updated. Expert experience E in the experience pool expert Since the lift-drag ratio is the main reward, the experience of the inner and outer wing segment sweep angles of the two control quantities that achieve the maximum lift-drag ratio under the current state quantity is set. Step C-3: Give the reinforcement learning network architecture setting; An action network Actor and an evaluation network Critic are established, the network Critic outputs the value function when inputting the state quantity s and the action quantity a, which is used to update the network Actor; The network Actor is composed of two initial hyperparameters and two networks Actor_e and Actor_t with the same structure, Actor_e gives the output action quantity a when inputting the state quantity s; Actor_t gives the output next action quantity a' when inputting the next state quantity s'; After the network architecture setting is completed, the initial value s0 of the state and the initial value X0 of the position state are determined according to the initial data given in step C-1: Wherein, Ma0 and a0 are derived and calculated from v0; Step C-4: Perform complete trajectory optimizer training in the reentry section under the DDPG algorithm architecture; The network parameter training steps are as follows: At time t, there is an action value a t-Δt , a state value s t and a position state value X t : Wherein, the subscript t represents the value of the variable at time t; variable H t ,α t Ma t ,χ 1t ,χ 2t Substitute into equation (3) to obtain the aerodynamic coefficient C at the current moment. Lt C Dt ; the aerodynamic coefficient C at the current moment Lt C Dt Substitute into equation (7) to obtain the lift acceleration L at the current moment. t D t ; will L t D t With X t Substituting into equation (5), the next position state value X is obtained using a Runge-Kutta fourth-order integrator. t+Δt Δt is the set specific time step; the position status value X t+Δt Chinese v t+Δt Substitute into equation (10) to obtain α t+Δt Substitute into the standard model to obtain Ma t+Δt ; will s t Get a from the network Actor t a t As shown in the following formula: a t = [Δχ 1t , Δχ 2t ] (17) The relationship is obtained: a t Substituting into equation (18), the next time step χ is obtained using a Runge-Kutta fourth-order integrator. 1t+Δt ,χ 2t+Δt Integrate location status value X t+Δt The parameter H has been obtained. t+Δt ,γ t+Δt Given s t+Δt ; Set action value a t With state value s t The input is fed into the network Critic, which provides the value function and updates the network Actor, completing the training process at time t. At t=0, the following equation holds: After the initial value is transmitted to start training, the network parameter training steps are repeatedly performed until the training end condition is reached, that is: After the training condition is met, the training process is recorded as one training, and the training number i is i+1; the upper limit of the training iteration i is set max When i=i max , the training is ended, and a trained reinforcement learning algorithm driven reentry phase trajectory optimization model is output, and trajectory optimization is performed according to the trained reinforcement learning algorithm driven reentry phase trajectory optimization model.

2. The method of claim 1, wherein: It also includes step D: considering the hypersonic variable aircraft reentry section flight under different terminal conditions, determining the specific characteristic parameters of the aircraft, and performing trajectory optimization according to the reentry trajectory optimization model driven by the trained reinforcement learning algorithm in step C, improving the accuracy and efficiency of the reentry trajectory optimization, and realizing high-precision variable decision control of the hypersonic variable aircraft in the reentry flight process.

3. The method of claim 1, wherein: When solving the aerodynamic force data in step B, the grid independence test is performed, the aerodynamic force data of the aircraft under similar working conditions is solved, and the correctness of the collected aerodynamic force data is verified by comparing the experimental data, and the correct aerodynamic force data is obtained.

4. The method of claim 1, wherein: The network architecture of Actor in C-4 includes five layers of input layer, three layers of hidden layer and output layer, except that sigmoid is used as the activation function between the last two layers, and relu is used as the activation function between the other layers, and the connection mode is full connection between neurons and neurons; The number of input neurons is set to the dimension of the state quantity, and the number of output neurons is set to the dimension of the action quantity.

5. The method of claim 1, wherein: The evaluation network Critic in C-4 is composed of two initial hyperparameters and networks with the same structure, namely Critic_e and Critic_t, which both give output value functions when inputting state quantity s and action quantity a; the update of Critic_e adopts normal update, and the update of Critic_t is in the form of soft update; the network architecture includes five layers of input layer, three layers of hidden layer and output layer, among which sigmoid is used as the activation function between the last two layers, and relu is used as the activation function between the other layers, and the connection mode is full connection between neurons; the number of input neurons is set to the total dimension of action quantity and state quantity, and the number of output neurons is the dimension of value function.

6. The method of claim 1, 4 or 5, wherein: The update of Actor_e adopts normal update, and the update of Actor_t is in the form of soft update; The normal update adopts the weight training form of gradient descent method, and the weight training is as follows: where L is a loss function, θ is a network hyper-parameter, E is an expectation function, R is a reward function, γ e is a discount factor, and Q is a value function. The normal update form is as follows: The soft update coefficient τ is set, the original network hyperparameter θ and the updated hyperparameter θ' are set, and the soft update form is as follows: θ'←τθ+(1-τ)θ',τ≤1 (23).

Citation Information

Patent Citations

  • Hypersonic aircraft reentry cooperative guidance method based on reinforcement learning

    CN114675545A

  • Hypervariant aircraft trajectory planning method based on penalty function sequence convex optimization

    CN117032275A