Reinforcement learning rocket load shedding guidance law design method considering time-varying characteristics of wind field

Through reinforcement learning and Markov decision-making process model, the rocket attitude angle instructions are corrected in real time, and the adaptability of time-varying characteristics of wind field in the rocket load reduction guidance law design is solved, achieving the reduction of aerodynamic load of the arrow body and the improvement of rocket reliability.

CN120372803AActive Publication Date: 2025-07-25CHINESE PEOPLES LIBERATION ARMY UNIT 63810
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510422153.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-05
Publication Date
2025-07-25
Estimated Expiration
2045-04-05

AI Technical Summary

Technical Problem

The existing rocket load reduction guidance law design method cannot adapt to the time-varying characteristics of the actual wind farm and the forecast wind farm, resulting in limited load reduction effect, affecting the reliability and carrying capacity of the rocket.

Method used

Through reinforcement learning methods, high-resolution high-altitude wind hourly forecast data and rising stage guidance Markov decision-making process model are used to correct the rocket's attitude angle command in real time, optimize the rocket load reduction guidance strategy, and adapt to the time-varying characteristics of the wind field.

Benefits of technology

When there is a deviation between the forecast wind farm and the actual wind farm, the rocket's attitude command will be corrected in real time to reduce the maximum aerodynamic load on the arrow body, and improve the reliability and carrying capacity of the rocket.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372803A_ABST
    Figure CN120372803A_ABST
Patent Text Reader

Abstract

The invention discloses a reinforcement learning rocket load shedding guidance law design method considering wind field time-varying characteristics, solves the problem that the existing rocket load shedding guidance law design cannot adapt to the deviation between an actual wind field and a forecast wind field, and belongs to the field of rocket load shedding guidance control. Comprising the following steps: generating high-resolution high-altitude wind hour-by-hour forecast data; constructing a rising stage guidance Markov decision process model; using high-resolution high-altitude wind hour-by-hour forecast data to simulate the uncertainty deviation between a forecast wind field and an actual wind field through a model; a rocket deloading guidance strategy is adopted to complete interaction sampling of a complete round of a rising stage guidance Markov decision process; calculating an updating gradient according to the sampling data, updating parameters of a rocket load shedding guidance strategy according to the updating gradient, optimizing the uncertainty deviation, and obtaining a rocket load shedding guidance law which enables the maximum atmospheric dynamic load expectation to be minimum in the whole process of the ascending section; according to the method, real-time correction of the rocket attitude instruction under the condition that the forecast wind field and the actual wind field have deviation is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of load reduction guidance and control for launch vehicles, and relates to a design method for a reinforcement learning rocket load reduction guidance law considering the time-varying characteristics of the wind field. Background Art

[0002] For a heavy launch vehicle with a high slenderness ratio, the internal force bending moment caused by the aerodynamic load during its ascent in the atmosphere is likely to cause the structural instability or even damage of the rocket body, reducing the reliability of the rocket. At the same time, in the structural design of the launch vehicle, restricted by the aerodynamic load, a conservative design is generally adopted, increasing the structural mass of the launch vehicle and reducing its carrying capacity. With the development of launch vehicle launches towards low cost, high frequency, and high reliability, it is necessary to specifically design the load reduction guidance law for the ascent stage of the launch vehicle, that is, the generation law for calculating the expected attitude of the rocket in real time from the rocket state variables, so as to reduce the aerodynamic load on the rocket.

[0003] The existing load reduction guidance law design relies on accurate wind field prediction data: specifically, based on the wind field at the launch moment generated by the prediction, a trajectory optimization method is used to regenerate the rocket attitude angle command that satisfies the maximum aerodynamic load constraint under this wind field.

[0004] However, the actual wind field has typical time-varying characteristics, that is, the wind speed and direction at each altitude have uncertain changes over time, which makes the existing load reduction guidance law design method that regenerates the rocket attitude angle command using a fixed predicted wind field unable to adapt to the deviation between the actual wind field and the predicted wind field, resulting in limited load reduction effect. Summary of the Invention

[0005] To solve the above technical problems, the present invention provides a design method for a reinforcement learning rocket load reduction guidance law considering the time-varying characteristics of the wind field. This method is applicable to the load reduction guidance and control of launch vehicles considering the time-varying characteristics of the wind field; through reinforcement learning, it can real-time correct the rocket attitude angle command during the rocket flight, adapt to the uncertainty deviation between the actual wind field and the predicted wind field, thereby reducing the maximum aerodynamic load on the rocket body and improving the reliability of the rocket.

[0006] The present invention builds a training framework for the rocket load reduction wind correction guidance strategy that adapts to the deviation between the predicted wind field and the actual wind field through high-resolution hourly prediction data generation of high-altitude wind, construction of a Markov decision process model for ascent stage guidance, simulation of wind field time-varying uncertainty, and gradient optimization of rocket load reduction guidance strategy update, realizing real-time correction of the rocket attitude command in the case of deviation between the predicted wind field and the actual wind field, thereby reducing the maximum aerodynamic load on the rocket body. This method is applicable to the load reduction guidance and control of launch vehicles considering the time-varying characteristics of the wind field.

[0007] The object of the present invention is specifically realized through the following technical solutions:

[0008] The present invention discloses a design method for a reinforcement learning rocket load reduction guidance law considering the time-varying characteristics of the wind field. The method includes:

[0009] Step 1: Obtain the target data required for the rocket load reduction guidance law, and generate high-resolution hourly forecast data of upper-air winds based on the WRF model.

[0010] Step 2: Through the centroid motion dynamics equation of the rocket during the ascent stage, take a complete flight process of the rocket during the ascent stage within the atmosphere as one episode of the ascent-stage guidance Markov decision-making process, and construct an ascent-stage guidance Markov decision-making process model.

[0011] Step 3: Record the total number of wind fields generated by the high-resolution hourly forecast data of upper-air winds. Based on the high-resolution hourly forecast data of upper-air winds, at the beginning of each episode of the ascent-stage guidance Markov decision-making process, randomly generate a wind field number within the total number of wind fields, select the wind field function in the ascent-stage guidance Markov decision-making process model corresponding to this number as the forecast wind field for this episode, and calculate the difference between the forecast wind field and the actual wind field to obtain the uncertainty deviation.

[0012] Step 4: Introduce a value function network in the policy network to represent the rocket load reduction guidance policy; interact through the rocket load reduction guidance policy and the ascent-stage guidance Markov decision-making process model to complete the sampling of a complete episode of the ascent-stage guidance Markov decision-making process and obtain sampling data; estimate the advantage function based on the sampling data, calculate the update gradient from the advantage function, update the parameters of the rocket load reduction guidance policy with the update gradient, and optimize the uncertainty deviation based on the rocket load reduction guidance policy after parameter update to obtain a rocket load reduction guidance law that minimizes the expected value of the maximum aerodynamic load throughout the ascent stage.

[0013] In Step 1, the generation method of the high-resolution hourly forecast data of upper-air winds includes:

[0014] Obtain the target data required for the rocket load reduction guidance law, where the target data is global forecast model data, terrain data, and observation data in the same time period.

[0015] Use dynamic downscaling technology to process the global forecast model data and terrain data received by the WRF model to obtain the boundary field and initial field required for driving the WRF model.

[0016] Combine the three-dimensional variational data assimilation method 3DVAR to perform multi-source heterogeneous data assimilation on the real-time upper-air wind sounding data, wind profiler radar data, and microwave radiometer data in the observation data, and update the initial field.

[0017] Start the forward integration of the WRF model using the updated initial field to obtain hourly forecast data of upper-air winds for the next 72 hours at the current time; correct the hourly forecast data of upper-air winds for the next 72 hours at the current time using the model result correction method in combination with the observed data to generate high-resolution hourly forecast data of upper-air winds.

[0018] In step one, the cost function used by the three-dimensional variational data assimilation method 3DVAR is:

[0019]

[0020] B = (x b - x true )(x b - x true ) T ;

[0021] In the formula, J(x) is the solution of the cost function, where the minimum solution is the analysis field, that is, the updated initial field; x represents the analysis variable at the model grid point, x b represents the background field, B represents the background field error covariance matrix, y0 represents the observed data, Hx represents the observation operator, R represents the observation error covariance matrix, the superscripts T and -1 represent the transpose and inverse of the matrix respectively, and x true represents the true field.

[0022] In step one, the model result correction method is:

[0023]

[0024] In the formula, is the corrected field at time t1 + n, where t1 is the time of the observed data, and n = 1, 2, 3,... is the difference between the correction time and the time of the observed data; b0, b1, and b2 are the first model coefficient, the second model coefficient, and the third model coefficient automatically established using the least squares method; is the WRF model forecast result at time t1 + n; is the error between the observed data at time t1 and the WRF model forecast result.

[0025] In step two, through the dynamic equation of the centroid motion of the rocket during the ascending stage, a complete flight process of the rocket in the atmosphere during the ascending stage is regarded as a round of the ascending stage guidance Markov decision process. The method for constructing the ascending stage guidance Markov decision process model is:

[0026] Construct the dynamic equation of the centroid motion of the rocket during the ascending stage from the current flight time, state variables, control variables, and wind field function of the rocket; among them, the constructed dynamic equation of the centroid motion of the rocket during the ascending stage is: In the formula, is the time derivative of the centroid motion state quantity during the rocket ascent stage; t is the current flight time of the rocket; x rocket is the centroid motion state quantity during the rocket ascent stage; u is the centroid motion control quantity during the rocket ascent stage; V wind (h) is the wind field function varying with altitude;

[0027] Select the state quantity x in the centroid motion dynamics equation during the rocket ascent stage rocket as the state quantity s of the ascent guidance Markov decision process model k ;

[0028] Select the control quantity u in the centroid motion dynamics equation during the rocket ascent stage as the action quantity a of the ascent guidance Markov decision process model k ;

[0029] Determine the discrete time step dT of the ascent guidance Markov decision process. Under the assumption that the wind field is known, discretize the continuous centroid motion dynamics equation of the launch vehicle ascent stage according to the determined discrete time step. Then, the state transition probability p(s k+1 |s k ,a k ) of the ascent guidance Markov decision process is the conditional probability p(V wind (h)) with respect to the wind field function V wind ;

[0030] Take the maximum aerodynamic load Q k |α Tk | as the reward function r(s k ,a k );

[0031] Then, a complete flight process of the rocket in the atmosphere ascent stage is a round of the ascent guidance Markov decision process, and the ascent guidance Markov decision process model is obtained; among them, the ascent guidance Markov decision process model is:

[0032] s k =x rocket | t=kdT ;

[0033] a k =u| t=kdT ;

[0034]

[0035] where k is the k-th discrete moment of the Markov decision process, s k+1 is the state quantity at the (k + 1)-th discrete moment, and k + 1 represents the (k + 1)-th discrete moment of the Markov decision process; N is the total length of the round of the ascent guidance Markov decision process; Qk |α Tk |represents the maximum aerodynamic load in the time interval from kdT to (k + 1)dT.

[0036] In step three, record the total number of wind fields generated for the high-resolution hourly upper-air wind forecast data. Based on the high-resolution hourly upper-air wind forecast data, at the start of each episode of the ascent-phase guidance Markov decision process, randomly generate a wind field number within the total number of wind fields, and select the wind field function in the ascent-phase guidance Markov decision process model corresponding to this number as the forecast wind field for this episode. The method for calculating the difference between the forecast wind field and the actual wind field to obtain the uncertainty deviation includes:

[0037] S1, record the total number of wind fields M generated for the high-resolution hourly upper-air wind forecast data;

[0038] S2, at the start of each episode of the ascent-phase guidance Markov decision process, randomly generate a wind field number j within the range [1, M]; in this episode, regard the wind field function V wind (h) as the j-th wind field in the high-resolution hourly upper-air wind forecast data; and regard the j-th wind field as the forecast wind field for this episode;

[0039] S3, repeat step S2 to obtain all the forecast wind fields, calculate the differences between all the forecast wind fields and the actual wind field to obtain the uncertainty deviation, and store it.

[0040] In step four, introduce a value function network in the policy network to represent the rocket load reduction guidance strategy. By interacting the rocket load reduction guidance strategy with the ascent-phase guidance Markov decision process model, the method for completing the sampling of a complete episode of the ascent-phase guidance Markov decision process to obtain the sampling data includes:

[0041] Introduce a value function network in the policy network to represent the rocket load reduction guidance strategy, and generate a set of sampling trajectories by interacting the rocket load reduction guidance strategy with the ascent-phase guidance Markov decision process model;

[0042] Record the output probability of the policy network and the output data of the value function network corresponding to each sampling trajectory in the set of sampling trajectories; at the same time, calculate the temporal difference residual corresponding to each sampling trajectory, and obtain the general advantage estimate of the temporal difference residual through recursive calculation; sum the general advantage estimate and the output data of the value function network to obtain the target value to be fitted by the value function network, and complete the sampling of a complete episode of the ascent-phase guidance Markov decision process;

[0043] The sampling data consists of the output data of the value function network, the target value to be fitted by the value function network, the general advantage estimate, and the output probability of the policy network.

[0044] In step 4, the sampling trajectory set contains N step = N traj × N sampling trajectories where N step is the total number of state transition steps, N traj is the number of sampling trajectories of a complete round, N is the total length of a round of the ascending segment guidance Markov decision process; i is the sampling trajectory serial number;

[0045] The temporal difference residual is:

[0046]

[0047] where r k,i is the reward at the k-th discrete moment of the i-th sampling trajectory, γ is the discount factor of the cumulative reward, is the output data of the value function network, is the output data of the value function network at the (k + 1)-th discrete moment;

[0048] The generalized advantage estimation is:

[0049]

[0050] where l is the summation serial number of the generalized advantage estimation, N is the total length of a round of the ascending segment guidance Markov decision process, k is the k-th discrete moment of the Markov decision process, λ is the generalized advantage estimation decay coefficient, is the temporal difference residual at the (k + l)-th discrete moment;

[0051] The target value to be fitted by the value function network is:

[0052]

[0053] In step 4, the advantage functions estimated based on the sampling data include: the value function network advantage function and the policy network advantage function; among them,

[0054] The value function network advantage function is:

[0055]

[0056] The policy network advantage function is:

[0057]

[0058] where is the value of the value function network advantage function, is the set of sampling trajectories, k is the k-th discrete time of the Markov decision process, and i is the sampling trajectory number. is the output data of the value function network. is the target value to be fitted by the value function network.

[0059] is the advantage function value of the policy network, and γ k is the discount factor of the cumulative reward at the k-th discrete time, and η k,i (θ) is the ratio of the policy behavior output probabilities at the k-th discrete time of the i-th sampling trajectory. is the Generalized Advantage Estimation, clip is the clipping function, and ∈ is the trust region radius for the gradient update of the policy network parameters.

[0060] The ratio of the policy behavior output probabilities at the k-th discrete time of the i-th sampling trajectory, η k,i (θ) is specifically expressed as:

[0061]

[0062] where is the probability of the current sampling trajectory output by the policy network. is the probability of all sampling trajectories output by the policy network.

[0063] In step four, the update gradients calculated by the advantage function include: the update gradient of the value function network parameters and the update gradient of the policy network parameters, where

[0064] The update gradient of the value function network parameters is:

[0065]

[0066] According to the update gradient of the value function network parameters with the update step perform gradient descent update on the value function network parameters in the rocket load reduction guidance strategy;

[0067] The update gradient of the policy network parameters is:

[0068]

[0069] According to the update gradient of the policy network parameters with the update step α θ perform gradient ascent update on the policy network parameters in the rocket load reduction guidance strategy;

[0070] In the formula, is the update gradient of the value function network parameters; is the gradient of the value function network value with respect to its parameters. The gradient for updating the policy network parameters.

[0071] The beneficial effects of the present invention are as follows:

[0072] In view of the load reduction requirements of a launch vehicle with a high slenderness ratio during the ascending flight in the atmosphere, the present invention utilizes high-resolution hourly prediction data of upper-air winds. Through the construction of a Markov decision process model for ascent guidance, the simulation of the time-varying uncertainty of the wind field, and the optimization of the gradient for updating the load reduction guidance strategy of the rocket, a training framework for the load reduction wind correction guidance strategy of the rocket that adapts to the deviation between the predicted wind field and the actual wind field is established. Through reinforcement learning, the attitude angle command of the rocket can be corrected in real time during the rocket flight, adapting to the uncertainty deviation between the actual wind field and the predicted wind field, realizing the real-time correction of the rocket attitude command in the case of an uncertain deviation between the predicted wind field and the actual wind field, thereby reducing the maximum aerodynamic load on the rocket body, improving the reliability of the rocket, and being applicable to the load reduction guidance control of launch vehicles considering the time-varying characteristics of the wind field. BRIEF DESCRIPTION OF THE DRAWINGS

[0073] The present invention will be further described in detail below with reference to the drawings and embodiments.

[0074] Figure 1 It is a schematic diagram of the system operation flow chart based on the WRF model.

[0075] Figure 2 It is a schematic diagram of the statistical average wind field data.

[0076] Figure 3 It is a schematic diagram of the range altitude profile of the nominal trajectory.

[0077] Figure 4 It is a schematic diagram of the vertical velocity profile of the nominal trajectory.

[0078] Figure 5 It is a schematic diagram of the horizontal velocity profile of the nominal trajectory.

[0079] Figure 6 It is a schematic diagram of the standard ballistic angle of attack profile.

[0080] Figure 7 It is a schematic diagram of the standard ballistic inclination angle profile. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0081] The reinforcement learning method used in the design method of the reinforcement learning rocket load reduction guidance law considering the time-varying characteristics of the wind field provided by the embodiment of the present invention is an effective framework for training the optimal policy in a stochastic environment. After a Markov decision process model related to the given task and the source of environmental uncertainty are given, the training policy is used to maximize the expected cumulative reward obtained in the environment with uncertain terms. The present invention trains a rocket load reduction guidance strategy under the reinforcement learning framework, and this method includes:

[0082] Step 1: Obtain the target data required for the rocket load reduction guidance law, and generate high-resolution hourly upper-air wind forecast data based on the WRF model;

[0083] Step 2: Through the dynamic equation of the centroid motion in the rocket ascent stage, take a complete flight process of the rocket in the atmosphere during the ascent stage as a round of the ascent stage guidance Markov decision-making process, and construct an ascent stage guidance Markov decision-making process model;

[0084] Step 3: Record the total number of wind fields generated by the high-resolution hourly upper-air wind forecast data. Based on the high-resolution hourly upper-air wind forecast data, at the start of each round of the ascent stage guidance Markov decision-making process, randomly generate a wind field number within the total number of wind fields, select the wind field function in the ascent stage guidance Markov decision-making process model corresponding to this number as the forecast wind field for this round, and calculate the difference between the forecast wind field and the actual wind field to obtain the uncertainty deviation;

[0085] Step 4: Introduce a value function network into the policy network to represent the rocket load reduction guidance strategy. Through the interaction between the rocket load reduction guidance strategy and the ascent stage guidance Markov decision-making process model, complete the sampling of a complete round of the ascent stage guidance Markov decision-making process to obtain sampling data; estimate the advantage function based on the sampling data, calculate the update gradient from the advantage function, update the parameters of the rocket load reduction guidance strategy with the update gradient, and optimize the uncertainty deviation based on the rocket load reduction guidance strategy after parameter update to obtain a rocket load reduction guidance law that minimizes the expected value of the maximum aerodynamic load throughout the ascent stage.

[0086] Among them, the policy network guides the decision-making of the intelligent agent and provides the probability of selecting an action under a given state. The value function network is used to evaluate the long-term value of a state or an action, and helps the intelligent agent learn and optimize its strategy. The policy network is a neural network used to model the strategy of the intelligent agent, that is, the probability distribution of selecting an action under a given state; the value function network is a neural network used to estimate the expected cumulative return that can be obtained under a given state or after taking a certain action.

[0087] In Step 1, the generation method of the high-resolution hourly upper-air wind forecast data includes:

[0088] Obtain the target data required for the rocket load reduction guidance law, and the target data is the global forecast model data, terrain data, and observation data at the same time period;

[0089] Use the dynamic downscaling technology to process the global forecast model data and terrain data received by the WRF model to obtain the boundary field and initial field required for driving the WRF model;

[0090] Combined with the three-dimensional variational data assimilation method 3DVAR, multi-source heterogeneous data assimilation is carried out on the real-time upper-air wind sounding data, wind profiler radar data, and microwave radiometer data in the observed data to update the initial field;

[0091] Use the updated initial field to start the forward integration of the WRF model to obtain the hourly upper-air wind forecast data for the next 72 hours at the current time; Use the model result correction method combined with the observed data to correct the hourly upper-air wind forecast data for the next 72 hours at the current time to generate high-resolution hourly upper-air wind forecast data.

[0092] The global forecast model data and terrain data are to obtain GFS, ECMWF, or GRAPES data from the corresponding data sources. GFS data can usually be downloaded from the National Centers for Environmental Prediction (NCEP) of the United States. ECMWF data needs to be obtained through ECMWF member states or through purchase. GRAPES data needs to be obtained through relevant research institutions of the China Meteorological Administration.

[0093] As Figure 1 shown, for a WRF model-based system in a specific region, hourly upper-air wind forecast data with a height distribution (layer thickness of 250 m) within the specific region can be obtained. The WRF model-based system is based on the WRF (Weather Research Forecasting) model. Using dynamic downscaling technology, the international / domestic mainstream numerical forecast results (GFS / ECMWF / GRAPES) are processed to obtain the boundary field and initial field required for model driving. Then, combined with the three-dimensional variational data assimilation method, multi-source heterogeneous data assimilation is carried out on the real-time upper-air wind sounding data, wind profiler radar data, and microwave radiometer to update the model initial field. The WRF model is integrated forward to obtain the hourly upper-air wind forecast data for the next 72 hours. Use the model result correction method combined with the real-time upper-air wind data to correct the hourly upper-air wind forecast data, and finally output high-resolution hourly upper-air wind forecast data.

[0094] For example: (1) Model data preparation:

[0095] After the target data is downloaded, perform target data format conversion. The WRF model uses a specific data format. Utilize the ungrib.exe tool in WRF and the processing system (WPS) to extract GRIB format data and convert it into an intermediate file that WRF can recognize. Then configure WPS to process the data provided to WRF. Edit the `namelist.wps` file to specify information such as the time range, spatial range, resolution, and vertical levels of the data. Run WPS, where `. / geogrid.exe`: Creates the model grid according to the configuration defined in `namelist.wps`. `. / ungrib.exe`: Decodes the GRIB data into an intermediate file. `. / metgrid.exe`: Merges the intermediate files at different times and interpolates them onto the model grid. Configure the WRF model, edit the `namelist.input` file, including the simulation time period, physical parameter options, dynamic framework options, etc. Ensure that the WRF model can read the `met_em*` files output from WPS. Run the WRF model, use `real.exe` to process the `met_em*` files output from WPS to generate the initial field and boundary conditions.

[0096] (2) Data assimilation:

[0097] Prepare upper-air wind sounding data, wind profiler radar data, and microwave radiometer data for the same time period. First, perform quality control and preprocessing on the observed data. Use historical extremes, spatial consistency, and basic physical laws to eliminate incorrect observed data. Then, use obsproc.exe in the WRF model to write the observed data in ASCII format recognized by the assimilation system.

[0098] Use WRFDA to perform three-dimensional variational assimilation. 3DVAR mainly solves for the minimum value of the objective function. It mainly obtains the analysis field (the solution of the minimum value) by solving, so that the best fitting effect is achieved between the analysis field and both the background field and the observed field. And at this time, the solution with the minimum cost function is the updated initial field.

[0099] In step one, the cost function used by the three-dimensional variational data assimilation method 3DVAR is:

[0100]

[0101] B = (x b - x true )(x b - x true ) T ;

[0102] In the formula, J(x) is the solution of the cost function. Among them, the minimum solution is the analysis field, that is, the updated initial field; x represents the analysis variable at the model grid points, xb denotes the background field, B denotes the background field error covariance matrix, y0 denotes the observed data, Hx denotes the observation operator, R denotes the observation error covariance matrix, the superscripts T and -1 represent the transpose and inverse of the matrix respectively, and x true denotes the true field.

[0103] (3) Simulation forecasting and forecast result correction:

[0104] After obtaining the updated initial field, run `wrf.exe` to start the simulation. After the simulation is completed, use the observed data, such as real-time upper-air wind sounding data, to correct the model results, and finally output high-resolution hourly upper-air wind forecast data.

[0105] In step one, the method for correcting the model results is:

[0106]

[0107] In the formula, is the corrected field at time t1 + n, where t1 is the time of the observed data, and n = 1, 2, 3,... is the difference between the correction time and the time of the observed data; b0, b1, and b2 are the first model coefficient, the second model coefficient, and the third model coefficient automatically established using the least squares method; is the WRF model forecast result at time t1 + n; is the error between the observed data at time t1 and the WRF model forecast result.

[0108] In step two, through the dynamic equation of the rocket's center-of-mass motion during the ascending stage, a complete flight process of the rocket in the atmosphere during the ascending stage is regarded as one round of the ascending-stage guidance Markov decision process. The method for constructing the ascending-stage guidance Markov decision process model is:

[0109] Construct the dynamic equation of the rocket's center-of-mass motion during the ascending stage from the rocket's current flight time, state variables, control variables, and wind field function; among them, the constructed dynamic equation of the rocket's center-of-mass motion during the ascending stage is: In the formula, is the time derivative of the state variable of the rocket's center-of-mass motion during the ascending stage; t is the rocket's current flight time; x rocket is the state variable of the rocket's center-of-mass motion during the ascending stage; u is the control variable of the rocket's center-of-mass motion during the ascending stage; V wind (h) is the wind field function varying with altitude;

[0110] Select the state variable x rocket in the dynamic equation of the rocket's center-of-mass motion during the ascending stage as the state variable s k of the ascending-stage guidance Markov decision process model;

[0111] The control variable u in the dynamic equation of the centroid motion during the rocket ascent stage is selected as the action variable a of the ascent-stage guidance Markov decision process model. k ;

[0112] Determine the discrete time step dT of the ascent-stage guidance Markov decision process. Under the assumption that the wind field is known, discretize the continuous dynamic equation of the centroid motion of the launch vehicle during the ascent stage according to the determined discrete time step. Then, the state transition probability p(s k+1 |s k , a k ) of the ascent-stage guidance Markov decision process is the conditional probability p(V wind (h)) with respect to the wind field function V wind (h);

[0113] The reward function of the ascent-stage guidance Markov decision process should represent the requirement of reducing the aerodynamic load. Therefore, the reward function is designed as the negative value of the maximum aerodynamic load during the entire ascent stage flight. Then, the maximum aerodynamic load Q k |α Tk | is used as the reward function r(s k , a k );

[0114] Then, a complete flight process of the rocket during the ascent stage within the atmosphere is one episode of the ascent-stage guidance Markov decision process, and the ascent-stage guidance Markov decision process model is obtained. Among them, the ascent-stage guidance Markov decision process model is:

[0115] s k = x rocket | t=kdT ;

[0116] a k = u| t=kdT ;

[0117]

[0118] In the formula, k is the k-th discrete moment of the Markov decision process, s k+1 is the state variable at the (k + 1)-th discrete moment, and k + 1 represents the (k + 1)-th discrete moment of the Markov decision process; N is the total length of the episode of the ascent-stage guidance Markov decision process; Q k |α Tk | represents the maximum aerodynamic load in the time interval from kdT to (k + 1)dT.

[0119] In step 3, record the total number of wind fields generated for the hourly high-resolution upper-air wind forecast data. Based on the hourly high-resolution upper-air wind forecast data, at the beginning of each episode of the ascent-phase guidance Markov decision process, randomly generate a wind field number within the total number of wind fields, and select the wind field function in the ascent-phase guidance Markov decision process model corresponding to this number as the forecast wind field for this episode. The method for calculating the difference between the forecast wind field and the actual wind field to obtain the uncertainty deviation includes:

[0120] S1. Record the total number of wind fields M generated for the hourly high-resolution upper-air wind forecast data;

[0121] S2. At the beginning of each episode of the ascent-phase guidance Markov decision process, randomly generate a wind field number j within the range of [1, M]; in this episode, use the wind field function V wind (h) in the ascent-phase guidance Markov decision process model as the j-th wind field in the hourly high-resolution upper-air wind forecast data; and use the j-th wind field as the forecast wind field for this episode;

[0122] S3. Repeat step S2 to obtain all forecast wind fields, calculate the difference between all forecast wind fields and the actual wind field to obtain the uncertainty deviation, and store it.

[0123] In step 4, introduce a value function network into the Policy Network to represent the rocket load reduction guidance strategy. By interacting the rocket load reduction guidance strategy with the ascent-phase guidance Markov decision process model, the method for completing the sampling of a complete episode of the ascent-phase guidance Markov decision process and obtaining the sampling data includes:

[0124] Introduce the value network in π θ in the policy network to represent the rocket load reduction guidance strategy where θ and θ v are both network parameters to be trained, and the total number of state transitions generated by the interaction between the rocket load reduction guidance strategy and the ascent-phase guidance Markov decision process model is N step The sampling trajectory set The total number of state transitions N in step contains N traj complete episode trajectories, so there is That is contains N step = N traj × N state transition data pairs of discrete sampling steps

[0125] Record the output probabilities of the policy network corresponding to each sampling trajectory in the sampling trajectory set and the output data of the value network At the same time, calculate each sampling trajectory The corresponding temporal difference residual And obtain the generalized advantage estimate of the temporal difference residual through recursive calculation The generalized advantage estimate And the output of the value network Sum them up to obtain the target value to be fitted by the value network Output by the value network And the target value to be fitted by the value network Generalized advantage estimate And the output probability of the policy network Form the sampling data

[0126] In step 4, the sampling trajectory set Is:

[0127]

[0128] In the formula, the sampling trajectory set Contains N step = N traj × N sampling trajectories N step Is the total number of state transitions, N traj Is the number of complete round sampling trajectories, N is the total length of the round of the ascending section guidance Markov decision process; i is the sampling trajectory serial number;

[0129] The temporal difference residual Is:

[0130]

[0131] In the formula, r k,i Is the reward at the k-th discrete moment of the i-th sampling trajectory, γ is the discount factor of the cumulative reward, Is the output data of the value network, Is the output data of the value function network at the (k + 1)-th discrete moment;

[0132] The generalized advantage estimate Is:

[0133]

[0134] In the formula, l is the summation serial number of the generalized advantage estimate, N is the total length of the round of the ascending section guidance Markov decision process, k is the k-th discrete moment of the Markov decision process, λ is the decay coefficient of the generalized advantage estimate, is the temporal difference residual at the (k + l)-th discrete time instant;

[0135] The target value to be fitted by the value network is:

[0136]

[0137] In step four, the method for estimating the advantage function based on the sampled data includes:

[0138] Based on the output of the value network in the sampled data and the target value to be fitted by the value network the value network advantage function for updating the value network parameters is obtained as:

[0139]

[0140] Based on the general advantage estimation in the sampled data and the output probability of the sampled policy network the policy network advantage function for updating the policy network parameters is obtained as:

[0141]

[0142] In the formula, is the value function network advantage function value, is the set of sampled trajectories, k is the k-th discrete time instant of the Markov decision process, i is the sampled trajectory number, is the output data of the value function network, is the target value to be fitted by the value function network;

[0143] is the policy network advantage function value, γ k is the discount factor of the cumulative reward at the k-th discrete time instant, η k,i (θ) is the ratio of the policy behavior output probabilities at the k-th discrete time instant of the i-th sampled trajectory, is the general advantage estimation, clip is the clipping function, ∈ is the trust region radius for updating the policy network parameter gradient;

[0144] The ratio of the policy behavior output probabilities η k,i (θ) at the k-th discrete time instant of the i-th sampled trajectory is specifically expressed as:

[0145]

[0146] In the formula, is the probability of the current sampled trajectory output by the policy network, is the probability of all sampled trajectories output by the policy network.

[0147] In step 4, the method of calculating the update gradient by the advantage function and updating the parameters of the rocket load reduction guidance strategy includes:

[0148] The value network parameter update gradient calculated by the value network advantage function is:

[0149]

[0150] According to the value network parameter update gradient with the update step perform gradient descent update on the value network parameters in the rocket load reduction guidance strategy;

[0151] The policy network parameter update gradient calculated by the policy network advantage function is:

[0152]

[0153] According to the policy network parameter update gradient with the update step α θ perform gradient ascent update on the policy network parameters in the rocket load reduction guidance strategy;

[0154] In the formula, is the value function network parameter update gradient; is the gradient of the value function network value with respect to its parameters, is the policy network parameter update gradient.

[0155] In order to verify the technical solution of the present invention, a specific application example is provided for illustration:

[0156] Considering the statistical wind field data in a certain place, using the nominal trajectory model of the dynamics of the ascending section of a certain launch vehicle, and using this method to conduct numerical simulation experiments to verify the reduction effect of this method on the maximum aerodynamic load on the rocket body.

[0157] Take Figure 2 the seasonal statistical average wind field data at 23.48° east longitude and 60.80° north latitude as the input data, and use the Figures 3 to 7 shown nominal trajectory model of the launch vehicle. The comparison of the maximum aerodynamic load on the rocket body obtained by using this method and not using this method (i.e., directly tracking the nominal trajectory model) under the seasonal wind field is shown in Table 1:

[0158] Table 1

[0159] wind farm Maximum aerodynamic load (Pa) using the present invention Maximum aerodynamic load (Pa) without using the present invention Load reduction ratio Spring 478 1196 60.03% Summer 524 1083 51.62% Autumn 512 1015 49.56% Winter 469 1182 60.32%

[0160] It can be seen from Table 1 that the technical solution disclosed by the present invention has a good load reduction effect.

[0161] The beneficial effects of the embodiments of the present invention are as follows:

[0162] In view of the load reduction requirements of a launch vehicle with a high slenderness ratio during the ascending flight in the atmosphere, the present invention utilizes high-resolution hourly forecast data of upper-level winds. By constructing a Markov decision process model for ascending guidance, simulating the time-varying uncertainty of the wind field, and optimizing the update gradient of the rocket load reduction guidance strategy, a training framework for the rocket load reduction wind correction guidance strategy that adapts to the deviation between the forecast wind field and the actual wind field is established. Through reinforcement learning, the attitude angle command of the rocket can be corrected in real time during the rocket flight, adapting to the uncertainty deviation between the actual wind field and the forecast wind field, realizing the real-time correction of the rocket attitude command in the case of an uncertain deviation between the forecast wind field and the actual wind field, thereby reducing the maximum aerodynamic load on the rocket body, improving the rocket reliability, and being applicable to the load reduction guidance control of a launch vehicle considering the time-varying characteristics of the wind field.

[0163] The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A design method of an enhanced learning rocket load reduction guidance law considering the time-varying characteristics of the wind field, characterized in that, The method includes: Step 1: Obtain the target data required for the rocket load-reducing guidance law, and generate high-resolution hourly upper-air wind forecast data based on the WRF model; Step 2: Through the centroid motion dynamics equation of the rocket during the ascent stage, take a complete flight process of the rocket in the atmosphere during the ascent stage as one episode of the ascent-stage guidance Markov decision process, and construct an ascent-stage guidance Markov decision process model; Step 3: Record the total number of wind fields generated for the high-resolution hourly upper-air wind forecast data. Based on the high-resolution hourly upper-air wind forecast data, at the beginning of each episode of the ascent-stage guidance Markov decision process, randomly generate a wind field number within the total number of wind fields, select the wind field function in the ascent-stage guidance Markov decision process model corresponding to this number as the forecast wind field for this episode, and calculate the difference between the forecast wind field and the actual wind field to obtain the uncertainty deviation; Step 4: Introduce a value function network in the policy network to represent the rocket load-reducing guidance strategy; interact with the ascent-stage guidance Markov decision process model through the rocket load-reducing guidance strategy to complete the sampling of a complete episode of the ascent-stage guidance Markov decision process and obtain sampling data; estimate the advantage function based on the sampling data, calculate the update gradient from the advantage function, update the parameters of the rocket load-reducing guidance strategy with the update gradient, and optimize the uncertainty deviation based on the rocket load-reducing guidance strategy with updated parameters to obtain a rocket load-reducing guidance law that minimizes the expected value of the maximum aerodynamic load throughout the ascent stage.

2. The method according to claim 1, characterized in that, In Step 1, the method for generating the high-resolution hourly upper-air wind forecast data includes: Obtain the target data required for the rocket load-reducing guidance law, where the target data is global forecast model data, terrain data, and observation data for the same time period; Use dynamic downscaling technology to process the global forecast model data and terrain data received by the WRF model to obtain the boundary field and initial field required for driving the WRF model; Combined with the three-dimensional variational data assimilation method 3DVAR, assimilate the multi-source heterogeneous data of real-time upper-air wind sounding data, wind profiler radar data, and microwave radiometer data in the observation data to update the initial field; Start the WRF model for forward integration using the updated initial field to obtain the hourly upper-air wind forecast data for the next 72h at the current moment; use the model result correction method combined with the observation data to correct the hourly upper-air wind forecast data for the next 72h at the current moment to generate high-resolution hourly upper-air wind forecast data.

3. The method according to claim 2, wherein In Step 1, the cost function used in the three-dimensional variational data assimilation method 3DVAR is: B = (x b - x true )(x b - x true ) T ; In the formula, J(x) is the solution of the cost function, where the smallest solution is the analysis field, that is, the updated initial field; x represents the analysis variable at the model grid points, x b represents the background field, B represents the background field error covariance matrix, y0 represents the observed data, Hx represents the observation operator, R represents the observation error covariance matrix, and the superscripts T and -1 represent the transpose and inverse of the matrix respectively, x true represents the true field.

4. The method according to claim 3, characterized in that, In Step 1, the model result correction method is: In the formula, is the corrected field at time t1 + n, where t1 is the observation data time, and n = 1, 2, 3,... is the difference between the correction time and the observation data time; b0, b1, and b2 are the first model coefficient, the second model coefficient, and the third model coefficient automatically established by the least squares method; is the WRF model forecast result at time t1 + n; is the error between the observation data at time t1 and the WRF model forecast result.

5. The method according to claim 4, wherein In Step 2, through the centroid motion dynamics equation of the rocket during the ascent stage, take a complete flight process of the rocket in the atmosphere during the ascent stage as one episode of the ascent-stage guidance Markov decision process, and the method for constructing the ascent-stage guidance Markov decision process model is: Construct the dynamic equation of the rocket's center-of-mass motion during the ascending stage based on the rocket's current flight time, state variables, control variables, and wind field function; among them, the constructed dynamic equation of the rocket's center-of-mass motion during the ascending stage is as follows: In the formula, is the time derivative of the state variables of the rocket's center-of-mass motion during the ascending stage; t is the rocket's current flight time; x rocket is the state variable of the rocket's center-of-mass motion during the ascending stage; u is the control variable of the rocket's center-of-mass motion during the ascending stage; V wind (h) is the wind field function that varies with altitude; Select the state variable x in the dynamic equation of the centroid motion during the rocket ascent stage rocket as the state variable s of the Markov decision process model for ascent guidance k ; The control variable \(u\) in the dynamic equation of the rocket's center-of-mass motion during the ascent phase is selected as the action variable \(a\) of the ascent-phase guidance Markov decision process model k ; Determine the discrete time step dT of the Markov decision process for the ascent guidance. Under the assumption that the wind field is known, discretize the continuous dynamic equations of the centroid motion of the launch vehicle during the ascent stage according to the determined discrete time step. Then, the state transition probability p(s k+1 |s k ,a k ) of the Markov decision process for the ascent guidance is the conditional probability p(V wind (h)) with respect to the wind field function V wind (h); Take the maximum aerodynamic load Q k |α Tk | as the reward function r(s k ,a k ); Then, a complete flight process of the rocket in the atmosphere during the ascent stage is one episode of the ascent-stage guidance Markov decision process, and an ascent-stage guidance Markov decision process model is obtained; among them, the ascent-stage guidance Markov decision process model is: s k = x rocket | t=kdT ; a k = u| t=kdT ; where k is the k-th discrete time of the Markov decision process, and s k+1 is the state quantity at the (k + 1)-th discrete time, and k + 1 represents the (k + 1)-th discrete time of the Markov decision process; N is the total length of the episode of the Markov decision process for ascent phase guidance; Q k |α Tk | represents the maximum aerodynamic load in the time interval from kdT to (k + 1)dT.

6. The method according to claim 5, wherein In step 3, record the total number of wind fields generated for the high-resolution hourly upper-air wind forecast data. Based on the high-resolution hourly upper-air wind forecast data, at the start of each episode of the ascent-phase guidance Markov decision process, randomly generate a wind field number within the total number of wind fields, and select the wind field function in the ascent-phase guidance Markov decision process model corresponding to this number as the forecast wind field for this episode. The method for calculating the difference between the forecast wind field and the actual wind field to obtain the uncertainty deviation includes: S1. Record the total number of wind fields M generated for the high-resolution hourly upper-air wind forecast data; S2. At the start of each round of the ascent-phase guidance Markov decision process, randomly generate a wind field number j within the range [1, M]; in this round, use the wind field function V wind (h) in the ascent-phase guidance Markov decision process model as the j-th wind field in the high-resolution hourly upper-air wind forecast data; and use the j-th wind field as the forecast wind field for this round. S3. Repeat step S2 to obtain all forecast wind fields, calculate the differences between all forecast wind fields and the actual wind field to obtain the uncertainty deviation, and store it.

7. The method according to claim 6, wherein In step 4, introduce a value function network into the policy network to represent the rocket load reduction guidance strategy. By interacting the rocket load reduction guidance strategy with the ascent-phase guidance Markov decision process model, complete the sampling of a complete episode of the ascent-phase guidance Markov decision process. The method for obtaining the sampling data includes: Introduce a value function network into the policy network to represent the rocket load reduction guidance strategy, and interact the rocket load reduction guidance strategy with the ascent-phase guidance Markov decision process model to generate a set of sampling trajectories; Record the output probability of the policy network and the output data of the value function network corresponding to each sampling trajectory in the set of sampling trajectories; at the same time, calculate the temporal difference residual corresponding to each sampling trajectory, and obtain the generalized advantage estimate of the temporal difference residual through recursive calculation; sum the generalized advantage estimate and the output data of the value function network to obtain the target value to be fitted by the value function network, and complete the sampling of a complete episode of the ascent-phase guidance Markov decision process; The sampling data consists of the output data of the value function network, the target value to be fitted by the value function network, the generalized advantage estimate, and the output probability of the policy network.

8. The method according to claim 7, wherein In step 4, the sampling trajectory set contains N step = N traj × N sampling trajectories where N step is the total number of state transition steps, N traj is the number of sampling trajectories of a complete round, N is the total length of a round of the ascending section guidance Markov decision process; i is the sampling trajectory serial number; Temporal difference residual is as follows: where r k,i is the reward at the k-th discrete time of the i-th sampling trajectory, γ is the discount factor of the cumulative reward, is the output data of the value function network, is the output data of the value function network at the (k + 1)-th discrete time; General Advantage Estimation is as follows: where \(l\) is the summation index of the general advantage estimation, \(N\) is the total length of the episodes of the upward segment guidance Markov decision process, \(k\) is the \(k\) - th discrete time instant of the Markov decision process, \(\lambda\) is the decay coefficient of the general advantage estimation, is the temporal - difference residual at the \((k + l)\) - th discrete time instant; Target value to be fitted by value function network is:

9. The method according to claim 8, characterized in that, In step 4, the advantage functions estimated based on the sampling data include: the value function network advantage function and the policy network advantage function; among them, The value function network advantage function is: The policy network advantage function is: Wherein, is the value of the advantage function of the value function network, is the set of sampled trajectories, k is the k-th discrete time of the Markov decision process, and i is the serial number of the sampled trajectory, is the output data of the value function network, is the target value to be fitted by the value function network; is the advantage function value of the policy network, γ k is the discount factor of the cumulative reward at the k-th discrete time, η k,i (θ) is the policy behavior output probability ratio at the k-th discrete time of the i-th sampled trajectory, is the Generalized Advantage Estimation, clip is the clipping function, ∈ is the trust region radius for the policy network parameter gradient update; The policy behavior output probability ratio η at the k-th discrete moment of the i-th sampling trajectory k,i (θ) is specifically expressed as: Among them, is the probability of the currently sampled trajectory output by the policy network, is the probability of all sampled trajectory sets output by the policy network.

10. The method according to claim 9, characterized in that, In step 4, the update gradients calculated from the advantage functions include: the value function network parameter update gradient and the policy network parameter update gradient, where, The value function network parameter update gradient is: Update the gradient according to the value function network parameters With the update step size Perform gradient descent update on the value function network parameters in the rocket load reduction guidance strategy; The policy network parameter update gradient is: Update the gradient according to the policy network parameters With an update step size α θ Perform gradient ascent update on the policy network parameters in the rocket load reduction guidance strategy; In the formula, is the update gradient of the value function network parameters; is the gradient of the value function network value with respect to its parameters, is the update gradient of the policy network parameters.

Citation Information

Patent Citations

  • Aircraft deloading guidance method and device, and storage medium

    CN117872731A

  • Energy-saving optimization method and apparatus for terminal air conditioning system of integrated data center cabinet

    WO2023116742A1