A reinforcement learning-based rocket load reduction guidance law design method considering time-varying wind field characteristics

By using reinforcement learning and Markov decision process models, the rocket attitude angle command is corrected in real time, which solves the problem of aerodynamic load deviation caused by the time-varying characteristics of wind field and improves the reliability and carrying capacity of the launch vehicle.

CN120372803BActive Publication Date: 2025-11-14CHINESE PEOPLES LIBERATION ARMY UNIT 63810
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510422153.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-05
Publication Date
2025-11-14
Estimated Expiration
2045-04-05

AI Technical Summary

Technical Problem

Existing load reduction guidance law design methods cannot adapt to the time-varying characteristics of actual and predicted wind fields, resulting in large deviations in the aerodynamic loads experienced by the launch vehicle during the ascent phase in the atmosphere, which affects the reliability and carrying capacity of the rocket.

Method used

By employing reinforcement learning, and using high-resolution hourly high-altitude wind forecast data and a Markov decision process model for the ascent phase guidance, the rocket attitude angle command is corrected in real time to adapt to the time-varying characteristics of the wind field, optimize the rocket load reduction guidance strategy, and reduce the maximum aerodynamic load on the rocket body.

Benefits of technology

Under the time-varying characteristics of wind fields, the rocket attitude angle command is corrected in real time, reducing the maximum aerodynamic load on the rocket body and improving the reliability and carrying capacity of the rocket.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372803B_ABST
    Figure CN120372803B_ABST
Patent Text Reader

Abstract

This invention discloses a reinforcement learning-based rocket load reduction guidance law design method considering the time-varying characteristics of wind fields. It solves the problem that existing rocket load reduction guidance law designs cannot adapt to the deviation between actual and predicted wind fields, belonging to the field of rocket load reduction guidance and control. The method includes: generating high-resolution hourly upper-level wind forecast data; constructing a Markov decision process model for the ascent phase guidance; using the high-resolution hourly upper-level wind forecast data to simulate the uncertainty deviation between the predicted and actual wind fields through the model; completing the interactive sampling of the entire round of the ascent phase guidance Markov decision process using a rocket load reduction guidance strategy; calculating the update gradient based on the sampled data; updating the parameters of the rocket load reduction guidance strategy using the update gradient; optimizing the uncertainty deviation; and obtaining a rocket load reduction guidance law that minimizes the expected maximum aerodynamic load throughout the ascent phase. This invention enables real-time correction of rocket attitude commands when there is a deviation between the predicted and actual wind fields.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of launch vehicle load reduction guidance and control technology, and relates to a reinforcement learning-based rocket load reduction guidance law design method that considers the time-varying characteristics of wind fields. Background Technology

[0002] For heavy-lift launch vehicles with a high aspect ratio, the internal bending moments caused by aerodynamic loads during the ascent phase within the atmosphere can easily lead to structural instability or even failure, reducing the rocket's reliability. Furthermore, the structural design of launch vehicles is generally constrained by aerodynamic loads, resulting in conservative designs that increase structural mass and reduce payload capacity. As launch vehicle launches evolve towards lower cost, higher frequency, and higher reliability, it is necessary to specifically design the load reduction guidance law for the ascent phase—that is, the generation law of the desired rocket attitude calculated in real-time from rocket state variables—to reduce the aerodynamic loads on the rocket.

[0003] Existing load reduction guidance law designs rely on accurate wind field forecast data: specifically, based on the wind field at the launch time generated by the forecast, a trajectory optimization method is used to regenerate the rocket attitude angle command that satisfies the maximum aerodynamic load constraint under that wind field.

[0004] However, actual wind fields have typical time-varying characteristics, meaning that wind speed and direction at various altitudes change uncertainly over time. This makes the existing load reduction guidance law design method, which uses a fixed forecast wind field to regenerate rocket attitude angle commands, unable to adapt to the deviation between the actual and forecast wind fields, resulting in limited load reduction effects. Summary of the Invention

[0005] To address the aforementioned technical problems, this invention provides a reinforcement learning-based rocket load reduction guidance law design method that considers the time-varying characteristics of wind fields. This method is applicable to the load reduction guidance and control of launch vehicles that consider the time-varying characteristics of wind fields. Through reinforcement learning, the rocket attitude angle command can be corrected in real time during rocket flight to adapt to the uncertainty deviation between the actual wind field and the predicted wind field, thereby reducing the maximum aerodynamic load on the rocket body and improving rocket reliability.

[0006] This invention establishes a training framework for rocket load reduction wind correction guidance strategy that adapts to the deviation between the predicted and actual wind fields by generating high-resolution hourly high-altitude wind forecast data, constructing a Markov decision process model for the ascent phase guidance, simulating time-varying uncertainties in the wind field, and optimizing the gradient update of the rocket load reduction guidance strategy. This framework enables real-time correction of rocket attitude commands when there is a deviation between the predicted and actual wind fields, thereby reducing the maximum aerodynamic load on the rocket body. This method is applicable to the load reduction guidance control of launch vehicles that consider the time-varying characteristics of the wind field.

[0007] The objective of this invention is specifically achieved through the following technical solutions:

[0008] This invention discloses a reinforcement learning-based rocket load reduction guidance law design method considering the time-varying characteristics of wind fields. The method includes:

[0009] Step 1: Obtain the target data required for the rocket load reduction guidance law, and generate high-resolution hourly upper-altitude wind forecast data based on the WRF model;

[0010] Step 2: Using the dynamic equations of the rocket's center of mass motion during the ascent phase, a complete ascent phase flight process within the rocket's atmosphere is taken as one round of the ascent phase guidance Markov decision process, and an ascent phase guidance Markov decision process model is constructed.

[0011] Step 3: Record the total number of wind fields generated by generating high-resolution upper-level wind hourly forecast data. Based on the high-resolution upper-level wind hourly forecast data, at the beginning of each round of the ascending-stage guided Markov decision process, randomly generate a wind field number within the total number of wind fields. Select the wind field function in the ascending-stage guided Markov decision process model corresponding to the number as the forecast wind field for that round. Calculate the difference between the forecast wind field and the actual wind field to obtain the uncertainty bias.

[0012] Step four: Introduce a value function network to represent the rocket load reduction guidance strategy in the strategy network; through the interaction between the rocket load reduction guidance strategy and the ascent phase guidance Markov decision process model, complete the sampling of the entire round of the ascent phase guidance Markov decision process to obtain sampled data; estimate the dominance function based on the sampled data, calculate the update gradient from the dominance function, update the parameters of the rocket load reduction guidance strategy from the update gradient, optimize the uncertainty bias based on the updated parameters of the rocket load reduction guidance strategy, and obtain the rocket load reduction guidance law that minimizes the expected value of the maximum aerodynamic load throughout the ascent phase.

[0013] In step one, the method for generating high-resolution hourly upper-level wind forecast data includes:

[0014] The target data required for the rocket load reduction guidance law are global forecast model data, terrain data, and observation data for the same time period.

[0015] Dynamic downscaling techniques were used to process global forecast model data and topographic data received by the WRF model to obtain the boundary field and initial field required for driving the WRF model.

[0016] By combining the three-dimensional variational data assimilation method 3DVAR, the real-time upper-air wind detection data, wind profiler radar data and microwave radiometer data in the observation data are assimilated into multi-source heterogeneous data to update the initial field;

[0017] The WRF model is initiated with the updated initial field to perform forward integration, obtaining hourly upper-level wind forecasts for the next 72 hours from the current moment. The hourly upper-level wind forecasts for the next 72 hours from the current moment are then corrected using model result correction methods combined with observational data to generate high-resolution hourly upper-level wind forecasts.

[0018] In step one, the cost function used by the three-dimensional variational data assimilation method 3DVAR is:

[0019]

[0020] B = (x b -x true (x) b -x true ) T ;

[0021] In the formula, J(x) is the solution to the cost function, where the minimum solution is the analysis field, i.e., the updated initial field; x represents the analysis variable at the model grid point. b Let B represent the background field, y0 represent the observed data, Hx represent the observation operator, R represent the observation error covariance matrix, and the superscripts T and -1 represent the transpose and inverse of the matrix, respectively. true Represents the real field.

[0022] In step one, the method for correcting the pattern results is as follows:

[0023]

[0024] In the formula, It is the correction field at time t1+n, where t1 is the time of the observed data, n = 1, 2, 3, ... is the difference between the correction time and the time of the observed data; b0, b1, b2 are the first model coefficients, the second model coefficients, and the third model coefficients automatically established using the least squares method; It is the WRF model prediction result at time t1+n; It is the error between the observation data at time t1 and the prediction results of the WRF model.

[0025] In step two, the ascent phase guidance Markov decision process model is constructed by using the rocket's center of mass motion dynamics equations, treating a complete ascent phase flight process within the rocket's atmosphere as one round of the ascent phase guidance Markov decision process.

[0026] The dynamic equations of motion of the rocket's center of mass during the ascent phase are constructed using the rocket's current flight time, state variables, control variables, and wind field function. The constructed dynamic equations of motion of the rocket's center of mass during the ascent phase are as follows: In the formula, t is the time derivative of the motion state of the rocket's center of mass during the ascent phase; t is the current flight time of the rocket; x rocket V is the motion state of the rocket's center of mass during the ascent phase; u is the control quantity of the rocket's center of mass during the ascent phase; V wind (h) is the wind field function that varies with altitude;

[0027] The state variable x in the equation of motion of the center of mass during the rocket's ascent phase rocket The state variable s is selected as the state variable in the ascending-phase guided Markov decision process model. k ;

[0028] The control variable u in the equation of motion of the center of mass during the rocket's ascent phase is selected as the behavioral variable a in the Markov decision process model for ascent phase guidance. k ;

[0029] Given a known wind field and a determined discrete step length dT for the Markov decision process during ascent guidance, the continuous dynamic equations of motion of the launch vehicle's center of mass during ascent are discretized according to the determined discrete step length. Then, the state transition probability p(s) of the Markov decision process during ascent guidance is obtained. k+1 |s k ,a k That is, with respect to the wind field function V wind The conditional probability p(V) of (h) wind (h));

[0030] The maximum aerodynamic load Q experienced by the rocket during its ascent phase. k |α Tk |As the reward function r(s) k ,a k );

[0031] Therefore, a complete ascent phase flight process within the rocket's atmosphere constitutes one round of the ascent phase guidance Markov decision process, resulting in the ascent phase guidance Markov decision process model; where the ascent phase guidance Markov decision process model is:

[0032] s k =x rocket | t=kdT ;

[0033] a k =u| t=kdT ;

[0034]

[0035] In the formula, k is the k-th discrete time of the Markov decision process, and s k+1 Q is the state variable at the (k+1)th discrete time step, where k+1 represents the (k+1)th discrete time step of the Markov decision process; N is the total length of the rounds in the rising-segment guided Markov decision process;k |α Tk | represents the maximum aerodynamic load within the time interval from kdT to (k+1)dT.

[0036] In step three, the total number of wind fields generated by the high-resolution hourly upper-level wind forecast data is recorded. Based on the high-resolution hourly upper-level wind forecast data, at the beginning of each round of the ascending-phase guided Markov decision process, a wind field number is randomly generated within the total number of wind fields. The wind field function in the ascending-phase guided Markov decision process model corresponding to this number is selected as the forecast wind field for that round. The method for calculating the difference between the forecast wind field and the actual wind field to obtain the uncertainty bias includes:

[0037] S1, records the total number of wind fields M generated by producing high-resolution hourly upper-level wind forecast data;

[0038] S2, at the beginning of each round of the ascending-stage guided Markov decision process, a wind field number j is randomly generated within the range [1,M]; in that round, the wind field function V in the ascending-stage guided Markov decision process model is... wind (h) is the j-th wind field in the high-resolution hourly upper-level wind forecast data; and the j-th wind field is used as the forecast wind field for this round;

[0039] S3. Repeat step S2 to obtain all forecast wind fields, calculate the difference between all forecast wind fields and actual wind fields to obtain the uncertainty deviation, and store it.

[0040] In step four, a value function network is introduced into the policy network to represent the rocket load reduction guidance strategy. The rocket load reduction guidance strategy interacts with the ascent phase guidance Markov decision process model to complete the sampling of a full round of the ascent phase guidance Markov decision process. The methods for obtaining the sampled data include:

[0041] A value function network is introduced into the policy network to represent the rocket load reduction guidance strategy. The set of sampled trajectories is generated by the interaction between the rocket load reduction guidance strategy and the Markov decision process model of the ascent phase guidance.

[0042] Record the output probability of the policy network and the output data of the value function network for each sampling trajectory in the sampling trajectory set; at the same time, calculate the temporal difference residual corresponding to each sampling trajectory, and obtain the general advantage estimate of the temporal difference residual through recursive calculation; sum the general advantage estimate and the output data of the value function network to obtain the target value to be fitted by the value function network, and complete the sampling of a complete round of the rising phase guidance Markov decision process;

[0043] The sampled data consists of the output data of the value function network, the target value to be fitted by the value function network, the general advantage estimate, and the output probability of the policy network.

[0044] In step four, the set of sampling trajectories Contains N step =N traj ×N sampling trajectories In the formula, N step N represents the total number of state transition steps. traj N represents the number of complete round-based sampling trajectories, N is the total round length of the rising-segment guided Markov decision process, and i is the sampling trajectory number.

[0045] Temporal Differential Residual for:

[0046]

[0047] In the formula, r k,i Let γ be the reward at the k-th discrete moment of the i-th sampling trajectory, and γ be the discount factor for the cumulative reward. Output data for the value function network. Output data for the value function network at the (k+1)th discrete time step;

[0048] General advantage estimation for:

[0049]

[0050] In the formula, l is the general advantage estimate summation index, N is the total length of the rounds in the rising-segment guided Markov decision process, k is the k-th discrete time of the Markov decision process, and λ is the general advantage estimate decay coefficient. The time-series difference residual is the (k+l)th discrete time step.

[0051] The target value to be fitted by the value function network for:

[0052]

[0053] In step four, the advantage functions estimated based on the sampled data include: the value function network advantage function and the policy network advantage function; among which,

[0054] The value function network advantage function is:

[0055]

[0056] The policy network advantage function is:

[0057]

[0058] In the formula, The value function network advantage function value, Let k be the set of sampled trajectories, k be the k-th discrete time step of the Markov decision process, and i be the sampled trajectory index. Output data for the value function network. The target value to be fitted by the value function network;

[0059] γ is the advantage function value of the policy network. k η is the discount factor for the cumulative reward at the k-th discrete time. k,i (θ) represents the probability ratio of the policy behavior output at the k-th discrete moment of the i-th sampled trajectory. For general advantage estimation, clip is the clipping function, and ∈ is the trust region radius for gradient updates of policy network parameters;

[0060] The probability ratio of the policy behavior output at the k-th discrete moment of the i-th sampling trajectory to η k,i (θ) is specifically represented as:

[0061]

[0062] in, The probability of the current sampled trajectory is output by the policy network. represents all probabilities of the set of sampled trajectories output by the policy network.

[0063] In step four, the update gradients calculated by the advantage function include: the value function network parameter update gradient and the policy network parameter update gradient, where,

[0064] The gradient for updating the network parameters of the value function is:

[0065]

[0066] Update gradients based on value function network parameters To update step size Gradient descent is used to update the network parameters of the value function in the rocket load reduction guidance strategy.

[0067] The gradient for updating the policy network parameters is:

[0068]

[0069] Update gradients based on policy network parameters To update step size α θ Gradient ascent update is performed on the strategy network parameters in the rocket load reduction guidance strategy.

[0070] In the formula, Update the gradient for the network parameters of the value function; Let the value function network take the gradient with respect to its parameters. Update gradients for policy network parameters.

[0071] The beneficial effects of this invention are:

[0072] This invention addresses the load reduction requirements of high-length-to-slender-ratio launch vehicles during their ascent phase within the atmosphere. Utilizing high-resolution hourly high-altitude wind forecast data, it constructs a Markov decision process model for ascent phase guidance, simulates time-varying uncertainties in the wind field, and optimizes the rocket load reduction guidance strategy update gradient. This establishes a training framework for a rocket load reduction wind correction guidance strategy that adapts to the deviation between the forecast and actual wind fields. Through reinforcement learning, it can correct rocket attitude angle commands in real time during flight, adapting to the uncertainty deviation between the actual and forecast wind fields. This enables real-time correction of rocket attitude commands even when there are uncertainties between the forecast and actual wind fields, thereby reducing the maximum aerodynamic load on the rocket body and improving rocket reliability. This invention is suitable for load reduction guidance and control of launch vehicles considering time-varying wind field characteristics. Attached Figure Description

[0073] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments.

[0074] Figure 1 This is a schematic diagram of the system operation process based on the WRF model.

[0075] Figure 2 This is a schematic diagram of statistical average wind field data.

[0076] Figure 3 This is a schematic diagram of the nominal trajectory range height profile.

[0077] Figure 4 This is a schematic diagram of the vertical velocity profile of the nominal trajectory.

[0078] Figure 5 This is a schematic diagram of the horizontal velocity profile of the nominal trajectory.

[0079] Figure 6 This is a schematic diagram of the standard ballistic angle of attack profile.

[0080] Figure 7 This is a schematic diagram of the standard ballistic inclination profile. Detailed Implementation

[0081] The reinforcement learning method used in the reinforcement learning-based rocket load reduction guidance law design method considering the time-varying characteristics of wind fields provided in this invention embodiment is an effective framework for training optimal policies under stochastic environments. Given a mission-related Markov decision process model and sources of environmental uncertainty, the method trains a policy to maximize the expected cumulative reward obtained under uncertain environments. This invention trains a rocket load reduction guidance strategy within a reinforcement learning framework, and the method includes:

[0082] Step 1: Obtain the target data required for the rocket load reduction guidance law, and generate high-resolution hourly upper-altitude wind forecast data based on the WRF model;

[0083] Step 2: Using the dynamic equations of the rocket's center of mass motion during the ascent phase, a complete ascent phase flight process within the rocket's atmosphere is taken as one round of the ascent phase guidance Markov decision process, and an ascent phase guidance Markov decision process model is constructed.

[0084] Step 3: Record the total number of wind fields generated by generating high-resolution upper-level wind hourly forecast data. Based on the high-resolution upper-level wind hourly forecast data, at the beginning of each round of the ascending-stage guided Markov decision process, randomly generate a wind field number within the total number of wind fields. Select the wind field function in the ascending-stage guided Markov decision process model corresponding to the number as the forecast wind field for that round. Calculate the difference between the forecast wind field and the actual wind field to obtain the uncertainty bias.

[0085] Step four involves introducing a Value Function Network (VFNetwork) into the Policy Network to represent the rocket load reduction guidance strategy. This VF guidance strategy interacts with the ascent phase guidance Markov decision process model to complete sampling of the entire ascent phase Markov decision process, obtaining sampled data. Based on the sampled data, the dominance function is estimated, and the updated gradient is calculated from the dominance function. The updated gradient is then used to update the parameters of the rocket load reduction guidance strategy. Based on the updated parameters, the uncertainty bias of the rocket load reduction guidance strategy is optimized to obtain the rocket load reduction guidance law that minimizes the expected value of the maximum aerodynamic load throughout the ascent phase.

[0086] The policy network guides the agent's decision-making, providing the probability of choosing an action in a given state. The value function network evaluates the long-term value of a state or action, helping the agent learn and optimize its policy. The policy network is a neural network used to model the agent's policy, i.e., the probability distribution of choosing an action in a given state; the value function network is a neural network used to estimate the expected cumulative reward that can be obtained after taking a certain action in a given state.

[0087] In step one, the method for generating high-resolution hourly upper-level wind forecast data includes:

[0088] The target data required for the rocket load reduction guidance law are global forecast model data, terrain data, and observation data for the same time period.

[0089] Dynamic downscaling techniques were used to process global forecast model data and topographic data received by the WRF model to obtain the boundary field and initial field required for driving the WRF model.

[0090] By combining the three-dimensional variational data assimilation method 3DVAR, the real-time upper-air wind detection data, wind profiler radar data and microwave radiometer data in the observation data are assimilated into multi-source heterogeneous data to update the initial field;

[0091] The WRF model is initiated with the updated initial field to perform forward integration, obtaining hourly upper-level wind forecasts for the next 72 hours from the current moment. The hourly upper-level wind forecasts for the next 72 hours from the current moment are then corrected using model result correction methods combined with observational data to generate high-resolution hourly upper-level wind forecasts.

[0092] Global forecast model data and topographic data are obtained from the corresponding data sources, such as GFS, ECMWF, or GRAPES. GFS data can usually be downloaded from the U.S. National Center for Environmental Prediction (NECP), ECMWF data needs to be obtained through ECMWF member countries or by purchase, and GRAPES data needs to be obtained from relevant research institutions of the China Meteorological Administration.

[0093] like Figure 1 As shown, a WRF-based system for a specific region can obtain hourly upper-level wind forecast data with a height distribution (layer thickness of 250m) within that region. The WRF-based system uses the WRF (Weather Research Forecasting) model as its foundation, employing dynamic downscaling techniques to process mainstream international / domestic numerical weather prediction results (GFS / ECMWF / GRAPES) to obtain the boundary and initial fields required for model driving. Then, it combines three-dimensional variational data assimilation methods to perform multi-source heterogeneous data assimilation from real-time upper-level wind sounding data, wind profiler radar data, and microwave radiometer data, updating the model's initial field. The WRF model is then forward-integrated to obtain hourly upper-level wind forecast data for the next 72 hours. Model result correction methods, combined with real-time upper-level wind data, are used to correct the hourly upper-level wind forecast data, ultimately outputting high-resolution hourly upper-level wind forecast data.

[0094] For example: (1) Preparation of pattern data:

[0095] After downloading the target data, perform target data format conversion. WRF mode uses a specific data format. The ungrib.exe tool in the WRF and Processing System (WPS) is used to extract GRIB format data and convert it into an intermediate file recognizable by WRF. Then, configure WPS to process the data provided to WRF. Edit the `namelist.wps` file to specify the time range, spatial range, resolution, vertical layers, and other information. Run WPS, where `. / geogrid.exe`: creates the model mesh according to the configuration defined in `namelist.wps`; `. / ungrib.exe`: decodes the GRIB data into an intermediate file; and `. / metgrid.exe`: merges intermediate files from different time periods and interpolates them onto the model mesh. Configure WRF mode by editing the `namelist.input` file, including the simulation time period, physical parameter options, dynamic frame options, etc. Ensure the WRF model can read the `met_em*` files output from WPS. Run the WRF mode and use `real.exe` to process the `met_em*` files output by WPS to generate the initial field and boundary conditions.

[0096] (2) Data assimilation:

[0097] To prepare upper-air wind sounding data, wind profiler radar data, and microwave radiometer data for the same time period, the observation data were first subjected to quality control and preprocessing. Erroneous observation data were removed using historical extreme values, spatial consistency, and basic physical laws. Then, obsproc.exe in WRF mode was used to write the observation data into ASCII format recognized by the assimilation system.

[0098] WRFDA is used to expand three-dimensional variational assimilation. 3DVAR mainly seeks the minimum value of the objective function. It is achieved by solving the analysis field (the solution with the minimum value) so that the analysis field achieves the best fit with both the background field and the observation field. At this point, the minimum solution of the cost function is also the updated initial field.

[0099] In step one, the cost function used by the three-dimensional variational data assimilation method 3DVAR is:

[0100]

[0101] B = (x b -x true (x) b -x true ) T ;

[0102] In the formula, J(x) is the solution to the cost function, where the minimum solution is the analysis field, i.e., the updated initial field; x represents the analysis variable at the model grid point.b Let B represent the background field, y0 represent the observed data, Hx represent the observation operator, R represent the observation error covariance matrix, and the superscripts T and -1 represent the transpose and inverse of the matrix, respectively. true Represents the real field.

[0103] (3) Simulation forecasts and corrections of forecast results:

[0104] After obtaining the updated initial field, run `wrf.exe` to start the simulation. Once the simulation is complete, use observational data, such as real-time upper-level wind detection data, to correct the model results, and finally output high-resolution hourly upper-level wind forecast data.

[0105] In step one, the method for correcting the pattern results is as follows:

[0106]

[0107] In the formula, It is the correction field at time t1+n, where t1 is the time of the observed data, n = 1, 2, 3, ... is the difference between the correction time and the time of the observed data; b0, b1, b2 are the first model coefficients, the second model coefficients, and the third model coefficients automatically established using the least squares method; It is the WRF model prediction result at time t1+n; It is the error between the observation data at time t1 and the prediction results of the WRF model.

[0108] In step two, the method for constructing the ascent phase guidance Markov decision process model is as follows: The model is constructed by using the dynamic equations of the rocket's center of mass motion during the ascent phase, treating a complete ascent phase flight process within the rocket's atmosphere as one round of the ascent phase guidance Markov decision process.

[0109] The dynamic equations of motion of the rocket's center of mass during the ascent phase are constructed using the rocket's current flight time, state variables, control variables, and wind field function. The constructed dynamic equations of motion of the rocket's center of mass during the ascent phase are as follows: In the formula, t is the time derivative of the motion state of the rocket's center of mass during the ascent phase; t is the current flight time of the rocket; x rocket V is the motion state of the rocket's center of mass during the ascent phase; u is the control quantity of the rocket's center of mass during the ascent phase; V wind (h) is the wind field function that varies with altitude;

[0110] The state variable x in the equation of motion of the center of mass during the rocket's ascent phase rocket The state variable s is selected as the state variable in the ascending-phase guided Markov decision process model. k ;

[0111] The control variable u in the equation of motion of the center of mass during the rocket's ascent phase is selected as the behavioral variable a in the Markov decision process model for ascent phase guidance. k ;

[0112] Given a known wind field and a determined discrete step length dT for the Markov decision process during ascent guidance, the continuous dynamic equations of motion of the launch vehicle's center of mass during ascent are discretized according to the determined discrete step length. Then, the state transition probability p(s) of the Markov decision process during ascent guidance is obtained. k+1 |s k ,a k That is, with respect to the wind field function V wind The conditional probability p(V) of (h) wind (h));

[0113] The reward function of the Markov decision process during the ascent phase should characterize the need to reduce aerodynamic loads. Therefore, the reward function is designed as the negative of the maximum aerodynamic load during the entire ascent phase. This allows us to define the maximum aerodynamic load Q experienced by the rocket during the ascent phase. k |α Tk |As the reward function r(s) k ,a k );

[0114] Therefore, a complete ascent phase flight process within the rocket's atmosphere constitutes one round of the ascent phase guidance Markov decision process, resulting in the ascent phase guidance Markov decision process model; where the ascent phase guidance Markov decision process model is:

[0115] s k =x rocket | t=kdT ;

[0116] a k =u| t=kdT ;

[0117]

[0118] In the formula, k is the k-th discrete time of the Markov decision process, and s k+1 Q is the state variable at the (k+1)th discrete time step, where k+1 represents the (k+1)th discrete time step of the Markov decision process; N is the total length of the rounds in the rising-segment guided Markov decision process; k |α Tk | represents the maximum aerodynamic load within the time interval from kdT to (k+1)dT.

[0119] In step three, the total number of wind fields generated by the high-resolution hourly upper-level wind forecast data is recorded. Based on the high-resolution hourly upper-level wind forecast data, at the beginning of each round of the ascending-phase guided Markov decision process, a wind field number is randomly generated within the total number of wind fields. The wind field function in the ascending-phase guided Markov decision process model corresponding to this number is selected as the forecast wind field for that round. The method for calculating the difference between the forecast wind field and the actual wind field to obtain the uncertainty bias includes:

[0120] S1, records the total number of wind fields M generated by producing high-resolution hourly upper-level wind forecast data;

[0121] S2, at the beginning of each round of the ascending-stage guided Markov decision process, a wind field number j is randomly generated within the range [1,M]; in that round, the wind field function V in the ascending-stage guided Markov decision process model is... wind (h) is the j-th wind field in the high-resolution hourly upper-level wind forecast data; and the j-th wind field is used as the forecast wind field for this round;

[0122] S3. Repeat step S2 to obtain all forecast wind fields, calculate the difference between all forecast wind fields and actual wind fields to obtain the uncertainty deviation, and store it.

[0123] In step four, a value function network is introduced into the policy network to represent the rocket load reduction guidance strategy. The rocket load reduction guidance strategy interacts with the ascent phase guidance Markov decision process model to complete the sampling of a full round of the ascent phase guidance Markov decision process. The methods for obtaining the sampled data include:

[0124] π in the policy network θ Introducing value networks Indicates rocket load reduction guidance strategy Where θ and θ v These are all parameters of the network to be trained, derived from the rocket load reduction guidance strategy. Interactive sampling with the ascending-phase guided Markov decision process model generates a total of N state transition steps. step Sampling trajectory set The total number of state transition steps N step Contains N traj A complete round trajectory, therefore there is Right now Contains N step =N traj The state transition data of ×N discrete sampling steps

[0125] Record the output probability of the policy network for each sampling trajectory in the set of sampling trajectories. and value network output data Simultaneously calculate each sampling trajectory Corresponding time-series differential residuals A general advantage estimate of the time-series difference residuals is obtained through recursive calculation. General advantage estimation and value network output Summing yields the target value to be fitted in the value network. Output from value network and the target value to be fitted in the value network General advantage estimation The output probability of the policy network The sampled data is composed of these components.

[0126] In step four, the set of sampling trajectories for:

[0127]

[0128] In the formula, the set of sampling trajectories Contains N step =N traj ×N sampling trajectories N step N represents the total number of state transition steps. traj N represents the number of complete round-based sampling trajectories, N is the total round length of the rising-segment guided Markov decision process, and i is the sampling trajectory number.

[0129] Temporal Differential Residual for:

[0130]

[0131] In the formula, r k,i Let γ be the reward at the k-th discrete moment of the i-th sampling trajectory, and γ be the discount factor for the cumulative reward. Output data to the value network. Output data for the value function network at the (k+1)th discrete time step;

[0132] General advantage estimation for:

[0133]

[0134] In the formula, l is the general advantage estimate summation index, N is the total length of the rounds in the rising-segment guided Markov decision process, k is the k-th discrete time of the Markov decision process, and λ is the general advantage estimate decay coefficient. The time-series difference residual is the (k+l)th discrete time step.

[0135] Value network target value to be fitted for:

[0136]

[0137] In step four, the methods for estimating the dominance function based on the sampled data include:

[0138] Based on the value network output in the sampled data and the target value to be fitted in the value network The value network advantage function obtained from the updated value network parameters is:

[0139]

[0140] Based on the general advantage estimation in the sampled data And sampling strategy network output probability The policy network advantage function that yields the updated policy network parameters is:

[0141]

[0142] In the formula, The value function network advantage function value, Let k be the set of sampled trajectories, k be the k-th discrete time step of the Markov decision process, and i be the sampled trajectory index. Output data for the value function network. The target value to be fitted by the value function network;

[0143] γ is the advantage function value of the policy network. k η is the discount factor for the cumulative reward at the k-th discrete time. k,i (θ) represents the probability ratio of the policy behavior output at the k-th discrete moment of the i-th sampled trajectory. For general advantage estimation, clip is the clipping function, and ∈ is the trust region radius for gradient updates of policy network parameters;

[0144] The probability ratio of the policy behavior output at the k-th discrete moment of the i-th sampling trajectory to η k,i (θ) is specifically represented as:

[0145]

[0146] In the formula, The probability of the current sampled trajectory is output by the policy network. represents all probabilities of the set of sampled trajectories output by the policy network.

[0147] In step four, the method for calculating the updated gradient from the dominance function and updating the parameters of the rocket load reduction guidance strategy from the updated gradient includes:

[0148] Gradient of value network parameter update calculated by the value network advantage function for:

[0149]

[0150] Update gradients based on value network parameters To update step size Gradient descent is used to update the value network parameters in the rocket load reduction guidance strategy.

[0151] Gradient of policy network parameter update calculated by the policy network advantage function for:

[0152]

[0153] Update gradients based on policy network parameters To update step size α θ Gradient ascent update is performed on the strategy network parameters in the rocket load reduction guidance strategy.

[0154] In the formula, Update the gradient for the network parameters of the value function; Let the value function network take the gradient with respect to its parameters. Update gradients for policy network parameters.

[0155] To verify the technical solution of the present invention, a specific application example is provided for illustration:

[0156] Considering the statistical wind field data of a certain location, and using the nominal trajectory model of the ascent phase of a certain launch vehicle, numerical simulation experiments were conducted using this method to verify the effect of this method on reducing the maximum aerodynamic load on the rocket body.

[0157] Will Figure 2 The seasonal statistical average wind field data at 23.48°E, 60.80°N was used as input data. Figures 3 to 7 Table 1 shows a comparison of the maximum aerodynamic loads on the launch vehicle body obtained using this method and without using this method (i.e., directly tracking the nominal trajectory model) under seasonal wind conditions, based on the nominal trajectory model shown.

[0158] Table 1

[0159] Wind field The maximum aerodynamic load (Pa) using the present invention The maximum aerodynamic load (Pa) of this invention was not employed. Load reduction ratio spring 478 1196 60.03% summer 524 1083 51.62% autumn 512 1015 49.56% winter 469 1182 60.32%

[0160] As can be seen from Table 1, the technical solution disclosed in this invention has a good load reduction effect.

[0161] The beneficial effects of the embodiments of the present invention are:

[0162] This invention addresses the load reduction requirements of high-length-to-slender-ratio launch vehicles during their ascent phase within the atmosphere. Utilizing high-resolution hourly high-altitude wind forecast data, it constructs a Markov decision process model for ascent phase guidance, simulates time-varying uncertainties in the wind field, and optimizes the rocket load reduction guidance strategy update gradient. This establishes a training framework for a rocket load reduction wind correction guidance strategy that adapts to the deviation between the forecast and actual wind fields. Through reinforcement learning, it can correct rocket attitude angle commands in real time during flight, adapting to the uncertainty deviation between the actual and forecast wind fields. This enables real-time correction of rocket attitude commands even when there are uncertainties between the forecast and actual wind fields, thereby reducing the maximum aerodynamic load on the rocket body and improving rocket reliability. This invention is suitable for load reduction guidance and control of launch vehicles considering time-varying wind field characteristics.

[0163] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A reinforcement learning-based rocket load reduction guidance law design method considering time-varying wind field characteristics, characterized in that, The method includes: Step 1: Obtain the target data required for the rocket load reduction guidance law, and generate high-resolution hourly upper-altitude wind forecast data based on the WRF model; Step 2: Using the dynamic equations of the rocket's center of mass motion during the ascent phase, a complete ascent phase flight process within the rocket's atmosphere is taken as one round of the ascent phase guidance Markov decision process, and an ascent phase guidance Markov decision process model is constructed. Step 3: Record the total number of wind fields generated by generating high-resolution upper-level wind hourly forecast data. Based on the high-resolution upper-level wind hourly forecast data, at the beginning of each round of the ascending-stage guided Markov decision process, randomly generate a wind field number within the total number of wind fields. Select the wind field function in the ascending-stage guided Markov decision process model corresponding to the number as the forecast wind field for that round. Calculate the difference between the forecast wind field and the actual wind field to obtain the uncertainty bias. Step four: Introduce a value function network to represent the rocket load reduction guidance strategy in the strategy network; through the interaction between the rocket load reduction guidance strategy and the ascent phase guidance Markov decision process model, complete the sampling of the entire round of the ascent phase guidance Markov decision process to obtain sampled data; estimate the dominance function based on the sampled data, calculate the update gradient from the dominance function, update the parameters of the rocket load reduction guidance strategy from the update gradient, optimize the uncertainty bias based on the updated parameters of the rocket load reduction guidance strategy, and obtain the rocket load reduction guidance law that minimizes the expected value of the maximum aerodynamic load throughout the ascent phase.

2. The method as described in claim 1, characterized in that, In step one, the method for generating high-resolution hourly upper-level wind forecast data includes: The target data required for the rocket load reduction guidance law are global forecast model data, terrain data, and observation data for the same time period. Dynamic downscaling techniques were used to process global forecast model data and topographic data received by the WRF model to obtain the boundary field and initial field required for driving the WRF model. By combining the three-dimensional variational data assimilation method 3DVAR, the real-time upper-air wind detection data, wind profiler radar data and microwave radiometer data in the observation data are assimilated into multi-source heterogeneous data to update the initial field; The WRF model is initiated with the updated initial field to perform forward integration, obtaining hourly upper-level wind forecasts for the next 72 hours from the current moment. The hourly upper-level wind forecasts for the next 72 hours from the current moment are then corrected using model result correction methods combined with observational data to generate high-resolution hourly upper-level wind forecasts.

3. The method as described in claim 2, characterized in that, In step one, the cost function used by the three-dimensional variational data assimilation method 3DVAR is: B=(x b -x true )(x b -x true ) T ; In the formula, J(x) is the solution to the cost function, where the minimum solution is the analysis field, i.e., the updated initial field; x represents the analysis variable at the model grid point. b Let B represent the background field, y0 represent the observed data, Hx represent the observation operator, R represent the observation error covariance matrix, and the superscripts T and -1 represent the transpose and inverse of the matrix, respectively. true Represents the real field.

4. The method as described in claim 3, characterized in that, In step one, the method for correcting the pattern results is as follows: In the formula, It is the correction field at time t1+n, where t1 is the time of the observed data, n = 1, 2, 3, ... is the difference between the correction time and the time of the observed data; b0, b1, b2 are the first model coefficients, the second model coefficients, and the third model coefficients automatically established using the least squares method; It is the WRF model prediction result at time t1+n; It is the error between the observation data at time t1 and the prediction results of the WRF model.

5. The method as described in claim 4, characterized in that, In step two, the ascent phase guidance Markov decision process model is constructed by using the rocket's center of mass motion dynamics equations, treating a complete ascent phase flight process within the rocket's atmosphere as one round of the ascent phase guidance Markov decision process. The dynamic equations of motion of the rocket's center of mass during the ascent phase are constructed using the rocket's current flight time, state variables, control variables, and wind field function. The constructed dynamic equations of motion of the rocket's center of mass during the ascent phase are as follows: In the formula, t is the time derivative of the motion state of the rocket's center of mass during the ascent phase; t is the current flight time of the rocket; x rocket V is the motion state of the rocket's center of mass during the ascent phase; u is the control quantity of the rocket's center of mass during the ascent phase; V wind (h) is the wind field function that varies with altitude; The state variable x in the equation of motion of the center of mass during the rocket's ascent phase rocket The state variable s is selected as the state variable in the ascending-phase guided Markov decision process model. k ; The control variable u in the equation of motion of the center of mass during the rocket's ascent phase is selected as the behavioral variable a in the Markov decision process model for ascent phase guidance. k ; Given a known wind field and a determined discrete step length dT for the Markov decision process during ascent guidance, the continuous dynamic equations of motion of the launch vehicle's center of mass during ascent are discretized according to the determined discrete step length. Then, the state transition probability p(s) of the Markov decision process during ascent guidance is obtained. k+1 |s k ,a k That is, with respect to the wind field function V wind The conditional probability p(V) of (h) wind (h)); The maximum aerodynamic load Q experienced by the rocket during its ascent phase. k |α Tk |As the reward function r(s) k ,a k ); Therefore, a complete ascent phase flight process within the rocket's atmosphere constitutes one round of the ascent phase guidance Markov decision process, resulting in the ascent phase guidance Markov decision process model; where the ascent phase guidance Markov decision process model is: s k =x rocket | t=kdT ; a k =u| t=kdT ; In the formula, k is the k-th discrete time of the Markov decision process, and s k+1 Q is the state variable at the (k+1)th discrete time step, where k+1 represents the (k+1)th discrete time step of the Markov decision process; N is the total length of the rounds in the rising-segment guided Markov decision process; k |α Tk | represents the maximum aerodynamic load within the time interval from kdT to (k+1)dT.

6. The method as described in claim 5, characterized in that, In step three, the total number of wind fields generated by the high-resolution hourly upper-level wind forecast data is recorded. Based on the high-resolution hourly upper-level wind forecast data, at the beginning of each round of the ascending-phase guided Markov decision process, a wind field number is randomly generated within the total number of wind fields. The wind field function in the ascending-phase guided Markov decision process model corresponding to this number is selected as the forecast wind field for that round. The method for calculating the difference between the forecast wind field and the actual wind field to obtain the uncertainty bias includes: S1, records the total number of wind fields M generated by producing high-resolution hourly upper-level wind forecast data; S2, at the beginning of each round of the ascending-stage guided Markov decision process, a wind field number j is randomly generated within the range [1,M]; in that round, the wind field function V in the ascending-stage guided Markov decision process model is... wind (h) is the j-th wind field in the high-resolution hourly upper-level wind forecast data; and the j-th wind field is used as the forecast wind field for this round; S3. Repeat step S2 to obtain all forecast wind fields, calculate the difference between all forecast wind fields and actual wind fields to obtain the uncertainty deviation, and store it.

7. The method as described in claim 6, characterized in that, In step four, a value function network is introduced into the policy network to represent the rocket load reduction guidance strategy. The rocket load reduction guidance strategy interacts with the ascent phase guidance Markov decision process model to complete the sampling of a full round of the ascent phase guidance Markov decision process. The methods for obtaining the sampled data include: A value function network is introduced into the policy network to represent the rocket load reduction guidance strategy. The set of sampled trajectories is generated by the interaction between the rocket load reduction guidance strategy and the Markov decision process model of the ascent phase guidance. Record the output probability of the policy network and the output data of the value function network for each sampling trajectory in the sampling trajectory set; at the same time, calculate the temporal difference residual corresponding to each sampling trajectory, and obtain the general advantage estimate of the temporal difference residual through recursive calculation; sum the general advantage estimate and the output data of the value function network to obtain the target value to be fitted by the value function network, and complete the sampling of a complete round of the rising phase guidance Markov decision process; The sampled data consists of the output data of the value function network, the target value to be fitted by the value function network, the general advantage estimate, and the output probability of the policy network.

8. The method as described in claim 7, characterized in that, In step four, the set of sampling trajectories Contains N step =N traj ×N sampling trajectories In the formula, N step N represents the total number of state transition steps. traj N represents the number of complete round-based sampling trajectories, N is the total round length of the rising-segment guided Markov decision process, and i is the sampling trajectory number. Temporal Differential Residual for: In the formula, r k,i Let γ be the reward at the k-th discrete moment of the i-th sampling trajectory, and γ be the discount factor for the cumulative reward. Output data for the value function network. Output data for the value function network at the (k+1)th discrete time step; General advantage estimation for: In the formula, l is the general advantage estimate summation index, N is the total length of the rounds in the rising-segment guided Markov decision process, k is the k-th discrete time of the Markov decision process, and λ is the general advantage estimate decay coefficient. The time-series difference residual is the (k+l)th discrete time step. The target value to be fitted by the value function network for:

9. The method as described in claim 8, characterized in that, In step four, the advantage functions estimated based on the sampled data include: the value function network advantage function and the policy network advantage function; among which, The value function network advantage function is: The policy network advantage function is: In the formula, The value function network advantage function value, Let k be the set of sampled trajectories, k be the k-th discrete time step of the Markov decision process, and i be the sampled trajectory index. Output data for the value function network. The target value to be fitted to the value function network; γ is the advantage function value of the policy network. k η is the discount factor for the cumulative reward at the k-th discrete time. k,i (θ) represents the probability ratio of the policy behavior output at the k-th discrete moment of the i-th sampled trajectory. For general advantage estimation, clip is the clipping function, and ∈ is the trust region radius for gradient updates of policy network parameters; The probability ratio of the policy behavior output at the k-th discrete moment of the i-th sampling trajectory to η k,i (θ) is specifically represented as: in, The probability of the current sampled trajectory is output by the policy network. represents all probabilities of the set of sampled trajectories output by the policy network.

10. The method as described in claim 9, characterized in that, In step four, the update gradients calculated by the advantage function include: the value function network parameter update gradient and the policy network parameter update gradient, where, The gradient for updating the network parameters of the value function is: Update gradients based on value function network parameters To update step size Gradient descent is used to update the network parameters of the value function in the rocket load reduction guidance strategy. The gradient for updating the policy network parameters is: Update gradients based on policy network parameters To update step size α θ Gradient ascent update is performed on the strategy network parameters in the rocket load reduction guidance strategy. In the formula, Update the gradient for the network parameters of the value function; Let the value function network take the gradient with respect to its parameters. Update gradients for policy network parameters.

Citation Information

Patent Citations

  • Aircraft deloading guidance method and device, and storage medium

    CN117872731A

  • Energy-saving optimization method and apparatus for terminal air conditioning system of integrated data center cabinet

    WO2023116742A1