Reinforced learning-based performance recovery control method for vertical landing stage of short vertical propulsion system

Through the method based on reinforcement learning, the training agent optimizes the fuel flow rate and the angle of the imported guide vane of the short-limit propulsion system and solves the problems of high order, strong coupling and serious overshoot in the performance recovery control method of the existing medium- and medium- and short-limit propulsion system in the vertical landing stage, and realizes the stability of thrust output and attitude balance in the case of component degradation.

CN120010246AActive Publication Date: 2025-05-16TSINGHUA UNIVERSITY
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202411973967.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-16
Estimated Expiration
2044-12-30

AI Technical Summary

Technical Problem

The existing performance recovery control method of short vertical propulsion systems in the vertical landing stage has problems such as high order, strong coupling and serious overshoot, making it difficult to effectively restore attitude balance.

Method used

Using a reinforcement learning-based method, by initializing the comment deep neural network and the action deep neural network, combining the degradation factor as state augmentation, the agent is trained to optimize fuel flow and the angle of the imported guide vane of the lift fan to achieve model-free performance recovery control.

Benefits of technology

In the case of component degradation, maintaining the thrust output of the short vertical propulsion system helps the aircraft maintain balance in pitch attitude during the vertical landing phase, reducing the risk of thrust coupling and overshooting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120010246A_ABST
    Figure CN120010246A_ABST
Patent Text Reader

Abstract

The invention provides a reinforcement learning-based performance recovery control method for a vertical landing stage of a short vertical propulsion system. According to the optimization method, the degradation condition of the short vertical propulsion system is considered in the design of a controller through a reinforcement learning algorithm, and a plurality of thrust outputs of the short vertical propulsion system can still be kept stable under the disturbance condition of component degradation through a reinforcement learning pre-trained intelligent agent; the short vertical aircraft is helped to keep the balance of the pitching attitude under the condition that the short vertical aircraft degenerates in the vertical landing stage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of aeroengines, and in particular to a performance recovery control method for a short-to-vertical propulsion system in a vertical landing phase based on reinforcement learning. Background Art

[0002] For short takeoff and vertical landing aircraft and their propulsion systems, the aerodynamic control surfaces cannot function in the hovering state, and the altitude and attitude control can only rely entirely on the multi-source thrust generated by the two lift source nozzles in front and behind the propulsion system. However, in this process, if the rotating parts of the propulsion system have performance degradation, such as high-pressure turbine overheating, lift fan intake distortion, etc., it will cause thrust loss and thrust distribution changes. The loss of thrust will cause the speed of the altitude channel aircraft to change, and the change in thrust distribution will cause an imbalance in the pitch attitude, causing the aircraft to deviate from the designed landing point and may even cause the aircraft to stall. The current performance recovery control methods all compensate for the loss of thrust to keep the thrust of the aircraft unchanged and ensure that the altitude channel is not affected, but there are few attitude recovery control methods for short vertical propulsion systems. In addition, the current performance recovery control methods include model-based methods and model-free methods. Model-based methods rely heavily on model confidence and if model degradation is considered, the controller designed by the model-based method will have a particularly high order, and the amount of calculation is too large to be applied in practice; while model-free performance recovery control methods such as PID are difficult to consider model degradation in the control process, and can only control the parameters that need to be compensated in a one-to-one manner by compensating thrust with fuel, which is easy to cause thrust coupling and overshoot in the short-to-vertical propulsion system, and the simple increase in fuel will affect the attitude balance. Therefore, this paper proposes a model-free performance recovery control method for the short-to-vertical propulsion system based on reinforcement learning, which augments the degradation factor as a state into the intelligent agent training of reinforcement learning, and takes the control amount of the compensated thrust as the action of the intelligent agent, and then uses the intelligent agent obtained after offline reinforcement learning training as the controller, thereby realizing model-free performance recovery control considering degradation.

[0003] This patent provides a performance recovery control method for the vertical landing phase of a short-to-vertical propulsion system based on reinforcement learning. This optimization method takes the degradation of the short-to-vertical propulsion system into account in the design of the controller through a reinforcement learning algorithm. The intelligent agent trained offline through reinforcement learning ensures that under the disturbance of component degradation, the multiple thrust outputs of the short-to-vertical propulsion system can still remain stable, helping the short-to-vertical aircraft to maintain the balance of the pitch attitude online in the event of degradation during the vertical landing phase. Summary of the invention

[0004] In order to solve the problems of high order, strong coupling and severe overshoot of the performance recovery controller of the traditional short-to-vertical propulsion system in the vertical landing phase, a performance recovery control method of the short-to-vertical propulsion system based on reinforcement learning is proposed.

[0005] The purpose of the present invention is achieved through the following technical solutions.

[0006] The present invention discloses a method for controlling the performance of a short-to-vertical propulsion system based on reinforcement learning, comprising the following steps:

[0007] Step 1: Initialize the review deep neural network parameters Q(s,a|θ) in the agent Q ), where s is the state output from the short vertical propulsion system, including thrust sum, thrust ratio, lift fan efficiency factor, fan efficiency factor, high pressure compressor efficiency factor, high pressure turbine efficiency factor and low pressure turbine efficiency factor, that is, s = [TT, TS, η LF ,η FAN ,η HPC ,η HPT ,η LPT ]′, a is the action output from reinforcement learning, which includes fuel flow and lift fan inlet guide vane angle, a=[W f ,IGV]′,θ Q It is the collection of weights and bias parameters in the neural network. The purpose of the deep neural network is to learn an optimal θ Q Function, so that a Q value can be evaluated according to the state s and action a to represent the value of this state-action pair, and then go to step 2;

[0008] Step 2: Initialize the action deep neural network parameters μ(s|θ) in the agent μ ), where s is the state output from the short-to-vertical propulsion system, θ μ is a set of weights and bias parameters in a neural network. The purpose of an action deep neural network is to learn an optimal θ μ The function can then map the output action a according to the state s and proceed to step 3;

[0009] Step 3: Set the input of the short vertical propulsion system to the output a of the action deep neural network, that is, the short vertical propulsion system is used as the environment for the agent to explore, and the reward function of the short vertical propulsion system is calculated according to the parameter changes caused by the output of the action deep neural network. The formula is as follows:

[0010]

[0011] In the formula, t represents the simulation time, r represents the reward, represents the error in thrust and, that is, the performance recovery error of the altitude channel, represents the error of thrust ratio, i.e., the performance recovery error of attitude channel. f(T4) is the over-temperature penalty term, i.e., if it is not over-temperature, f(T4) = 0, if it is over-temperature, f(T4) = -100. After setting the reward, enter an exploration to generate initial experience and enter step 4.

[0012] Step 4: Perform random exploration based on the initialization parameters to obtain the initial state s1 = [TT, TS, η LF ,η FAN ,η HPC ,η HPT ,η LPT ]′, go to step 5;

[0013] Step 5: Select the action at the current time t based on the current strategy and exploration noise. The formula is as follows:

[0014]

[0015] In the formula, a t is the action at time t, μ(s t |θ μ ) is the current state s t In action deep neural network θ μ The output mapped under the parameter setting is is the exploration noise, t Input into the short vertical propulsion system and calculate the current reward r according to the reward function set in step 3 t , and obtain a new state s t+1 , the state-action pair (s t ,a t ,r t ,s t+1 ) is stored in the experience replay buffer, and then goes to step 6;

[0016] Step 6: Sample a small batch of state-action pairs N×(s t ,a t ,r t ,s t+1 ), where N represents the size of the batch. The input of this sampled batch is put into the review deep neural network to calculate the loss function. The calculation formula is as follows:

[0017]

[0018] Where i is the number of cycles in the internal batch, L is the loss function, and represents the target θ to be minimized Q That is, to comment on the parameters of the deep neural network, Q(s i ,a i |θ Q) is the current state-action pair (s i ,a i ) under the current neural network parameters, y i is the learning goal of the neural network, and its formula is as follows:

[0019]

[0020] Where i is the number of cycles of the internal batch, r(s i ,a i ) is the reward for the current sampling batch, γ is the decay coefficient, and θ Q′ Represents the parameters of the target neural network for commenting, that is, we want to reduce the loss function so that the deep neural network θ Q As close as possible to the target neural network θ Q′ , compared to the frequently updated main network, the target network θ Q′ It has a lower update frequency and uses a weighted approach to update parameters from the main network, as shown below:

[0021] θ Q′ =ρθ Q′ +(1-ρ)θ Q

[0022] Where ρ is a hyperparameter between 0 and 1, which determines the degree of soft update. Then the parameters of the action deep neural network are updated and the process goes to step 7.

[0023] Step 7: Update the strategy using the data sampled from the current batch. The formula is as follows:

[0024]

[0025] Where i is the number of cycles in the inner batch, L is the loss function, and θ Q is the parameter of the review deep neural network, θ μ are the parameters of the action deep neural network, s i Indicates the current state. Indicates that in the current state s i The action taken, α represents the learning rate, represents the loss gradient of Q with respect to action a, represents μ relative to θ μ The loss gradient is obtained by selecting Q(s,μ(s) in the current Q value function t |θ μ)) As the estimated value of the Q-value function, it can be seen that the policy gradient on which policy improvement depends is obtained by differentiating the Q-value function multiple times. The method for updating the parameters of the policy target network is to move upward along the mapping gradient of the value function Q, causing the action policy μ to change in the direction of increasing Q-value function, and then update the parameters of the deep neural network for actions:

[0026] θ μ′ = ρθ μ′ +(1 - ρ)θ μ

[0027] In the formula, ρ is a hyperparameter between 0 and 1, which determines the degree of soft update. Then, perform in-batch iteration. If i < N, continue to step 6 to update the neural network. If i = N, enter step 8;

[0028] Step 8: Judge the size relationship between the current time t and the total number of episodes M. If t < M, return to step 3. If t = M, the offline training ends. Configure the agent after the offline training into the environment. The input of the agent is the state of the short takeoff and vertical landing propulsion system s1 = [TT, TS, η LF , η FAN , η HPC , η HPT , η LPT , and the output is the control input a = [W f , IGV]′ of the short takeoff and vertical landing propulsion system, thus realizing the model-free performance recovery control design considering performance degradation. Description of the Drawings

[0029] Figure 1 These are the specific implementation steps of the content of the present invention, "A Performance Recovery Control Method for the Vertical Landing Phase of a Short Takeoff and Vertical Landing Propulsion System Based on Reinforcement Learning". Detailed Embodiment

[0030] Technical Solution:

[0031] This embodiment proposes a performance recovery control method for the vertical landing phase of a short takeoff and vertical landing propulsion system based on reinforcement learning, including the following steps:

[0032] Step 1: Initialize the parameters Q(s, a|θ Q ) of the critic deep neural network in the agent, where s is the state output from the short takeoff and vertical landing propulsion system. The state includes the sum of thrust, thrust ratio, lift fan efficiency factor, fan efficiency factor, high-pressure compressor efficiency factor, high-pressure turbine efficiency factor, and low-pressure turbine efficiency factor, that is, s = [TT, TS, η LF , η FAN , η HPC , η HPT , η LPT]′, a is the action output from reinforcement learning, which includes fuel flow and lift fan inlet guide vane angle, a=[W f ,IGV]′,θ Q It is the collection of weights and bias parameters in the neural network. The purpose of the deep neural network is to learn an optimal θ Q Function, so that a Q value can be evaluated according to the state s and action a to represent the value of this state-action pair, and then go to step 2;

[0033] Step 2: Initialize the action deep neural network parameters μ(s|θ) in the agent μ ), where s is the state output from the short-to-vertical propulsion system, θ μ is a set of weights and bias parameters in a neural network. The purpose of an action deep neural network is to learn an optimal θ μ The function can then map the output action a according to the state s and proceed to step 3;

[0034] Step 3: Set the input of the short vertical propulsion system to the output a of the action deep neural network, that is, the short vertical propulsion system is used as the environment for the agent to explore, and the reward function of the short vertical propulsion system is calculated according to the parameter changes caused by the output of the action deep neural network. The formula is as follows:

[0035]

[0036] In the formula, t represents the simulation time, r represents the reward, represents the error in thrust and, that is, the performance recovery error of the altitude channel, represents the error of thrust ratio, i.e., the performance recovery error of attitude channel. f(T4) is the over-temperature penalty term, i.e., if it is not over-temperature, f(T4) = 0, if it is over-temperature, f(T4) = -100. After setting the reward, enter an exploration to generate initial experience and enter step 4.

[0037] Step 4: Perform random exploration based on the initialization parameters to obtain the initial state s1 = [TT, TS, η LF ,η FaN ,η HPC ,η HPT ,η LPT ]′, go to step 5;

[0038] Step 5: Select the action at the current time t based on the current strategy and exploration noise. The formula is as follows:

[0039]

[0040] In the formula, a t is the action at time t, μ(s t |θμ ) is the current state s t In action deep neural network θ μ The output mapped under the parameter setting is is the exploration noise, t Input into the short vertical propulsion system and calculate the current reward r according to the reward function set in step 3 t , and obtain a new state s t+1 , the state-action pair (s t ,a t ,r t ,s t+1 ) is stored in the experience replay buffer, and then goes to step 6;

[0041] Step 6: Sample a small batch of state-action pairs N×(s t ,a t ,r t ,s t+1 ), where N represents the size of the batch. The input of this sampled batch is put into the review deep neural network to calculate the loss function. The calculation formula is as follows:

[0042]

[0043] Where i is the number of cycles in the internal batch, L is the loss function, and represents the target θ to be minimized Q That is, to comment on the parameters of the deep neural network, Q(s i ,a i |θ Q ) is the current state-action pair (s i ,a i ) under the current neural network parameters, y i is the learning goal of the neural network, and its formula is as follows:

[0044]

[0045] Where i is the number of cycles of the internal batch, r(s i ,a i ) is the reward for the current sampling batch, γ is the decay coefficient, and θ Q′ Represents the parameters of the target neural network for commenting, that is, we want to reduce the loss function so that the deep neural network θ Q As close as possible to the target neural network θ Q′ , compared to the frequently updated main network, the target network θ Q′ It has a lower update frequency and uses a weighted approach to update parameters from the main network, as shown below:

[0046] θ Q′=ρθ Q′ +(1 - ρ)θ Q

[0047] Where ρ is a hyperparameter between 0 and 1, which determines the degree of soft update, and then the parameters of the action deep neural network are updated, and step 7 is entered;

[0048] Step 7: Policy update is performed through the data obtained by current batch sampling, and the formula is as follows:

[0049]

[0050] Where i is the loop count of the internal batch, L is the loss function, θ Q is the parameter of the critic deep neural network, θ μ is the parameter of the action deep neural network, s i represents the current state, represents the action taken in the current state s i α represents the learning rate, represents the loss gradient of Q with respect to the action a, represents the loss gradient of μ with respect to θ μ By selecting Q(s, μ(s t |θ μ )) as the estimated value of the Q-value function, it can be seen that the policy gradient on which policy improvement depends is obtained by differentiating the Q-value function multiple times. The method of updating the parameters of the policy target network is to move upward along the mapping gradient of the value function Q, so that the action policy μ changes in the direction of increasing the Q-value function, and then the parameters of the action deep neural network are updated:

[0051] θ μ′ =ρθ μ′ +(1 - ρ)θ μ

[0052] Where ρ is a hyperparameter between 0 and 1, which determines the degree of soft update, and then in-batch iteration is performed. If i < N, go back to step 6 to continue neural network update. If i = N, enter step 8;

[0053] Step 8: Judge the size relationship between the current time t and the total number of episodes M. If t < M, return to step 3. If t = M, the offline training ends, and the agent after offline training is configured into the environment. The input of the agent is the state of the short vertical propulsion system s1 = [TT, TS, η LF , η FAN , η HPC , η HPT , η LPT , and the output is the control input a = [Wf ,IGV]′, thus achieving model-free performance recovery control design considering performance degradation.

[0054] Compared with the prior art, the reinforcement learning-based performance recovery control method for the vertical landing phase of a short-to-vertical propulsion system obtained in this embodiment can take into account the degradation situation without an accurate model, and realize the performance recovery of the height and attitude of the vertical landing phase of a model-free short-to-vertical propulsion system taking into account performance degradation.

Claims

1. A performance recovery control method for a short-to-vertical propulsion system during vertical landing based on reinforcement learning, characterized in that It includes the following steps: Step 1: Initialize the review deep neural network parameters Q(s, a|θ) in the agent Q ), where s is the state output from the short vertical propulsion system, including thrust sum, thrust ratio, lift fan efficiency factor, fan efficiency factor, high pressure compressor efficiency factor, high pressure turbine efficiency factor and low pressure turbine efficiency factor, that is, s = [TT, TS, η LF , η FAN , η HPC , η HPT , η LPT ]′, a is the action output from reinforcement learning, which includes fuel flow and lift fan inlet guide vane angle, a=[W f ,IGV]′,θ Q It is the collection of weights and bias parameters in the neural network. The purpose of the deep neural network is to learn an optimal θ Q Function, so that a Q value can be evaluated according to the state s and action a to represent the value of this state-action pair, and then go to step 2; Step 2: Initialize the parameters μ(s|θμ) of the action deep neural network in the agent, where s is the state output from the short takeoff and vertical landing propulsion system, and θμ is the set of parameters of weights and biases in the neural network. The purpose of the action deep neural network is to learn an optimal θμ function so as to map the output action a according to the state s, and then proceed to Step 3; Step 3: Set the input of the short takeoff and vertical landing propulsion system to the output a of the action deep neural network, that is, the short takeoff and vertical landing propulsion system serves as the environment for the agent to explore. Calculate the reward function based on the parameter changes caused by the output of the action deep neural network in the short takeoff and vertical landing propulsion system. The formula is as follows: In the formula, t represents the simulation time, r represents the reward, represents the error in thrust and, that is, the performance recovery error of the altitude channel, The error representing the thrust ratio, i.e., the performance recovery error of the attitude channel, f(T4) is the over-temperature penalty term, i.e., if it is not over-temperature, f(T4) = 0, if it is over-temperature, T(T4) = -100, after setting the reward, enter an exploration to generate initial experience, and enter step 4; Step 4: Perform random exploration based on the initialization parameters to obtain the initial state s1 = [TT, TS, η LF , η FAN , η HPC , η HPT , η LPT ]′, go to step 5; Step 5: Select the action at the current moment t according to the current policy and exploration noise. The formula is as follows: In the formula, a t is the action at time t, μ(s t |θ μ ) is the current state s t In action deep neural network θ μ The output mapped under the parameter setting is is the exploration noise, t Input into the short vertical propulsion system and calculate the current reward r according to the reward function set in step 3 t , and obtain a new state s t+1 , the state-action pair (s t , a t , r t ,s t+1 ) is stored in the experience replay buffer, and then goes to step 6; Step 6: Sample a small batch of state-action pairs N×(s t , a t , r t ,s t+1 ), where N represents the size of the batch. The input of this sampled batch is put into the review deep neural network to calculate the loss function. The calculation formula is as follows: Where i is the number of cycles in the internal batch, L is the loss function, and represents the target θ to be minimized Q That is, to comment on the parameters of the deep neural network, Q(s i , a i |θ Q ) is the current state-action pair (s i , a i ) under the current neural network parameters, y i is the learning goal of the neural network, and its formula is as follows: Where i is the number of cycles of the internal batch, r(s i , a i ) is the reward for the current sampling batch, γ is the decay coefficient, and θ Q′ Represents the parameters of the target neural network for commenting, that is, we want to reduce the loss function so that the deep neural network θ Q As close as possible to the target neural network θ Q′ , compared to the frequently updated main network, the target network θ Q′ It has a lower update frequency and uses a weighted approach to update parameters from the main network, as shown below: i Q′ =ρθ Q′ +(1-p)θ Q In the formula, ρ is a hyperparameter between 0 and 1, which determines the degree of soft update. Then update the parameters of the action deep neural network and proceed to Step 7; Step 7: Update the policy through the data obtained by sampling in the current batch. The formula is as follows: Where i is the number of cycles in the internal batch, L is the loss function, and θ Q is the parameter of the review deep neural network, θ μ are the parameters of the action deep neural network, s i Indicates the current state. Indicates that in the current state s i The action taken, α represents the learning rate, represents the loss gradient of Q with respect to action a, represents the loss gradient of μ with respect to θμ, by selecting Q(s, μ(s) in the current Q-value function t |θ μ )) As the estimated value of the Q-value function, it can be seen that the policy gradient on which the policy improvement depends is derived from the multiple derivations of the Q-value function. The method to update the policy target network parameters is to move upward along the mapping gradient of the value function Q, so that the action policy μ changes in the direction of increasing the Q-value function, and then update the action deep neural network parameters: i μ′ =ρθ μ′ +(1-p)θ μ In the formula, ρ is a hyperparameter between 0 and 1, which determines the degree of soft update. Then perform iteration within the batch. If i < N, continue to return to Step 6 for neural network update. If i = N, proceed to Step 8; Step 8: Determine the relationship between the current time t and the total number of rounds M. If t < M, return to step 3. If t = M, the offline training is completed and the agent after offline training is configured into the environment. The input of the agent is the state of the short vertical propulsion system s1 = [TT, TS, η LF , η FAN , η HPC , η HPT , η LPT ], the output is the control input of the short vertical propulsion system a=[W f , IGV]′, thereby achieving model-free performance recovery control design considering performance degradation.

2. The method for controlling the performance of a short-to-vertical propulsion system in the vertical landing phase based on reinforcement learning according to claim 1 is characterized in that The degradation augmentation method in step 1 is characterized in that: the state quantity of the intelligent agent in the conventional propulsion system tracking process includes s=[TT, TS]′, and this method takes the degradation factor coupling into account in the state quantity, that is, the state quantity is augmented, specifically s=[TT, TS, η LF , η FAN , η HPC , η HPT , η LPT ]′, by augmenting the control quantity, the intelligent agent can select a more appropriate action output a according to the changes in thrust and thrust ratio caused by the changes in the degradation factors of different components during offline training.

3. The method for controlling the performance of a short-to-vertical propulsion system in the vertical landing phase based on reinforcement learning according to claim 1 is characterized in that The reward function facility method in step 3 is characterized in that conventional performance recovery control only considers eliminating the thrust and error of the altitude channel by compensating the fuel Unable to consider the error of attitude channel It is also impossible to explicitly characterize the performance degradation in the controller, and the overtemperature in the compensation process cannot be eliminated by model-free control. However, step 3 of the present invention sets the calculation formula This enables the agent to not only eliminate the error of the height channel It can also eliminate the attitude channel The temperature before the turbine is also used as a penalty item f(T4) to prevent overheating. At the same time, the calculation of the entire reward is performed through the augmented state quantity and action described in claim 2, and the propulsion system degradation factor is taken into account in the entire training process, which makes up for the shortcomings of the current method.

Citation Information

Patent Citations

  • Turbofan engine direct thrust intelligent control method based on reinforcement learning

    CN114527654A

  • Aero-engine transition state optimization control method based on reinforcement learning

    CN114675535A

  • Active-disturbance-rejection controller design method and device, and storage medium

    CN115903510A

  • Carrier-based aircraft automatic carrier landing intelligent guidance method and system based on reinforcement learning

    CN117742147A

  • Landing guidance method based on combination of expert data and reinforcement learning

    CN117828980A