A layered phase adaptive hybrid electric propulsion energy management method
Patent Information
- Application Number
- CN202610686511.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-19
- Publication Date
- 2026-09-25
AI Technical Summary
但是,独立的强化学习方案在起飞、爬升等安全关键阶段难以提供严格可认证的约束满足保证,其黑箱特性使得行为可解释性和鲁棒性不足,且缺乏有效的安全回退机制,当学习策略因分布外输入而性能退化时,系统可能面临违规风险
(1)本发明提出的一种分层相位自适应混合电推进能量管理方法,通过分层模型预测控制与SAC强化学习的深度融合,并引入相位自适应的连续加权融合机制,能够在起飞、爬升等安全关键阶段由MPC主导输出,严格满足电池荷电状态、温度及功率变化率等硬约束,同时在巡航等稳态阶段让SAC策略主导,以熵正则化奖励函数驱动长期燃油经济性与电池荷电状态维持的联合优化,从而在全飞行任务中实现安全性与能效的动态平衡;
Smart Images

Figure CN122808971A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of energy management technology, and specifically to a hierarchical phase adaptive hybrid electric propulsion energy management method. Background Technology
[0002] In the field of energy management for hybrid electric propulsion aircraft, especially series hybrid electric propulsion distributed propulsion aircraft, existing technologies are mainly divided into two categories. The first category is optimization methods based on model predictive control (MRC). These methods solve constrained optimal control problems in the prediction time domain, effectively handling safety constraints such as battery state of charge and power limits, and ensuring the recursive feasibility of control. However, MRC methods typically rely on accurate predictions of future flight paths, and their performance degrades significantly in the presence of uncertainties such as wind disturbances and model mismatch. Furthermore, to meet hard constraints, their control strategies are often overly conservative, making it difficult to achieve optimal energy allocation and battery recharging during low-power demand phases such as cruise. In addition, the computational burden of online optimization for long-time domain or high-fidelity models is heavy, limiting their deployment in real-time airborne environments. The second category is reinforcement learning-based methods. Algorithms such as Soft Actor-Critic (SAC) can learn energy management strategies adapted to various mission profiles through offline training and achieve good fuel economy under steady-state conditions. However, independent reinforcement learning schemes struggle to provide rigorously verifiable constraint satisfaction guarantees during safety-critical phases such as takeoff and climb. Their black-box nature results in insufficient behavioral interpretability and robustness, and the lack of effective safety rollback mechanisms means the system may face compliance risks when the learning strategy degrades due to distributed external inputs. Therefore, current technologies have yet to provide a unified energy management framework that can guarantee strict safety during critical phases, achieve energy efficiency optimization throughout the entire flight mission, and possess real-time feasibility and robust rollback capabilities. Summary of the Invention
[0003] To provide a unified energy management framework that ensures strict safety during critical phases, optimizes energy efficiency throughout the entire flight mission, and possesses real-time feasibility and robust rollback capability, this invention proposes a hierarchical phase-adaptive hybrid electric propulsion energy management method, comprising the following steps: S1: In each control cycle, acquire the current status information of the aircraft and the prediction information of future flight missions. The status information includes battery state of charge, battery temperature, power requirements and flight phase parameters. S2: Using the hierarchical model predictive control module, long-term high-level planning is performed based on the predictive information to generate a global energy reference trajectory. Then, short-term low-level tracking control is performed based on the global energy reference trajectory and the current state information to calculate the reference control signal. S3: Utilize a pre-trained SAC reinforcement learning model to generate random action signals based on the current state information; S4: Dynamically determine the hybrid weight based on the current flight phase. In the safety-critical phase, the weight is biased towards the reference control signal, and in the steady-state cruise phase, the weight is biased towards the random action signal. Then, the two are combined into the final control command through a weighted fusion method. S5: Issue the final control command to the propulsion system for execution; S6: During the operation of the reinforcement learning model, the output of the model is continuously monitored. When the deviation between the output and the reference control signal exceeds a preset threshold, or when the output exceeds the safe action boundary, the model is switched to pure model prediction control mode for safe rollback, and the propulsion system is controlled solely by the reference control signal. The final control command is only used to allocate the output power of the engine and battery respectively, while the speed of the aircraft is generated by an independent feedback controller based on thrust-drag dynamics.
[0004] This invention utilizes phase adaptive weighted fusion of layered MPC and SAC, along with independent speed decoupling control. During the safety-critical phase, MPC takes the lead to ensure strict constraint satisfaction, while SAC takes the lead to improve fuel economy during the cruise phase. At the same time, output deviation monitoring and a safety backoff mechanism ensure system robustness.
[0005] Furthermore, in step S2, the prediction time domain of the long-time domain high-level planning covers the entire remaining flight mission, the prediction time domain of the short-time domain low-level tracking control is a single control cycle, and the sampling period of the high-level planning is a preset multiple range of the sampling period of the low-level tracking control.
[0006] Furthermore, in step S2, the hierarchical model predictive control module uses the battery state of charge, battery temperature, and power change rate as hard constraints when calculating the baseline control signal, and solves a constrained quadratic programming problem in each control cycle.
[0007] Furthermore, in step S3, the reward function in the pre-trained SAC reinforcement learning model includes fuel consumption reward, battery state of charge maintenance reward, constraint violation penalty, and entropy regularization term.
[0008] Furthermore, in step S4, the mixed weights are dynamically generated through a continuous function and processed using exponential smoothing filtering and rate of change constraints.
[0009] Furthermore, in step S4, in the weighted fusion method, the mixed weight is obtained by looking up a predefined weight table according to the current flight phase. When the flight phase is identified as being in a region with ambiguous boundaries, the mixed weight is set to a preset global conservative value, so that the reference control signal dominates.
[0010] Furthermore, in step S6, after detecting that the output exceeds the safe action boundary and switching to pure model prediction control mode, the following steps are also included: The output of the reinforcement learning model is continuously monitored. When the output continuously meets the safety action boundary within a preset time window and the deviation from the baseline control signal is less than a preset threshold, the propulsion system is restored to control via weighted fusion.
[0011] Furthermore, the independent feedback controller is a proportional-integral controller, the speed of the aircraft is preset according to the current flight phase, and the speed of the aircraft is smoothly transitioned when switching phases.
[0012] Furthermore, the SAC reinforcement learning model is updated through a priority experience replay mechanism, wherein the priority of the experience samples is determined by a weighted sum of the temporal difference error and the deviation term of the action signal output by the reinforcement learning model relative to the reference control signal.
[0013] Compared with the prior art, the present invention has at least the following beneficial effects: (1) The hierarchical phase adaptive hybrid electric propulsion energy management method proposed in this invention deeply integrates hierarchical model predictive control and SAC reinforcement learning, and introduces a phase adaptive continuous weighted fusion mechanism. It can be dominated by MPC output during safety-critical phases such as takeoff and climb, strictly satisfying hard constraints such as battery state of charge, temperature and power change rate. At the same time, it allows SAC strategy to dominate during steady-state phases such as cruise, and drives the joint optimization of long-term fuel economy and battery state of charge maintenance with entropy regularization reward function, thereby achieving a dynamic balance between safety and energy efficiency in the entire flight mission. (2) By continuously monitoring the deviation between the SAC output and the reference control signal and the safety boundary, when an abnormality is detected, the system immediately switches to pure MPC mode for safe rollback, and automatically switches back to fusion control after the conditions return to normal, thereby improving the fault tolerance and robustness of the system. (3) The energy management command is limited to allocating the output power of the engine and the battery respectively, while the flight speed is generated by an independent proportional-integral controller based on thrust-drag dynamics, reducing the online calculation burden. Combined with the conservative backoff strategy of weight smooth transition and boundary ambiguity stage, the continuity of control commands and physical consistency are guaranteed. Attached Figure Description
[0014] Figure 1 This is a step diagram of a hierarchical phase-adaptive hybrid electric propulsion energy management method. Detailed Implementation
[0015] The following are specific embodiments of the present invention, which are described in conjunction with the accompanying drawings to further illustrate the technical solutions of the present invention. However, the present invention is not limited to these embodiments.
[0016] In the field of energy management for hybrid electric propulsion aircraft, especially for series hybrid electric distributed propulsion (DEP) configurations, achieving power distribution and energy optimization throughout the entire flight mission cycle while meeting stringent safety constraints has always been a technical challenge. Existing technologies mainly fall into two categories: one is based on model predictive control (MPC), which solves constrained optimization problems online and can effectively guarantee the satisfaction of hard constraints such as battery state of charge, temperature, and power limits. However, its control strategies are often conservative, making it difficult to fully utilize the optimization space for energy recovery and battery recharging during low-demand phases such as cruise. Furthermore, it has a heavy computational burden and is highly dependent on flight path prediction. The other category is based on reinforcement learning (RL), which learns long-term optimal strategies through offline training, exhibiting good adaptability and efficiency optimization capabilities. However, it struggles to provide strict deterministic safety guarantees during safety-critical phases such as takeoff and climb, and its strategy behavior is difficult to interpret, lacking a reliable backoff mechanism. Therefore, existing technologies have not yet been able to simultaneously address safety and energy optimality throughout the entire mission cycle within the same framework, and also lack an adaptive fusion mechanism that can smoothly switch control weights according to flight phases. To address the aforementioned technical issues, such as Figure 1 As shown, this invention proposes a hierarchical phase-adaptive hybrid electric propulsion energy management method, comprising the following steps: S1: In each control cycle, acquire the current status information of the aircraft and the prediction information of future flight missions. The status information includes battery state of charge, battery temperature, power requirements and flight phase parameters. S2: Using the hierarchical model predictive control module, long-term high-level planning is performed based on the predictive information to generate a global energy reference trajectory. Then, short-term low-level tracking control is performed based on the global energy reference trajectory and the current state information to calculate the reference control signal. S3: Utilize a pre-trained SAC reinforcement learning model to generate random action signals based on the current state information; S4: Dynamically determine the hybrid weight based on the current flight phase. In the safety-critical phase, the weight is biased towards the reference control signal, and in the steady-state cruise phase, the weight is biased towards the random action signal. Then, the two are combined into the final control command through a weighted fusion method. S5: Issue the final control command to the propulsion system for execution; S6: During the operation of the reinforcement learning model, the output of the model is continuously monitored. When the deviation between the output and the reference control signal exceeds a preset threshold, or when the output exceeds the safe action boundary, the model is switched to pure model prediction control mode for safe rollback, and the propulsion system is controlled solely by the reference control signal.
[0017] The final control command is used only to allocate the output power of the engine and battery, while the speed of the aircraft is generated by an independent feedback controller based on thrust-drag dynamics.
[0018] In this embodiment, a series-type hybrid electric propulsion distributed propulsion aircraft is used as an example to describe the method of the present invention in detail. The propulsion system architecture of the aircraft mainly includes an engine-generator set, a battery pack, a power distribution system, and multiple distributed electric drive thrusters. The engine-generator set consumes fuel to generate electricity, the battery pack provides peak power as an auxiliary energy source and can be charged by the engine-generator set under low load conditions, and the power distribution system distributes power to each thruster according to control commands.
[0019] In this embodiment, the energy management problem of the aircraft is modeled as a discrete-time dynamic system. A state vector is defined. This includes, but is not limited to, parameters such as the aircraft's position coordinates, speed, altitude, battery state of charge (SOC), battery temperature, engine speed, and recovered energy reserves. Control inputs. This includes the total power command of the propulsion system, the power distribution ratio between the engine-generator set and the battery, and the voltage commands for each distributed thruster. External input. This includes the predicted recoverable power, atmospheric density variations, and propeller slipstream effects within the prediction time domain. The system state transition satisfies the state equation. ,in This is the discrete-time dynamic function of the system, which comprehensively reflects the physical relationships such as the energy conversion efficiency of the propulsion system, the charging and discharging characteristics of the battery, and the aerodynamic-propulsion coupling effect of the aircraft.
[0020] To simultaneously satisfy safety constraints and long-term energy optimality throughout the entire flight mission cycle, this invention employs a control architecture that integrates hierarchical model predictive control and reinforcement learning. This method performs the following operations in each control cycle.
[0021] First, at the start of each control cycle, the aircraft's onboard data acquisition system obtains current status information and future flight mission predictions. Current status information primarily includes battery state of charge, battery temperature, current power requirements, and flight phase parameters provided by the flight management system. Flight phase parameters are determined in real-time using simple threshold logic, such as determining whether the aircraft is in the takeoff, climb, cruise, descent, or landing phase based on current altitude, airspeed, vertical speed, and thrust setpoints. Future flight mission predictions are derived from pre-loaded flight plans in the flight management system, including remaining waypoints, estimated duration of each segment, altitude profiles, and predicted atmospheric environmental parameters.
[0022] After obtaining the above information, the hierarchical model predictive control module begins operation. This module consists of a high-level planning layer and a low-level tracking control layer, which operate at different time scales.
[0023] The high-level planning layer operates at a slower sampling rate, with a sampling period of, for example, 1 second, and its prediction time domain covers the entire remaining flight mission. Based on a simplified coarse-grained system dynamics model, the high-level planning layer solves the open-loop optimal control problem. Its objective function includes energy consumption costs and terminal costs at each time step, while constraints include global energy balance, waypoint arrival constraints, and a lower bound constraint on battery state of charge. By solving this optimization problem, the high-level planning layer generates a global energy reference trajectory, which provides the optimal average power allocation scheme and battery state of charge reference curves for each stage of the future flight mission.
[0024] The lower-level tracking control layer operates at a faster sampling rate, for example, a sampling period of 0.1 to 0.2 seconds. Its prediction time domain is typically set to a single control cycle or several control cycles, significantly shorter than the prediction time domain of the higher-level layers. The lower-level tracking control layer uses the global energy reference trajectory generated by the higher-level planning layer as the tracking target, and combines this with current precise flight state information to solve a quadratic programming-type tracking problem within a shorter prediction time domain. The objective function of this tracking problem includes a tracking deviation term from the reference trajectory and a penalty term for control variable changes. The constraints employ a refined system model and strict hard constraints, including instantaneous upper and lower limits of battery state of charge, a safe boundary for battery temperature, a power change rate limit, and actuator saturation limits. Within each control cycle, the lower-level tracking control layer calculates a baseline control signal that satisfies all hard constraints by solving this constrained quadratic programming problem. This baseline control signal represents the power allocation command that the model predictive control module considers safe and feasible under the current operating conditions.
[0025] Meanwhile, the pre-trained SAC reinforcement learning model operates in parallel. This model is a deep reinforcement learning model based on the maximum entropy reinforcement learning framework. Its policy network outputs a Gaussian distribution of mean and covariance matrices, from which random action signals are sampled. These random action signals are represented as... The SAC model satisfies action boundary constraints (e.g., power output range limits). Unlike traditional deterministic policies, the SAC model encourages policy stochasticity through entropy regularization, thereby gaining better exploration capabilities and robustness to model uncertainty during training. In this implementation, the SAC model employs a dual-Q network structure to mitigate Q-value overestimation and uses an adaptive temperature coefficient to maintain the desired entropy level. The SAC model's observation space includes battery state of charge, instantaneous power demand, battery temperature, instantaneous fuel consumption rate, flight phase indicators, previous control commands, and current recovery efficiency.
[0026] Furthermore, the reward function of the SAC model is designed to guide the model in learning strategies that align with the mission objectives. Specifically, the reward function includes a fuel consumption reward term, which is negative and proportional to the instantaneous fuel consumption rate, encouraging the model to reduce fuel consumption; a battery state-of-charge (SOC) maintenance reward term, which encourages the battery SOC to remain within its optimal operating range, avoiding over-discharge or over-charge; a constraint violation penalty term, which imposes a large penalty when the model's output action causes the battery SOC to fall below the lower limit, the battery temperature to exceed the safety threshold, or the power change to exceed the allowable range; and an entropy regularization term, which encourages exploratory behavior. The overall form of the reward function can be expressed as a linearly weighted combination of the terms, with the weights varying depending on the flight phase. For example, during takeoff and climb, the weight of the constraint violation penalty term increases significantly to ensure safety is prioritized; while during cruise, the weight of the fuel consumption reward term increases relatively to optimize economy.
[0027] The SAC model was trained offline before the flight mission. The training process took place in a high-fidelity distributed electric propulsion aircraft simulation environment, which fully modeled the nonlinear dynamic characteristics of the series hybrid electric propulsion system, including engine fuel consumption characteristics, battery charging and discharging internal resistance and thermal effects, motor efficiency, aerodynamic interference between distributed thrusters, and regenerative braking energy recovery. Training covered the complete flight mission profile, employing a priority experience playback mechanism to update model parameters. The priority of experience samples was determined by a weighted sum of their temporal difference errors and the deviation of the model's output action from the baseline control signal. This priority setting strategy encouraged the model to focus more on state regions where there was a significant discrepancy between model predictive control and reinforcement learning decisions—regions with potentially large learning gains.
[0028] The reference control signals output by the model predictive control module are obtained respectively. and the random action signal output by the SAC model Afterward, the system enters the hybrid fusion phase. The core of hybrid fusion lies in dynamically determining a hybrid weight between 0 and 1 based on the current flight phase. Then, the two signals are combined into the final control command by weighted summation. .
[0029] Mixed weights The determination strategy embodies the core idea of phase adaptation. During safety-critical flight phases, such as takeoff and climb, due to high thrust demands, drastic dynamic changes, and extremely low tolerance for violations of battery state of charge and power limits, [the strategy is as follows]. Setting the weighting to a small value, such as 0.3, biases the weighted result towards the baseline control signal output by the model predictive control module, meaning that model predictive control dominates the control weights, and the SAC model only provides minor corrections. During the steady-state cruise phase, due to relatively relaxed constraints, airspeed and altitude remain stable, and the primary objective is long-term efficiency optimization. At this point, [the weighting is adjusted accordingly]. Setting it to a larger value, such as 0.8, will bias the weighting result towards the random action signal output by the SAC model, thus fully leveraging the advantages of reinforcement learning in efficiency optimization and adaptation.
[0030] In this embodiment, the specific values of the hybrid weights are determined in advance through simulation iteration optimization. The initial weight ratios reference engineering practice experience of integrating hierarchical model predictive control and reinforcement learning in similar safety-critical systems, i.e., model predictive control is assigned higher weights in hard constraint handling scenarios, and reinforcement learning is assigned higher weights in performance optimization scenarios. Subsequently, a large number of closed-loop simulations are performed in a high-fidelity simulation environment, using constraint violation rate, fuel consumption, trajectory tracking error, etc., as evaluation indicators, and the weight coefficients of each flight phase are adjusted in increments of 0.1 until a satisfactory balance between safety and efficiency is achieved under different mission profiles. For ambiguous regions near the boundaries of flight phases, sensor noise or deviations from the nominal operating conditions may lead to unreliable phase identification. In this case, the system switches the hybrid weights to a preset global conservative value, for example... In other words, model predictive control takes the lead, ensuring the system can still operate safely when there is uncertainty in stage identification. In addition, to prevent abrupt changes in weights from causing jumps in control variables, the mixed weights undergo exponential smoothing filtering and rate of change limiting during the generation process, thus making the transition smoother.
[0031] Final control command This command is used only to allocate the output power of the engine-generator set and the battery pack, and does not directly participate in the aircraft's closed-loop speed control. In other words, this command determines how much of the total power demand of the propulsion system is handled by the engine-generator set and how much is supplemented or absorbed by the battery pack, thus achieving power distribution. The aircraft's speed control is handled independently by a separate proportional-integral (PI) feedback controller. This PI controller calculates the required total thrust command based on the preset speed target value for the current flight phase and the deviation between the current airspeed and the target airspeed. This thrust command is then converted into the total power demand of the propulsion system through thrust-drag dynamics, thus forming a cascaded structure with the energy management layer. This decoupled design eliminates the need for the energy management module to solve complex trajectory tracking problems online, greatly reducing computational complexity. Simultaneously, the speed target value undergoes smooth transition processing during flight phase switching, such as through first-order inertial filtering or S-curve planning, to avoid drastic fluctuations in propulsion system power caused by abrupt changes in the speed command during phase transitions.
[0032] Once the final control command is generated, it is sent to the aircraft's propulsion system for execution. The power distribution system adjusts the output power of the engine-generator set and the charging and discharging power of the battery pack according to this command, and distributes electrical energy to each distributed thruster.
[0033] In the energy management method of this invention, in addition to the core control process described above, an independent safety monitoring and rollback mechanism is also provided. This mechanism continuously monitors the output of the reinforcement learning model throughout its operation. Specifically, the monitoring includes two aspects: First, monitoring the deviation between the random action signal output by the SAC model and the baseline control signal output by the model predictive control module. When this deviation exceeds a preset threshold, it indicates a serious discrepancy between the reinforcement learning model's decision and the conservative, safe baseline decision, potentially indicating a risk. Second, monitoring whether the action signal output by the SAC model exceeds preset safe action boundaries, such as the maximum permissible power change rate or the maximum battery discharge power.
[0034] When any of the above monitoring conditions are triggered, the system immediately switches from weighted fusion control mode to pure model predictive control mode for a safety rollback. In pure model predictive control mode, the final control command is directly determined by the reference control signal output by the model predictive control module, i.e., it is forcibly set. The output of the SAC model is completely bypassed. At this point, the propulsion system is entirely controlled by model predictive control commands that satisfy all hard constraints, thereby ensuring flight safety.
[0035] After the system enters the safety fallback mode, the monitoring module continues to monitor the output status of the SAC model. When the model's output consistently meets the following two conditions within a preset time window: the output signal always falls within the safe action boundary, and the deviation between the output signal and the baseline control signal remains less than a preset threshold, the system will automatically exit the pure model predictive control mode and revert to the weighted fusion control mode. This automatic recovery mechanism avoids permanent locking into the conservative mode, which would lead to long-term energy efficiency losses, while ensuring that control is only re-granted to the reinforcement learning model when it has indeed recovered to a reliable state.
[0036] The aforementioned safety fallback mechanism provides enhanced security for learning-based control methods. Unlike constraint penalties or safety exploration techniques used only during the training phase, it runs as an independent background process in real-time during actual online control, providing deterministic safety redundancy in the event of any anomalous output from the reinforcement learning model.
[0037] In summary, this invention, through the deep fusion of hierarchical model predictive control and SAC reinforcement learning, and the introduction of a phase-adaptive continuous weighted fusion mechanism, enables MPC to dominate output during safety-critical phases such as takeoff and climb, strictly satisfying hard constraints such as battery state of charge, temperature, and power change rate. Simultaneously, during steady-state phases such as cruise, SAC strategy takes the lead, using an entropy-regularized reward function to drive the joint optimization of long-term fuel economy and battery state of charge maintenance, thereby achieving a dynamic balance between safety and energy efficiency throughout the entire flight mission. Furthermore, by continuously monitoring the deviation between SAC output and the reference control signal and the safety boundary, when an anomaly is detected, the system immediately switches to pure MPC mode for safety backoff, and automatically switches back to fused control after conditions return to normal, significantly improving the system's fault tolerance and robustness. In addition, limiting energy management commands to allocating the output power of the engine and battery respectively, while flight speed is generated by an independent proportional-integral controller based on thrust-drag dynamics, effectively reduces the online computational burden. Combined with a weighted smooth transition and a conservative backoff strategy during boundary ambiguity phases, the continuity and physical consistency of control commands are ensured. Ultimately, this method can reduce fuel consumption and emissions, extend battery cycle life, and has good real-time deployment capabilities and adaptability to different flight missions while meeting stringent flight safety requirements.
[0038] It should be noted that all directional indications (such as up, down, left, right, front, back, etc.) in the embodiments of the present invention are only used to explain the relative positional relationship and movement of each component in a certain specific posture (as shown in the figure). If the specific posture changes, the directional indication will also change accordingly.
[0039] Furthermore, in this invention, descriptions involving terms such as "first," "second," and "a" are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0040] In this invention, unless otherwise explicitly specified and limited, the terms "connection," "fixed," etc., should be interpreted broadly. For example, "fixed" can mean a fixed connection, a detachable connection, or an integral part; it can mean a mechanical connection or an electrical connection; it can mean a direct connection or an indirect connection through an intermediate medium; it can mean the internal communication of two components or the interaction between two components, unless otherwise explicitly limited. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0041] Furthermore, the technical solutions of the various embodiments of the present invention can be combined with each other, but only if they are feasible for those skilled in the art. If the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such combination of technical solutions does not exist and is not within the scope of protection claimed by the present invention.
Claims
1. A layered phase-adaptive hybrid electric propulsion energy management method, characterized in that, Including the following steps: S1: In each control cycle, acquire the current status information of the aircraft and the prediction information of future flight missions. The status information includes battery state of charge, battery temperature, power requirements and flight phase parameters. S2: Using the hierarchical model predictive control module, long-term high-level planning is performed based on the predictive information to generate a global energy reference trajectory. Then, short-term low-level tracking control is performed based on the global energy reference trajectory and the current state information to calculate the reference control signal. S3: Utilize a pre-trained SAC reinforcement learning model to generate random action signals based on the current state information; S4: Dynamically determine the hybrid weight based on the current flight phase. In the safety-critical phase, the weight is biased towards the reference control signal, and in the steady-state cruise phase, the weight is biased towards the random action signal. Then, the two are combined into the final control command through a weighted fusion method. S5: Issue the final control command to the propulsion system for execution; S6: During the operation of the reinforcement learning model, the output of the model is continuously monitored. When the deviation between the output and the reference control signal exceeds a preset threshold, or when the output exceeds the safe action boundary, the model is switched to pure model prediction control mode for safe rollback, and the propulsion system is controlled solely by the reference control signal. The final control command is only used to allocate the output power of the engine and battery respectively, while the speed of the aircraft is generated by an independent feedback controller based on thrust-drag dynamics.
2. The layered phase adaptive hybrid electric propulsion energy management method as described in claim 1, characterized in that, In step S2, the prediction time domain of the long-time domain high-level planning covers the entire remaining flight mission, the prediction time domain of the short-time domain low-level tracking control is a single control cycle, and the sampling period of the high-level planning is a preset multiple range of the sampling period of the low-level tracking control.
3. The layered phase adaptive hybrid electric propulsion energy management method as described in claim 1, characterized in that, In step S2, the hierarchical model predictive control module uses the battery state of charge, battery temperature and power change rate as hard constraints when calculating the baseline control signal, and solves a constrained quadratic programming problem in each control cycle.
4. The layered phase adaptive hybrid electric propulsion energy management method as described in claim 1, characterized in that, In step S3, the reward function in the pre-trained SAC reinforcement learning model includes fuel consumption reward, battery state of charge maintenance reward, constraint violation penalty, and entropy regularization term.
5. The layered phase adaptive hybrid electric propulsion energy management method as described in claim 1, characterized in that, In step S4, the mixed weights are dynamically generated through a continuous function and processed using exponential smoothing filtering and rate of change constraints.
6. The layered phase adaptive hybrid electric propulsion energy management method as described in claim 1, characterized in that, In step S4, the weighted fusion method obtains the mixed weight from a predefined weight table based on the current flight phase. When the flight phase is identified as being in a region with ambiguous boundaries, the mixed weight is set to a preset global conservative value, so that the reference control signal dominates.
7. The layered phase adaptive hybrid electric propulsion energy management method as described in claim 1, characterized in that, In step S6, after detecting that the output exceeds the safe action boundary and switching to pure model prediction control mode, the following steps are also included: The output of the reinforcement learning model is continuously monitored. When the output continuously meets the safety action boundary within a preset time window and the deviation from the baseline control signal is less than a preset threshold, the propulsion system is restored to control via weighted fusion.
8. The layered phase adaptive hybrid electric propulsion energy management method as described in claim 1, characterized in that, The independent feedback controller is a proportional-integral controller. The speed of the aircraft is preset according to the current flight phase, and the speed of the aircraft is smoothly transitioned when switching phases.
9. The layered phase adaptive hybrid electric propulsion energy management method as described in claim 1, characterized in that, The SAC reinforcement learning model is updated through a priority experience replay mechanism, wherein the priority of the experience samples is determined by a weighted sum of the temporal difference error and the deviation term of the action signal output by the reinforcement learning model relative to the reference control signal.