A range-extending hybrid propulsion dual-source dynamic coupling energy management method
By adopting a dual-source dynamic coupling energy management method for extended-range hybrid propulsion with a full-process design, and combining flight scenario modeling and reinforcement learning to optimize power allocation, the problem of low energy efficiency and insufficient adaptability of energy management strategies under complex operating conditions in existing technologies has been solved, and efficient energy utilization and stable control of flight equipment have been achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
- Filing Date
- 2026-02-06
- Publication Date
- 2026-04-10
AI Technical Summary
Existing technologies have low energy management efficiency under complex flight conditions. The dual-source power distribution does not take into account differences in flight load and dynamic response, resulting in mode switching conflicts, discontinuous power supply, and start-stop disturbances, which affect flight stability and control experience. Furthermore, the model design is not adapted to flight dynamics characteristics, resulting in high engineering load and reduced control accuracy when conditions change abruptly.
Through a fully designed dual-source dynamic coupling energy management method for range-extended hybrid propulsion, combining flight scenario modeling, coefficient optimization, and intelligent allocation, and linking with the flight scenario thermal management mechanism, the power allocation coefficients of the battery and range extender are optimized based on reinforcement learning. This accurately adapts to the dual-source dynamic characteristics and flight load requirements, achieving efficient operation of the range extender and reasonable charging and discharging of the battery. It also incorporates the range extender differential response model and power smoothing constraints.
Significantly improves the energy efficiency of flight equipment, extends the flight range, simplifies the model structure to reduce computational load, avoids power fluctuations and overheating, extends the lifespan of core components, and enhances flight safety and reliability.
Smart Images

Figure CN121671871B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to control technology for new energy flight equipment, specifically a dual-source dynamic coupling energy management method for range-extended hybrid propulsion. Background Technology
[0002] As the global aviation industry accelerates its transformation towards low-carbon and high-efficiency operations, new energy flight equipment has become the core carrier for green upgrades in low-altitude flight. Among these, pure electric flight equipment has become the mainstream research and development direction due to its zero-emission advantage. However, limited by battery energy density and the lack of in-flight refueling infrastructure, it still faces bottlenecks such as range anxiety and insufficient adaptability in long-range, high-load scenarios (such as climb and headwind flight). Range-extended electric flight equipment, using an internal combustion engine to drive a range extender to refuel the battery, overcomes the range limitations of pure electric flight equipment and reduces reliance on large-capacity batteries, achieving an optimal balance between performance and cost. This has become a key technological route for the green transformation of low-altitude flight. Range-extended flight equipment adopts a dual-source power supply structure of battery and range extender, requiring an energy management system to switch operating modes to adapt to varying flight load demands. The core function of the energy management system is to coordinate the power distribution between the two sources, ensuring a balance between energy efficiency, battery life, and flight stability. However, batteries and range extenders have inherent differences in response speed and power output characteristics. Furthermore, the load fluctuations during flight (such as power demand during takeoff which can be 2-3 times that of cruise) are significant, leading to problems such as mode switching timing conflicts, insufficient power supply continuity, and range extender start-stop disturbances in practical applications of dual-source systems. Especially under dynamic flight conditions such as frequent takeoffs and landings, variable flight attitudes, and airflow disturbances, achieving response coordination and path scheduling based on dual-source characteristics and flight load requirements has become a core bottleneck restricting the performance improvement of energy management systems for range-extended flight equipment.
[0003] Currently, research on energy management strategies for hybrid propulsion flight equipment mainly focuses on optimizing energy allocation through approaches such as integrating scenario adaptation requirements, coupling thermal management systems, or combining energy prediction to improve energy efficiency or flight reliability. Chinese invention patent application number CN202411345935.3, entitled "A Method for Generating a Coupling System Strategy for Energy and Thermal Management of a Flying Car," models the fuel cell subsystem, lithium battery subsystem, and cooling subsystem of a hydrogen fuel cell hybrid flying car, constructs an energy and thermal management coupling model, and solves it using the PPO algorithm. This achieves dynamic adjustment of hydrogen fuel cell operating parameters, improving energy utilization efficiency and extending system lifespan. However, this patent focuses on the thermal management coupling optimization of the hydrogen fuel cell power system and does not design a dynamic coordination mechanism for power allocation in the dual-source power structure of the range extender and battery. It also fails to consider the thermal inertial response hysteresis characteristics of the range extender under flight scenarios, resulting in insufficient dual-source power matching accuracy under high load fluctuation conditions. Furthermore, it does not specifically suppress start-stop disturbances of the range extender, limiting the adaptability for ensuring the lifespan of core components. Chinese invention patent application number CN202411198543.9, entitled "An Energy Management Method for Solar-Hydrogen Hybrid Flying Car Based on Solar Energy Distribution Map," predicts solar radiation intensity using a BP neural network, plans flight trajectories using an improved Astar algorithm, establishes a coupled trajectory and energy management model, and solves it using a sequential quadratic programming algorithm. This achieves a reasonable allocation of power from each power source, reduces hydrogen energy consumption, and balances the state of charge of the power battery. However, this patent relies on the preconditions of a solar energy distribution map and trajectory planning, and its operating condition adaptation depends on preset energy prediction data. It suffers from dynamic response delays and cannot cope with sudden changes in operating conditions during flight (such as sudden airflow or emergency attitude adjustments). Furthermore, it does not construct a simplified dual-source coupling model adapted to the characteristics of the range extender. In engineering implementation, it is necessary to process both trajectory and energy data simultaneously, resulting in a large computational load. Chinese invention patent application number CN202511292004.6, entitled "An Energy Management Control System and Method for a Fuel Cell Electric Aircraft," utilizes a multi-source energy supply structure consisting of fuel cell components, lithium batteries, and solar panels to achieve on-demand energy distribution and automatic switching of operating modes, thereby improving energy supply flexibility and energy utilization efficiency and extending aircraft endurance. However, this patent focuses on the switching of multi-source energy supply modes and basic energy distribution, without dynamically optimizing the power distribution coefficient between the range extender and the battery in conjunction with the real-time power demand of the flight scenario. It also lacks a dynamic response constraint and thermal management linkage mechanism for the range extender, which can easily lead to range extender power fluctuations and battery overheating risks in high-power flight scenarios. Furthermore, it lacks sufficient control over the smoothness of power distribution under transient conditions, affecting flight stability.
[0004] Compared with traditional energy management strategies, existing technologies have not achieved optimal energy efficiency under complex flight conditions and have many drawbacks: dual-source power distribution does not take into account differences in flight load and dynamic response, resulting in mode switching conflicts, discontinuous power supply, and start-stop disturbances, affecting flight stability and control experience; the model design is not adapted to flight dynamics characteristics, resulting in high engineering implementation load and reduced control accuracy during sudden changes in operating conditions; and the lack of a flight scenario thermal management mechanism makes core components prone to overheating or power fluctuations, shortening service life and increasing maintenance costs. Summary of the Invention
[0005] To address the shortcomings of existing technologies and solve the problems of low energy efficiency and insufficient adaptability of energy management strategies when dealing with complex flight scenarios, this invention aims to provide a dual-source dynamic coupling energy management method for range-extended hybrid propulsion systems. This method focuses on achieving optimal energy efficiency in flight scenarios. Through a comprehensive design process of "flight scenario modeling - coefficient optimization - intelligent allocation," it links the flight scenario thermal management mechanism and optimizes the power allocation coefficients between the battery and the range extender based on reinforcement learning. This precisely adapts to the dynamic characteristics of the dual sources and the flight load requirements, ensuring that the range extender always operates within its high-efficiency range and that battery charging and discharging are more rational. Ultimately, this significantly improves the energy efficiency of flight equipment while also considering flight adaptability, system stability, and the lifespan of core components.
[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0007] A method for dynamic coupling energy management of dual-source hybrid propulsion systems with extended range, the method comprising:
[0008] 1) First, a global physical model of the dual-source power system and flight scenario is constructed to build a longitudinal motion model of flight to quantify the mapping relationship between flight drag and propulsion system torque. At the same time, a dual-source dynamic coupling model is established. Among them, the battery model is used to characterize the dynamic changes of SOC (State of Charge) and transient response characteristics such as thermal coupling internal resistance. The range extender model is used to quantify thermal inertial hysteresis characteristics and construct a first-order inertial response relationship, and is adapted to the characteristic parameters of high-power flight scenarios. By deploying sensors at key locations in the system, real-time state signals such as flight speed, acceleration, flight attitude angle, flight altitude, atmospheric density, battery temperature and SOC are collected. Based on the collected signals, the real-time power deviation is calculated to quantify the asymmetry of the dual-source response.
[0009] 2) Based on the modeling results and sensor acquisition signals obtained in step 1), a reinforcement learning physical constraint reward function is first designed. This reward function includes a flight scenario efficiency reward term, a range extender start-stop suppression penalty term, and a battery thermal degradation penalty term. The efficiency reward term is adapted to the flight load characteristics, the start-stop suppression penalty term is related to the range extender start-stop event count, and the thermal degradation penalty term is related to the battery internal resistance model and the temperature characteristics of the flight scenario. Then, the QMPSO (Quantum-behaved Multi-objective Particle Swarm Optimization) algorithm is used to optimize and train the coefficients of the reward function, with fuel consumption, power deviation, and range extender start-stop count as the optimization objective functions, and the optimized reward function coefficients are output. Finally, the measured power demand is compared with the model predicted power demand, and the power matching error (RMSE) is calculated to verify the coupling accuracy of the dual-source dynamic coupling model, providing a model foundation for subsequent reinforcement learning optimization.
[0010] 3) Based on the optimized reward function obtained in step 2) and the validated dual-source dynamic coupling model, the CDO-DDPG (Coupled Dynamic Optimization - Deep Deterministic Policy Gradient) reinforcement learning algorithm is used to optimize the power allocation coefficient. The state space of this reinforcement learning algorithm is set as SOC, flight speed, power demand, battery temperature, and power deviation, and the action space is set as battery target power and range extender target power. The range extender differential response model and power smoothing constraints are embedded in the algorithm network to achieve smooth control of the power allocation process, avoid power fluctuations caused by sudden changes in flight load, and finally output the optimal power allocation coefficient between the battery and the range extender.
[0011] 4) Based on the optimal power allocation coefficient output by reinforcement learning in step 3), the total power is decomposed collaboratively in combination with the real-time power demand of the system to determine the target power of the battery and the target power of the range extender and execute control commands; a flight scenario thermal management linkage mechanism is introduced to dynamically adjust the maximum allowable power of the battery and the range extender based on the dual-source model according to the battery temperature and range extender temperature collected by the sensors, so as to avoid the core components from overheating during high-power flight and improve the system control stability and flight safety.
[0012] 5) Based on the stable control effect achieved in step 4), a comparative verification experiment was designed to evaluate the performance of the dual-source power coordinated control strategy of the present invention. A traditional rule-based power allocation strategy was selected as the comparison strategy. A whole model of the flight equipment was built in the aerospace simulation software, and a unified test flight condition was set. Under the same test condition, the core performance indicators of the strategy of the present invention and the comparison strategy were collected respectively, including fuel consumption rate, battery SOC fluctuation range, range extender start-stop times and power matching error (RMSE). Through multi-dimensional comparative analysis of each performance indicator, the superiority of the strategy of the present invention in terms of flight economy, operational stability and system life guarantee was verified.
[0013] Further, step 1) specifically includes:
[0014] 1.1) Construct a dynamic model of the flight scenario, define the drag that needs to be overcome during flight, and establish the dynamic balance equations and the mapping relationship between the torque and driving force of the propulsion system:
[0015] Aerodynamic drag:
[0016] Gravitational component drag:
[0017] Climbing resistance:
[0018] Acceleration resistance:
[0019] In the formula, The total mass of the flight equipment, It is the acceleration due to gravity. The climbing drag coefficient, Flight trajectory tilt angle, The air density corresponding to the flight altitude. This is the aerodynamic drag coefficient. For wing reference area, For the instantaneous speed of the flight equipment, This is the rotational mass conversion factor. To accelerate the flight equipment.
[0020] Establish the flight dynamics equilibrium equations:
[0021]
[0022] In the formula, To drive the overall system performance.
[0023] Based on the mechanical relationship between torque and force:
[0024]
[0025] In the formula, To propel the system to output torque, For the mechanical efficiency of the transmission system, The equivalent radius of the propeller.
[0026] The torque demand model of the propulsion system can be obtained by combining the following steps:
[0027]
[0028] 1.2) Construct a dual-source coupling model of battery and range extender to characterize their dynamic characteristics and coupling relationship, adapting to the load requirements of flight scenarios:
[0029]
[0030] In the formula, Open circuit voltage, For reference open-circuit voltage, For temperature coefficient, For battery temperature, For reference temperature, This is the equivalent internal resistance of the battery. For reference internal resistance, The internal resistance temperature index. The deviation coefficient, This refers to the number of charge-discharge cycles. This is the cyclic influence coefficient. For rated capacity, For the current capacity, The self-discharge coefficient, This is the calendar aging time. Calendar aging factor, The cyclic aging coefficient, For battery output power, This refers to the battery terminal voltage. This represents the battery terminal current.
[0031] Considering the thermal inertial hysteresis characteristics, based on a first-order inertial system, the actual output power of the dynamic response range extender satisfies the target power as follows:
[0032]
[0033] In the formula, The time constant of the range extender system (characterizing the power response lag due to thermal inertia, and dynamically adjusted to adapt to flight scenarios). This represents the actual output power of the range extender. For the target power of the range extender, For flight scenario adaptation coefficients.
[0034] Establish dual-source power coupling:
[0035]
[0036] In the formula, For the total required power, This represents the actual output power of the range extender. Quantify the dual-source response asymmetry for dynamic power deviation rate.
[0037] 1.3) By collecting signals such as flight speed, acceleration, flight attitude angle, flight altitude, atmospheric density, battery temperature and SOC through sensors, and combining them with the real-time aerodynamic parameters output by the atmospheric data computer, the real-time power demand is calculated and the power deviation is solved, providing input data for subsequent optimization.
[0038] Further, step 2) specifically includes:
[0039] 2.1) Based on the model established in step 1) and the sensor-acquired signals, design a multi-objective reward function that considers energy consumption, start-stop losses, and aging effects in flight scenarios:
[0040]
[0041] In the formula, These are the weighting coefficients. For energy consumption during flight scenarios, This is for battery aging. This is a penalty item for starting and stopping the range extender.
[0042] Specifically, the energy consumption item for flight scenarios is defined as:
[0043]
[0044] In the formula, This is the fuel energy consumption weighting coefficient. For the range extender fuel mass flow rate, This is the battery energy consumption weighting coefficient. Because it is a low-calorific-value fuel, To improve the efficiency of the drive motor.
[0045] Battery aging is defined as follows:
[0046]
[0047] In the formula, The aging penalty weighting coefficient, These are the fitting coefficients. Depth of charge / discharge, This refers to the charge / discharge rate. This is the aging impact index.
[0048] The start-stop penalty item is defined as follows:
[0049]
[0050] In the formula, The start / stop penalty weighting coefficient. This refers to the fixed energy loss during a single start-stop cycle. For the start time, At the moment when the power is stable, This refers to the number of start-stop cycles of the range extender. This refers to the dynamic power loss during the start-stop transition phase.
[0051] Substituting the components, the final reward function is:
[0052]
[0053] 2.2) The above weighting coefficients are optimized using the Quantum Particle Swarm Optimization (QMPSO) algorithm. The specific process is as follows:
[0054] A multi-objective function is constructed with the objectives of minimizing fuel consumption, power deviation, and the number of range extender start-stop cycles:
[0055]
[0056] In the formula, Represents total fuel consumption. Represents the total power deviation. This represents the total number of starts and stops.
[0057] Derivation of rotation angle increment based on updating particle position using quantum rotation angle:
[0058]
[0059] In the formula, For inertial weights, As a learning factor, for Random numbers These are the individual and globally optimal rotation angles, respectively. This is a quantum perturbation term.
[0060] Based on the above iterative update, the optimal coefficients are output: .
[0061] 2.3) The root mean square error (RMSE) was used to verify the accuracy of the dual-source coupling model, and the results were compared with the measured power demand. Compared with model predictions We can obtain:
[0062]
[0063] when ( When the preset threshold is reached, the model is considered valid and can be used for subsequent optimization.
[0064] Furthermore, step 3) specifically includes:
[0065] 3.1) Based on the dual-source dynamic coupling model established in step 1) and the reward function designed in step 2), a CDO-DDPG algorithm structure with high coupling between the model and the algorithm is constructed. Integrating the dual-source dynamic characteristics and the requirements of the flight scenario, the state space is reconstructed and defined as follows:
[0066]
[0067] In the formula, For battery power, For flight speed, For total power requirements, For battery temperature, This refers to power deviation.
[0068] Define power distribution coefficients in the action space. Then the target power of the battery and the range extender are respectively:
[0069]
[0070] In the formula, For the target power of the battery, For the target power of the range extender, This is the deviation compensation coefficient.
[0071] The action space is By adjusting Achieve dynamic power allocation.
[0072] 3.2) Embed the range extender differential response model based on DDPG, derive the CDO-DDPG formula, embed the start-stop loss model into the Q function, and define the state-action value function:
[0073]
[0074] In the formula, This refers to the optimized reward function described above. As a discount factor, These are the target policy and the target Q-network, respectively. For Critic network parameters, For Actor network parameters, Let be the mathematical expectation of the state transition.
[0075] The Actor network maximizes the cumulative reward through the policy gradient, and the gradient is derived as follows:
[0076]
[0077] In the formula, D is the experience replay pool. and These are the parameters for the Critic and Actor networks, respectively. The mathematical expectation is based on the sampled state space in the experience replay pool. The gradient of the Q function with respect to the power allocation coefficients. Output the gradient of the parameters for the Actor network.
[0078] By minimizing the time-series difference error update, the error is derived. :
[0079]
[0080] loss function for:
[0081]
[0082] In the formula, This represents the sample size.
[0083] To ensure a smooth power transition of the range extender and the optimal power allocation coefficient at the final output, a power smoothing constraint is added to the strategy, based on the range extender model from step 1).
[0084]
[0085] In the formula, The maximum allowable power change rate of the range extender, ultimately outputting the optimal distribution coefficient. .
[0086] Further, step 4) specifically includes:
[0087] 4.1) Based on the optimal allocation coefficients output in step 3). Substituting the power function and combining it with the motor efficiency model, the torque command for the drive motor is derived. The relationship between motor power and torque is:
[0088]
[0089] The thrust reverser motor torque is:
[0090]
[0091] In the formula, To drive the motor speed, This refers to the motor efficiency.
[0092] 4.2) Based on battery temperature and range extender temperature Adjust the maximum allowable power in real time and derive the formula for the maximum allowable power:
[0093]
[0094] In the formula, The maximum allowable power at the reference temperature. The maximum permissible power of the range extender at the reference temperature. This refers to the battery's safe operating temperature range. For the safe operating temperature range of the range extender, The low-temperature power attenuation coefficient, This is the high-temperature power attenuation coefficient. If... or Then adjust the allocation coefficient. To ensure system stability.
[0095] Furthermore, step 5) specifically includes:
[0096] 5.1) Based on the stable control effect achieved in step 4), a traditional rule-based power distribution strategy is selected as a comparison strategy. A whole model of the flight equipment is built in the aviation simulation software, and a unified test flight condition is set. Under the same test condition, the core performance indicators of the strategy of the present invention and the comparison strategy are collected respectively, including fuel consumption rate, battery SOC fluctuation range, range extender start-stop times and power matching error (RMSE).
[0097] The formulas for the core evaluation indicators are derived as follows:
[0098]
[0099] In the formula, For fuel density, This represents the total fuel consumption during the test period. For driving mileage, This refers to fuel consumption rate.
[0100]
[0101] In the formula, , The maximum and minimum SOC values during the test period. This represents the fluctuation range of SOC.
[0102]
[0103] In the formula, This is the initial temperature of the battery. This is the battery's end temperature.
[0104]
[0105] In the formula, For the duration of the test period, This is an indicator function (it takes the value 1 if the condition is met, otherwise it takes the value 0).
[0106] Through multi-dimensional comparative analysis of the above indicators, the advantages of the strategy of the present invention in terms of flight energy efficiency, operational stability and component life guarantee are verified.
[0107] The beneficial effects of this invention are:
[0108] 1. Adapt to the characteristics of flight scenarios, establish a flight dynamics model and a dual-source power coupling dynamic model, and combine reinforcement learning algorithms to optimize the power allocation coefficient, so that the range extender operates in the high-efficiency range under all flight conditions, avoids energy waste caused by high-rate charging and discharging of batteries, significantly improves the energy utilization efficiency of flight equipment, and extends the driving range.
[0109] 2. By adopting a simplified modeling approach to eliminate redundant parameters and specifically optimize the adaptability to flight scenarios, the model structure is simplified while the control accuracy is ensured through model verification and algorithm iteration. This solves the problems of high engineering implementation difficulty and high computational load caused by existing multi-model coupled designs, and adapts to the computing power requirements of the airborne computing platform of flight equipment.
[0110] 3. Dynamic constraints are embedded in power distribution to ensure smooth transition of dual-source power and avoid energy loss and mechanical shock caused by sudden changes in flight load; the thermal management mechanism of the flight scenario is linked to keep the battery and range extender in the high-efficiency temperature range, reduce efficiency degradation caused by overheating, extend the service life of core components, reduce aviation operation and maintenance costs, and improve flight safety and reliability.
[0111] This invention relates to an energy management strategy for range-extended flight equipment in flight scenarios. It achieves coordinated power allocation between the battery and the range extender through dynamic coupling modeling and reinforcement learning, adapting to the load requirements of flight scenarios, improving system energy efficiency and extending component life. Attached Figure Description
[0112] Figure 1 This is a model diagram of the aircraft built for the simulation of this invention;
[0113] Figure 2 This is a diagram illustrating the operating conditions of the present invention.
[0114] Figure 3 This is an overall flowchart of the method of the present invention. Detailed Implementation
[0115] Figure 1 This is a model diagram of the aircraft built for simulation in this invention; such as... Figure 1As shown, the power coordination control system of the dual-source power system of this invention mainly includes a pilot model, a mode judgment and mode selection module, a power demand, motor power and target speed calculation module, and a power execution and energy storage module. The pilot model generates a power request signal for the system based on flight mission or operating condition requirements; the mode judgment and mode selection module determines the current operating mode based on the system's operating status; the power demand, motor power and target speed calculation module decomposes and calculates the system's power demand under the selected mode. The power execution and energy storage module includes a starter motor (generator), an engine, a battery model, and a drive motor. The battery model describes the battery's energy storage and output characteristics, the drive motor converts electrical energy into mechanical energy output, and the engine and starter motor (generator) provide auxiliary energy to the system when needed. The flight dynamics model and transmission module describe the relationship between the system's power output and flight motion state, thereby achieving power coordination and control of the dual-source power system during flight. Figure 2 This is a diagram illustrating the operating conditions of the present invention; such as... Figure 2 The diagram shows a schematic of the test condition curve used in this embodiment of the invention. The horizontal axis represents time, and the vertical axis represents the system power demand. This test condition includes multiple high-power demand plateau segments and low-power fluctuation segments throughout the test process. The power demand exhibits obvious step changes over time, superimposed with certain random fluctuations, to simulate typical operating conditions such as acceleration, steady-state cruise, and load changes during flight. By simulating and verifying the dual-source power system under this test condition, the control effect and adaptability of the energy management strategy of this invention under different power demand conditions can be comprehensively evaluated.
[0116] like Figure 3 As shown, a dual-source dynamic coupling energy management method for range-extended hybrid propulsion includes:
[0117] 1) First, a global physical model of the dual-source power system and flight scenario is constructed to build a longitudinal motion model of flight to quantify the mapping relationship between flight drag and propulsion system torque. At the same time, a dual-source dynamic coupling model is established. Among them, the battery model is used to characterize the dynamic changes of SOC (State of Charge) and transient response characteristics such as thermal coupling internal resistance. The range extender model is used to quantify thermal inertial hysteresis characteristics and construct a first-order inertial response relationship, and is adapted to the characteristic parameters of high-power flight scenarios. By deploying sensors at key locations in the system, real-time state signals such as flight speed, acceleration, flight attitude angle, flight altitude, atmospheric density, battery temperature and SOC are collected. Based on the collected signals, the real-time power deviation is calculated to quantify the asymmetry of the dual-source response.
[0118] Specifically, it includes:
[0119] 1.1) Construct a dynamic model of the flight scenario, define the drag that needs to be overcome during flight, and establish the dynamic balance equations and the mapping relationship between the torque and driving force of the propulsion system:
[0120] Aerodynamic drag:
[0121] Gravitational component drag:
[0122] Climbing resistance:
[0123] Acceleration resistance:
[0124] In the formula, The total mass of the flight equipment, It is the acceleration due to gravity. The climbing drag coefficient, Flight trajectory tilt angle, The air density corresponding to the flight altitude. This is the aerodynamic drag coefficient. For wing reference area, For the instantaneous speed of the flight equipment, This is the rotational mass conversion factor. To accelerate the flight equipment.
[0125] Establish the flight dynamics equilibrium equations:
[0126]
[0127] In the formula, To drive the overall system performance.
[0128] Based on the mechanical relationship between torque and force:
[0129]
[0130] In the formula, To propel the system to output torque, For the mechanical efficiency of the transmission system, The equivalent radius of the propeller.
[0131] The torque demand model of the propulsion system can be obtained by combining the following steps:
[0132]
[0133] 1.2) Construct a dual-source coupling model of battery and range extender to characterize their dynamic characteristics and coupling relationship, adapting to the load requirements of flight scenarios:
[0134]
[0135] In the formula, Open circuit voltage, For reference open-circuit voltage, For temperature coefficient, For battery temperature, For reference temperature, This is the equivalent internal resistance of the battery. For reference internal resistance, The internal resistance temperature index. The deviation coefficient, This refers to the number of charge-discharge cycles. This is the cyclic influence coefficient. For rated capacity, For the current capacity, The self-discharge coefficient, This is the calendar aging time. Calendar aging factor, The cyclic aging coefficient, For battery output power, This refers to the battery terminal voltage. This represents the battery terminal current.
[0136] Considering the thermal inertial hysteresis characteristics, based on a first-order inertial system, the actual output power of the dynamic response range extender satisfies the target power:
[0137]
[0138] In the formula, The time constant of the range extender system (characterizing the power response lag due to thermal inertia, and dynamically adjusted to adapt to flight scenarios). This represents the actual output power of the range extender. For the target power of the range extender, For flight scenario adaptation coefficients.
[0139] Establish dual-source power coupling:
[0140]
[0141] In the formula, For the total required power, This represents the actual output power of the range extender. Quantify the dual-source response asymmetry for dynamic power deviation rate.
[0142] 1.3) By collecting signals such as flight speed, acceleration, flight attitude angle, flight altitude, atmospheric density, battery temperature and SOC through sensors, and combining them with the real-time aerodynamic parameters output by the atmospheric data computer, the real-time power demand is calculated and the power deviation is solved, providing input data for subsequent optimization.
[0143] 2) Based on the modeling results and sensor acquisition signals obtained in step 1), a reinforcement learning physical constraint reward function is first designed. This reward function includes a flight scenario efficiency reward term, a range extender start-stop suppression penalty term, and a battery thermal degradation penalty term. The efficiency reward term is adapted to the flight load characteristics, the start-stop suppression penalty term is related to the range extender start-stop event count, and the thermal degradation penalty term is related to the battery internal resistance model and the temperature characteristics of the flight scenario. Then, the QMPSO (Quantum-behaved Multi-objective Particle Swarm Optimization) algorithm is used to optimize and train the coefficients of the reward function, with fuel consumption, power deviation, and range extender start-stop count as the optimization objective functions, and the optimized reward function coefficients are output. Finally, the measured power demand is compared with the model predicted power demand, and the power matching error (RMSE) is calculated to verify the coupling accuracy of the dual-source dynamic coupling model, providing a model foundation for subsequent reinforcement learning optimization.
[0144] Specifically, it includes:
[0145] 2.1) Based on the model established in step 1) and the sensor-acquired signals, design a multi-objective reward function that considers energy consumption, start-stop losses, and aging effects in flight scenarios:
[0146]
[0147] In the formula, These are the weighting coefficients. For energy consumption during flight scenarios, This is for battery aging. This is a penalty item for starting and stopping the range extender.
[0148] Specifically, the energy consumption item for flight scenarios is defined as:
[0149]
[0150] In the formula, This is the fuel energy consumption weighting coefficient. For the range extender fuel mass flow rate, This is the battery energy consumption weighting coefficient. Because it is a low-calorific-value fuel, To improve the efficiency of the drive motor.
[0151] Battery aging is defined as follows:
[0152]
[0153] In the formula, The aging penalty weighting coefficient, These are the fitting coefficients. Depth of charge / discharge, This refers to the charge / discharge rate. This is the aging impact index.
[0154] The start-stop penalty item is defined as follows:
[0155]
[0156] In the formula, The start / stop penalty weighting coefficient. This refers to the fixed energy loss during a single start-stop cycle. For the start time, At the moment when the power is stable, This refers to the number of start-stop cycles of the range extender. This refers to the dynamic power loss during the start-stop transition phase.
[0157] Substituting the components, the final reward function is:
[0158]
[0159] 2.2) The above weighting coefficients are optimized using the Quantum Particle Swarm Optimization (QMPSO) algorithm. The specific process is as follows:
[0160] A multi-objective function is constructed with the objectives of minimizing fuel consumption, power deviation, and the number of range extender start-stop cycles:
[0161]
[0162] In the formula, Represents total fuel consumption. Represents the total power deviation. This represents the total number of starts and stops.
[0163] Derivation of rotation angle increment based on updating particle position using quantum rotation angle:
[0164]
[0165] In the formula, For inertial weights, As a learning factor, for Random numbers These are the individual and globally optimal rotation angles, respectively. This is a quantum perturbation term.
[0166] Based on the above iterative update, the optimal coefficients are output: .
[0167] 2.3) The root mean square error (RMSE) was used to verify the accuracy of the dual-source coupling model, and the results were compared with the measured power demand. Compared with model predictions We can obtain:
[0168]
[0169] when ( When the preset threshold is reached, the model is considered valid and can be used for subsequent optimization.
[0170] 3) Based on the optimized reward function obtained in step 2) and the validated dual-source dynamic coupling model, the CDO-DDPG (Coupled Dynamic Optimization - Deep Deterministic Policy Gradient) reinforcement learning algorithm is used to optimize the power allocation coefficient. The state space of this reinforcement learning algorithm is set as SOC, flight speed, power demand, battery temperature, and power deviation, and the action space is set as battery target power and range extender target power. The range extender differential response model and power smoothing constraints are embedded in the algorithm network to achieve smooth control of the power allocation process, avoid power fluctuations caused by sudden changes in flight load, and finally output the optimal power allocation coefficient between the battery and the range extender.
[0171] Specifically, it includes:
[0172] 3.1) Based on the dual-source dynamic coupling model established in step 1) and the reward function designed in step 2), a CDO-DDPG algorithm structure with high coupling between the model and the algorithm is constructed. Integrating the dual-source dynamic characteristics and the requirements of the flight scenario, the state space is reconstructed and defined as follows:
[0173]
[0174] In the formula, For battery power, For flight speed, For total power requirements, For battery temperature, This refers to power deviation.
[0175] Define power distribution coefficients in the action space. Then the target power of the battery and the range extender are respectively:
[0176]
[0177] In the formula, For the target power of the battery, For the target power of the range extender, This is the deviation compensation coefficient.
[0178] The action space is By adjusting Achieve dynamic power allocation.
[0179] 3.2) Embed the range extender differential response model based on DDPG, derive the CDO-DDPG formula, embed the start-stop loss model into the Q function, and define the state-action value function:
[0180]
[0181] In the formula, This refers to the optimized reward function described above. As a discount factor, These are the target policy and the target Q-network, respectively. For Critic network parameters, For Actor network parameters, Let be the mathematical expectation of the state transition.
[0182] The Actor network maximizes the cumulative reward through the policy gradient, and the gradient is derived as follows:
[0183]
[0184] In the formula, D is the experience replay pool. and These are the parameters for the Critic and Actor networks, respectively. The mathematical expectation is based on the sampled state space in the experience replay pool. The gradient of the Q function with respect to the power allocation coefficients. Output the gradient of the parameters for the Actor network.
[0185] By minimizing the time-series difference error update, the error is derived. :
[0186]
[0187] loss function for:
[0188]
[0189] In the formula, This represents the sample size.
[0190] To ensure a smooth power transition of the range extender and the optimal power allocation coefficient at the final output, a power smoothing constraint is added to the strategy, based on the range extender model from step 1).
[0191]
[0192] In the formula, The maximum allowable power change rate of the range extender, ultimately outputting the optimal distribution coefficient. .
[0193] 4) Based on the optimal power allocation coefficient output by reinforcement learning in step 3), the total power is decomposed collaboratively in combination with the real-time power demand of the system to determine the target power of the battery and the target power of the range extender and execute control commands; a flight scenario thermal management linkage mechanism is introduced to dynamically adjust the maximum allowable power of the battery and the range extender based on the dual-source model according to the battery temperature and range extender temperature collected by the sensors, so as to avoid the core components from overheating during high-power flight and improve the system control stability and flight safety.
[0194] Specifically, it includes:
[0195] 4.1) Based on the optimal allocation coefficients output in step 3). Substituting the power function and combining it with the motor efficiency model, the torque command for the drive motor is derived. The relationship between motor power and torque is:
[0196]
[0197] The torque of the reverse thrust motor is:
[0198]
[0199] In the formula, To drive the motor speed, For motor efficiency.
[0200] 4.2) Based on battery temperature and range extender temperature Adjust the maximum allowable power in real time and derive the formula for the maximum allowable power:
[0201]
[0202] In the formula, The maximum allowable power at the reference temperature. The maximum permissible power of the range extender at the reference temperature. For the battery's safe operating temperature range, For the safe operating temperature range of the range extender, The low-temperature power attenuation coefficient, This is the high-temperature power attenuation coefficient. If... or Then adjust the allocation coefficient. To ensure system stability.
[0203] 5) Based on the stable control effect achieved in step 4), a comparative verification experiment was designed to evaluate the performance of the dual-source power coordinated control strategy of the present invention. A traditional rule-based power allocation strategy was selected as the comparison strategy. A whole model of the flight equipment was built in the aerospace simulation software, and a unified test flight condition was set. Under the same test condition, the core performance indicators of the strategy of the present invention and the comparison strategy were collected respectively, including fuel consumption rate, battery SOC fluctuation range, range extender start-stop times and power matching error (RMSE). Through multi-dimensional comparative analysis of each performance indicator, the superiority of the strategy of the present invention in terms of flight economy, operational stability and system life guarantee was verified.
[0204] Specifically, it includes:
[0205] 5.1) Based on the stable control effect achieved in step 4), a traditional regular power distribution strategy is selected as a comparison strategy. A whole model of the flight equipment is built in the aviation simulation software, and a unified test flight condition is set. Under the same test condition, the core performance indicators of the strategy of the present invention and the comparison strategy are collected respectively, including fuel consumption rate, battery SOC fluctuation range, range extender start-stop times and power matching error (RMSE).
[0206] The formulas for the core evaluation indicators are derived as follows:
[0207]
[0208] In the formula, For fuel density, This represents the total fuel consumption during the test period. For driving mileage, This refers to fuel consumption rate.
[0209]
[0210] In the formula, , The maximum and minimum SOC values during the test period. This represents the fluctuation range of SOC.
[0211]
[0212] In the formula, This is the initial temperature of the battery. This is the battery's end temperature.
[0213]
[0214] In the formula, For the duration of the test period, This is an indicator function (it takes the value 1 if the condition is met, otherwise it takes the value 0).
[0215] Through multi-dimensional comparative analysis of the above indicators, the advantages of the strategy of the present invention in terms of flight energy efficiency, operational stability and component life guarantee are verified.
[0216] This invention relates to an energy management strategy for range-extended flight equipment in flight scenarios. It achieves coordinated power allocation between the battery and the range extender through dynamic coupling modeling and reinforcement learning, adapting to the load requirements of flight scenarios, improving system energy efficiency and extending component life.
Claims
1. A method for energy management of a range-extended hybrid propulsion dual-source dynamic coupling, characterized in that, The method comprises the following steps: 1) The global physical modeling of the dual-source power system and the flight scene is performed, the flight longitudinal motion model is constructed to quantify the mapping relationship between the flight resistance and the propulsion system torque, and the dual-source dynamic coupling model is established; wherein the battery model is used to represent the SOC dynamic change and the thermal coupling internal resistance transient response characteristic, the range extender model is used to quantify the thermal inertia hysteresis characteristic and construct a first-order inertia response relationship, and the characteristic parameters suitable for the high-power flight scene; through the deployment of sensors, the flight speed, acceleration, flight attitude angle, flight height, atmospheric density, battery temperature and SOC state signals are collected in real time, and the real-time power deviation is calculated based on the collected signals to quantify the asymmetry of the dual-source response; 2) Based on the modeling results and sensor acquisition signals obtained in step 1), first, a reinforcement learning physical constraint reward function is designed, which includes a flight scene efficiency reward item, a range extender start-stop suppression penalty item and a battery thermal degradation penalty item, wherein the efficiency reward item is adapted to the flight load characteristic, the start-stop suppression penalty item is associated with the range extender start-stop event count, and the thermal degradation penalty item is associated with the battery internal resistance model and the flight scene temperature characteristic; then, the QMPSO algorithm is used to optimize and train the coefficients of the reward function, taking the fuel consumption, power deviation and range extender start-stop times as the optimization objective function, and outputting the optimized reward function coefficients; finally, the measured power demand is compared with the model predicted power demand, the power matching error is calculated, and the coupling accuracy of the dual-source dynamic coupling model is verified, providing a model basis for subsequent reinforcement learning optimization; 3) Based on the optimized reward function obtained in step 2) and the verified dual-source dynamic coupling model, the CDO-DDPG reinforcement learning algorithm is used to optimize the power distribution coefficient; the state space of the reinforcement learning algorithm is set as SOC, flight speed, power demand, battery temperature and power deviation, and the action space is set as the battery target power and the range extender target power; the range extender differential response model and the power smoothing constraint are embedded in the algorithm network to realize the smoothness control of the power distribution process, avoid the power fluctuation caused by the flight load mutation, and finally output the optimal power distribution coefficient of the battery and the range extender; 4) Based on the optimal power distribution coefficient output by the reinforcement learning in step 3), the total power is cooperatively decomposed in combination with the real-time power demand of the system to determine the battery target power and the range extender target power and execute the control instructions; the flight scene thermal management linkage mechanism is introduced, the maximum allowable power of the battery and the range extender is dynamically adjusted based on the dual-source model according to the battery temperature and the range extender temperature collected by the sensor, and the over-temperature operation phenomenon of the core components during high-power flight is avoided. 5) Based on the stable control effect realized in step 4), a comparative verification experiment is designed to evaluate the performance of the dual-source power collaborative control strategy; the traditional rule-based power distribution strategy is selected as the comparison strategy, a flight equipment whole machine model is built in the aviation simulation software, and a unified test flight working condition is set; under the same test working condition, the core performance indicators of the dual-source power collaborative control strategy and the comparison strategy are collected, including fuel consumption rate, battery SOC fluctuation amplitude, range extender start-stop times and power matching error; through multi-dimensional comparative analysis of each performance indicator, the superiority of the dual-source power collaborative control strategy in flight economy, operation stability and system life guarantee is verified.
2. The extended-range hybrid propulsion dual-source dynamic coupling energy management method of claim 1, wherein, The step 1) specifically comprises: 1.1) Constructing a flight scene dynamics model, defining the resistance to be overcome in the flight process, establishing a dynamics balance equation and a mapping relationship between the propulsion system torque and the driving force: Aerodynamic drag: ; Gravity component resistance: ; Climbing resistance: ; acceleration resistance: ; wherein, is the total mass of the flying device, is the acceleration of gravity, is the climb drag coefficient, is the flight trajectory inclination angle, is the air density corresponding to the flight altitude, is the aerodynamic drag coefficient, is the wing reference area, is the instantaneous speed of the flying device, is the rotational mass conversion coefficient, is the acceleration of the flying device; Establishing a flight dynamics balance equation: ; In the formula, Ptot is the total drive power of the propulsion system; From the mechanics relationship between torque and force: ; wherein Tout is the output torque of the propulsion system, η is the mechanical efficiency of the transmission system, R is the equivalent radius of the propeller. The propulsion system torque demand model is obtained by combining: ; 1.2) Constructing a battery and range extender dual-source coupling model to represent the dynamic characteristics and coupling relationship of the two, and to adapt to the load demand of the flight scene: ; wherein, is the open circuit voltage, is the reference open circuit voltage, is the temperature coefficient, is the battery temperature, is the reference temperature, is the battery equivalent internal resistance, is the reference internal resistance, is the internal resistance temperature exponent, is the deviation coefficient, is the number of charge and discharge cycles, is the cycle impact coefficient, is the rated capacity, is the current capacity, is the self-discharge coefficient, is the calendar aging time, is the calendar aging coefficient, is the cycle aging coefficient, is the battery output power, is the battery terminal voltage, is the battery terminal current; Considering the thermal inertia hysteresis characteristic, the actual output power and target power of the range extender are derived based on a first-order inertia system to satisfy: ; In the formula, is the time constant of the range extender system, representing the power response lag characteristic caused by thermal inertia, adapting to the dynamic adjustment of the flight scene, is the actual output power of the range extender, is the target power of the range extender, is the flight scene adaptation coefficient; The dual-source power coupling relationship is established: ; In the formula, Ptotal is the total demand power, Pout is the actual output power of the range extender, Pdyn is the dynamic power deviation rate quantifying the asymmetry of dual-source response; 1.3) Collecting flight speed, acceleration, flight attitude angle, flight height, atmospheric density, battery temperature and SOC signals through sensors, combining real-time aerodynamic parameters output by an atmospheric data computer to calculate real-time power demand and solve power deviation, providing input data for subsequent optimization.
3. The extended-range hybrid propulsion dual-source dynamic- coupling energy management method of claim 2, wherein, The step 2) specifically comprises: 2.1) Based on the model established in step 1) and the sensor collected signals, a multi-objective reward function considering flight scene energy consumption, start-stop loss and aging effect is designed: ; wherein, is a weight coefficient, is a flight scenario energy consumption term, is a battery aging term, is a range extender start-stop penalty term; The flight scene energy consumption term is defined as: ; In the formula, is the fuel energy consumption weight coefficient, is the fuel mass flow of the range extender, is the battery energy consumption weight coefficient, is the low heat value of the fuel, is the drive motor efficiency; The battery aging term is defined as: ; wherein is an aging penalty weight coefficient, is a fitting coefficient, is a charge and discharge depth, is a charge and discharge rate, is an aging influence index; The start-stop penalty term is defined as: ; In the formula, is the start-stop penalty weight coefficient, is the single start-stop fixed energy loss, is the start time, is the power stabilization time, is the number of times of starting the range extender, is the dynamic power loss in the start-stop transition stage; Substituting each component, the final reward function is: ; 2.2) The above weight coefficients are optimized by using quantum particle swarm optimization algorithm The specific process is as follows: A multi-objective function is constructed to minimize fuel consumption, power deviation and range extender start-stop times: ; wherein represents total fuel consumption, represents total power deviation, represents total start-stop number; The particle position is updated based on the quantum rotation angle, and the rotation angle increment is derived: ; wherein is an inertia weight, is a learning factor, is a random number, are the individual and global optimal rotation angles, respectively, is a quantum perturbation term; According to the iterative update of the above formula, the optimal coefficient is output: ; 2.3) The accuracy of the dual-source coupled model was verified by the root mean square error, and the measured power demand was compared with the model predicted value and the model predicted value , we get: ; When Time, is a preset threshold value, the model verification is qualified, and is used for subsequent optimization.
4. The extended-range hybrid propulsion dual-source dynamic coupling energy management method of claim 3, wherein, The step 3) specifically comprises: 3.1) Based on the dual-source dynamic coupling model established in step 1) and the reward function designed in step 2), a CDO-DDPG algorithm structure with high coupling between model and algorithm is constructed, which integrates dual-source dynamic characteristics and flight scene demand, and reconstructs the state space, defining: ; wherein, is the battery power, is the flight speed, is the total power demand, is the battery temperature, is the power deviation; In the action space define the power allocation coefficient , then the battery and the range extender target power are respectively: ; In the formula, is the target power of the battery, is the target power of the range extender, is the deviation compensation coefficient; The action space is , by adjusting to achieve power dynamic allocation; 3.2) Embedding the range extender differential response model in DDPG, deriving the CDO-DDPG formula, embedding the start-stop loss model in the Q function, and defining the state-action value function: ; wherein is the reward function optimized above, is the discount factor, are the target policy and the target Q network, respectively, are the Critic network parameters, are the Actor network parameters, is the mathematical expectation of the state transition; The actor network maximizes the cumulative reward through policy gradient, and the gradient is derived: ; where D is an experience replay pool, is a mathematical expectation based on the sampled state space in the experience replay pool, is a gradient of the Q function with respect to the power allocation coefficient, is a gradient of the Actor network output with respect to the parameter; updating the error by minimizing a timing difference error : ; Loss function is: ; In the formula, is the sample size; To ensure smooth transition of the range extender power and finally output the optimal distribution coefficient, the power smoothing constraint is added to the strategy combined with the range extender model in step 1): ; In the formula, Pmax is the maximum allowable power variation of the range extender, and fopt is the optimal final output distribution coefficient. .
5. The extended-range hybrid propulsion dual-source dynamic coupling energy management method of claim 4, wherein, The step 4) specifically comprises: 4.1) Based on the optimal distribution coefficient output in step 3) Substitute the power function, combined with the motor efficiency model to derive the drive motor torque instruction, the relationship between motor power and torque is: ; The backstepping motor torque is: ; wherein is the motor speed, is the motor efficiency; 4.2) Based on battery temperature and range extender temperature Adjust the maximum allowed power in real time, derive the maximum allowed power formula: ; In the formula, Pmax is the maximum allowable power at the reference temperature, Pmax is the maximum allowable power of the range extender at the reference temperature, Tmax is the safe operating temperature range of the battery, Tmax is the safe operating temperature range of the range extender, is the low-temperature power attenuation coefficient, is the high-temperature power attenuation coefficient; if or , the distribution coefficient is adjusted to ensure system stability.
6. The extended-range hybrid propulsion dual-source dynamic coupling energy management method of claim 5, wherein, The step 5) specifically comprises: 5.1) Based on the stable control effect realized in step 4), the traditional rule-based power distribution strategy is selected as the comparison strategy, and the flight equipment whole machine model is built in the aviation simulation software, and a unified test flight condition is set. Under the same test condition, the core performance indicators of the dual-source power collaborative control strategy and the comparison strategy are collected respectively, including fuel consumption rate, battery SOC fluctuation amplitude, range extender start-stop times and power matching error; The formula derivation of the core evaluation index is as follows: ; wherein is the fuel density, is the total fuel consumption over the test period, is the driving range, is the fuel consumption rate; ; In the formula, , is the maximum and minimum value of SOC in the test cycle, is the amplitude of SOC fluctuation; ; wherein T0 is the initial battery temperature, Tf is the final battery temperature; ; wherein is the length of the test period, is an indicator function that takes the value 1 if the condition is met and 0 otherwise.
Citation Information
Patent Citations
A method for generating a flying car energy and thermal management coupling system strategy
CN119142215B
Solar energy and hydrogen energy hybrid power hovercar energy management method based on solar energy distribution map
CN119272954A
Energy management control system and method for fuel cell electric aircraft
CN120810548A
Multi-power-source hybrid power system and energy management method thereof
CN112069600A
Unmanned aerial vehicle hybrid power energy management method based on ECMS and PPO algorithms
CN117313311A