A Hybrid Unmanned Aerial Vehicle Energy Management System and Method Based on Reinforcement Learning

CN122570935APending Publication Date: 2026-08-14黄德靖
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-15
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

现有能量管理策略中,基于规则或离线优化的方法难以应对复杂多变的飞行工况,缺乏在线自适应能力;而基于强化学习的方法虽具有自学习优化潜力,但其固有的探索特性可能产生越限危险动作,且网络推理耗时存在不确定性,在安全攸关的飞行场景中,面临求解超时导致控制中断的致命隐患

Benefits of technology

[0114] By combining CEEMDAN decomposition with TCN-LSTM time series modeling, the real-time perception and adaptive capabilities of the energy management strategy are significantly improved. CEEMDAN decomposes the drastically fluctuating demand power into intrinsic mode functions of different scales, enabling the network to capture short-term changes and long-term trends in the load, thereby obtaining a high-quality state representation. On this basis, the SAC algorithm continuously explores online with the help of the maximum entropy mechanism, and can continuously optimize power allocation in flight without relying on an accurate system model, so that the strategy can better adapt to battery aging, drastic environmental changes and diverse flight mission profiles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122570935A_ABST
    Figure CN122570935A_ABST
Patent Text Reader

Abstract

This invention discloses a hybrid unmanned aerial vehicle (UAV) energy management system and method based on reinforcement learning, relating to the UAV field. The system includes a battery system, a fuel power system, a state feature extraction module, a signal decomposition module, an intelligent decision-making module, a safety filter, an experience replay pool, and a hybrid power controller. By combining CEEMDAN decomposition with TCN-LSTM time-series modeling, the real-time perception and adaptive capabilities of the energy management strategy are significantly improved. CEEMDAN decomposes drastically fluctuating power demand into intrinsic mode functions of different scales, enabling the network to capture short-term changes and long-term trends in load, thereby obtaining high-quality state representations. Based on this, the SAC algorithm continuously explores online using the maximum entropy mechanism, continuously optimizing power allocation during flight without relying on an accurate system model, allowing the strategy to better adapt to battery aging, drastic environmental changes, and diverse flight mission profiles.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of unmanned aerial vehicles (UAVs), and specifically to a hybrid UAV energy management system and method based on reinforcement learning. Background Technology

[0002] As the application of drones in logistics, inspection, and monitoring continues to deepen, the requirements for long endurance, high power output, and flight safety are becoming increasingly stringent. Hybrid-powered drones, due to their ability to combine the high power density of batteries with the high energy density of prime movers (such as internal combustion engines and fuel cells), have become an important direction for meeting these requirements. The energy management system of this type of drone needs to dynamically decide the power distribution between the battery and fuel power systems based on real-time changes in flight status and load demands, in order to achieve multiple objectives such as improving energy efficiency, extending endurance, and ensuring the safety margin of the power system.

[0003] However, in actual operation, the power demand of drones is affected by factors such as weather conditions, flight attitude, and sudden load changes, exhibiting severe high-frequency non-stationary fluctuations. Meanwhile, key state variables such as battery state of charge (SOC), battery temperature, prime mover speed, and power change rate are all subject to strict physical safety boundaries. Inappropriate control strategies can lead to accidents such as battery overheating, deep over-discharge, accelerated battery lifespan degradation, or even in-flight power interruption. Existing energy management strategies, such as rule-based or offline optimization methods, struggle to cope with complex and variable flight conditions and lack online adaptive capabilities. While reinforcement learning-based methods possess self-learning optimization potential, their inherent exploratory nature may lead to dangerous over-limit actions, and the network inference time is uncertain, posing a fatal risk of timeout and control interruption in safety-critical flight scenarios.

[0004] Based on this, the present invention provides a hybrid unmanned aerial vehicle (UAV) energy management system and method based on reinforcement learning. Summary of the Invention

[0005] To address the problems mentioned in the background, this invention provides a hybrid unmanned aerial vehicle (UAV) energy management system and method based on reinforcement learning. By combining CEEMDAN decomposition with TCN-LSTM temporal modeling, the real-time perception and adaptive capabilities of the energy management strategy are significantly improved. CEEMDAN decomposes drastically fluctuating demand power into intrinsic mode functions of different scales, enabling the network to capture short-term changes and long-term trends in load, thereby obtaining high-quality state representation. On this basis, the SAC algorithm continuously explores online with the help of the maximum entropy mechanism, and can continuously optimize power allocation during flight without relying on an accurate system model, so that the strategy can better adapt to battery aging, drastic environmental changes, and diverse flight mission profiles.

[0006] To achieve the above objectives, the present invention provides the following technical solution:

[0007] A hybrid unmanned aerial vehicle (UAV) energy management system based on reinforcement learning includes a battery system, a fuel power system, a state feature extraction module, a signal decomposition module, an intelligent decision-making module, a safety filter, an experience playback pool, and a hybrid power controller.

[0008] The fuel-powered system includes a prime mover, a generator, and a rectifier. The output shaft of the prime mover is coaxially and rigidly connected to the input shaft of the generator. The generator is electrically connected to the battery system and the motor drive controller via the rectifier. The prime mover includes, but is not limited to, an engine and a fuel cell. The state feature extraction module is used to collect battery SOC, battery temperature, prime mover speed, real-time power demand, and flight speed, and to calculate power differential features. The state sequence input is constructed; the signal decomposition module is used to perform adaptive noise complete set empirical mode decomposition on the real-time demand power to obtain multiple intrinsic mode functions and residual terms, and reconstruct the power feature vector.

[0009] The intelligent decision-making module includes a TCN-LSTM feature encoding module, an Actor network, a dual Critic network, and a minimum value comparator. The TCN-LSTM feature encoding module is used to extract multi-scale temporal features from the state sequence and output a high-dimensional temporal feature vector. The output of the dual Critic network is connected to the minimum value comparator, and the output of the minimum value comparator is connected to the Actor network. The input of the safety filter is connected to the Actor network, the battery system, and the fuel power system, respectively, and is used to perform constraint verification and correction on the actions output by the Actor network. The experience replay pool is used to store the state vector, the safety-corrected actions, the immediate reward, and the state vector at the next moment. The hybrid power controller is electrically connected to the state feature extraction module, the signal decomposition module, the intelligent decision-making module, the safety filter, and the experience replay pool, respectively, and is used to regulate the power distribution between the battery system and the fuel power system.

[0010] Furthermore, the state feature extraction module normalizes each dimension of the state vector, making its values ​​range from [0,1] to [-1,1], and the constructed state sequence satisfies:

[0011]

[0012]

[0013] In the formula, Let be the state vector at time t. , For the state dimension; The length of the time window; The battery is in its state of charge. For real-time power demand; It is a power differential feature; This refers to the speed of the prime mover; Battery temperature; This refers to flight speed.

[0014] Furthermore, the signal decomposition module requires real-time power... Decompose it to satisfy:

[0015]

[0016] In the formula, The raw signal of the real-time power demand of the UAV at time t. Let k be the kth order eigenmode function. The term represents the residual, and K represents the total number of intrinsic mode functions. The decomposition process is as follows:

[0017] Add to the original power signal Using sub-white noise, multiple sets of noisy signals are constructed and empirical mode decomposition is performed on each signal. The ensemble average of the first-order components obtained from each decomposition is then calculated to obtain the first-order eigenmode function. ;

[0018] Calculate the first-order residual:

[0019]

[0020] Recursive calculation of the first Order residuals:

[0021]

[0022] And based on the previous residual term and the adaptive noise component, construct the first... The intrinsic mode functions are obtained; the decomposition terminates when the number of extreme points of the residual term is less than 2; the signal decomposition module combines the obtained intrinsic mode function sequence with the residual term to form a probability feature vector:

[0023]

[0024] And reconstruct the state vector:

[0025]

[0026] To replace the original demand power scalar input to the TCN-LSTM feature encoding module;

[0027] The TCN-LSTM feature encoding module is embedded in the front end of the Actor network, forming the core feature extraction structure of the Actor network. Its data processing procedure is as follows:

[0028]

[0029] in, For Actor network mapping functions, For network parameters, The time window length; the network performs the following operations in sequence:

[0030] The input layer receives sequence input.

[0031]

[0032] The TCN part performs dilated causal convolution on the input sequence to extract temporal features. ;

[0033] LSTM part with As input, update the hidden state step by step, and output the final hidden state. ;

[0034] Fully connected layer will Mapped to preliminary actions The specific operation process is as follows:

[0035] Probability distribution parameter prediction: The fully connected layer will hide the state. Mapped to the mean of a joint Gaussian distribution With log standard deviation ,satisfy:

[0036]

[0037]

[0038] In the formula, The weight matrix and bias terms for mean mapping are shown. The weight matrix and bias term for the log-standard deviation mapping; standard deviation This ensures that it remains positive.

[0039] Reparameterized sampling: Introducing noise sampled from the standard normal distribution. Construct a differentiable action sampling process to calculate unrestricted action samples. :

[0040]

[0041] In the formula, For element-wise multiplication;

[0042] Motion limiting: Unrestricted motion samples are mapped to the [-1,1] interval using the Tanh activation function, matching the physical execution boundaries of the prime mover and battery system to obtain preliminary motion. :

[0043]

[0044] The initial action It includes two dimensions: the target output command of the prime mover and the target battery power.

[0045] The TCN-LSTM feature encoding module includes multi-layer dilated causal convolutional blocks, LSTM layers, and fully connected layers. The dilation coefficient of the multi-layer dilated causal convolutional blocks... It increases exponentially, satisfying , where i is the number of layers in the convolutional block;

[0046] TCN residual structure, satisfying:

[0047]

[0048] In the formula, The output features of the current TCN block, The output features of the previous TCN block It is the dilated causal convolution mapping function;

[0049] The update formula for the LSTM layer satisfies:

[0050] Input Gate:

[0051] Forgotten Gate:

[0052] Output gate:

[0053] Candidate memories:

[0054] Memory update:

[0055] Hidden state:

[0056] in, For TCN output feature sequences in Input at any time For LSTM hidden states, For memory units, For the sigmoid function, This is an element-wise product.

[0057] Furthermore, the security filter employs a two-layer security correction mechanism, with the control barrier function CBF linear inequality constraint as the core and physical limit constraints as a fallback, to process the initial actions output by the Actor network. The specific process is as follows:

[0058] Safety sets and control barrier function definitions: For the four safety dimensions—battery SOC, battery temperature, prime mover speed, and battery power change rate—corresponding safety sets are defined respectively. With control barrier function ,satisfy:

[0059] }

[0060] In the formula, This is the system's real-time state vector. The dimension of the state vector;

[0061] Linear inequality constraint transformation: To ensure the system always remains within the safe set, based on the system dynamics model... Construct CBF derivative constraints and transform them into standard linear inequality constraints:

[0062]

[0063] In the formula, The initial action output by the Actor network. For the constraint coefficient matrix, Both are used to constrain boundary vectors, and are calculated and generated in real time based on the current system state.

[0064] Solving the quadratic programming problem for safe actions: With the objective of minimizing the deviation between the final safe action and the initial action, a quadratic programming objective function is constructed. Combining the linear inequality constraints and the physical limit catch-all constraints mentioned above, a complete quadratic programming problem is formed:

[0065]

[0066]

[0067] In the formula, The initial action at time t is the output of the Actor network. , The physical limit constraint boundary corresponding to the prime mover and battery system;

[0068] Safety action issuance: The above problem is solved in real time using a quadratic programming solver to obtain safety actions that satisfy all constraints. The command is then sent to the prime mover controller and battery management system for execution.

[0069] If the algorithm's solution time exceeds the specified period, the security filter will directly output the security action from the previous moment or a rule-based security action.

[0070] Furthermore, the constraint buffering strength of the control barrier function is extended by K-type functions. The adjustment mechanism allows for dynamic adjustment of the stringency of safety constraints based on the operating conditions of the drone.

[0071] When the prime mover is a fuel cell, the constraints on the safety filter also include:

[0072] Fuel cell stack temperature constraints:

[0073]

[0074] Fuel cell output power change rate constraint:

[0075] In the formula, For fuel cell stack temperature, This refers to the output power of the fuel cell.

[0076] A method for a hybrid unmanned aerial vehicle (UAV) energy management system based on reinforcement learning includes the following steps:

[0077] Step 1: Collect real-time drone operating status data, including battery state of charge (SOC) and power demand. power differential prime mover speed Battery temperature and flight speed The state vector is constructed, and the required power is decomposed into intrinsic mode functions and residual terms using CEEMDAN to reconstruct the state vector. ;

[0078] Step 2: Convert the state vector The input is fed into the SAC intelligent decision-making module, and after features are extracted by the TCN-LSTM feature encoding module, the Actor network outputs the initial action. The preliminary action includes the target output command of the prime mover. and target battery power ,Right now:

[0079]

[0080] Step 3: Initial actions The input safety filter constructs linear inequality constraints based on the control barrier function, combines them with physical limit catch-all constraints, and generates safe actions through quadratic programming.

[0081]

[0082] The safety actions are then sent to the prime mover controller and battery management system for execution.

[0083] Step 4: Receive immediate rewards after performing the security action and the state at the next moment ,

[0084] empirical samples Store in the experience replay pool;

[0085] Step 5: Randomly sample a small batch of experience samples from the experience replay pool, and use a dual Critic network to evaluate the value of actions through temporal difference learning. The two Critic networks output... and The minimum value is taken as the final Q value through a comparator to suppress overestimation;

[0086] Step 6: Update the online network parameters of the dual Critic network by minimizing the Bellman residual. After the two branches of the dual Critic network output the action value, the minimum value is taken as the final Q value through the minimum value comparator to suppress value overestimation; and update the TCN-LSTM feature encoding module and Actor network parameters with the goal of maximizing expected return and policy entropy.

[0087] Step 7: Repeat steps 1 to 6 until the policy converges to obtain the trained energy management policy.

[0088] Furthermore, the immediate reward function in step 4 satisfies:

[0089]

[0090] In the formula, For system energy loss, For battery reference SOC, This refers to the amount of battery health degradation. These are the weighting coefficients.

[0091] The system energy loss Equivalent power loss of prime mover Heat loss due to battery internal resistance Composition, namely:

[0092]

[0093] Wherein, the equivalent power loss of the prime mover Calculated based on the real-time fuel consumption rate of the prime mover and the lower calorific value of the fuel:

[0094]

[0095] In the formula, The fuel consumption rate at the current moment is determined by the prime mover speed. and output torque Obtained through mapping of prime mover characteristic maps, i.e. = ( , ); It has a low calorific value as a fuel;

[0096] The heat loss due to internal resistance of the battery Calculated based on battery charging and discharging current and battery equivalent internal resistance:

[0097]

[0098] In the formula, This represents the current charging and discharging current of the battery. This represents the equivalent internal resistance at the current battery temperature and SOC.

[0099] The objective value calculation of the dual-Critic network in step 5 satisfies the Bellman equation:

[0100]

[0101] In the formula, As a discount factor, The action value output of the dual Critic network, Temperature coefficient;

[0102] The loss function of the Actor network in step 6 satisfies:

[0103]

[0104] In the formula, For temperature coefficient, For Actor network strategies, This is an experience replay pool.

[0105] Furthermore, the experience replay pool employs a priority experience replay mechanism, based on the absolute value of the timing difference error. Each empirical sample is assigned a sampling priority, and the priority is related to... Positive correlation; during sampling, samples are randomly sampled according to priority weights, with priority given to replaying samples with larger errors.

[0106] The temperature coefficient Automatic adjustment by minimizing the objective function:

[0107]

[0108] In the formula, Let be the target entropy.

[0109] Furthermore, the dual-Critic network comprises an online network and a target network. The target network updates its parameters using a soft update method, with the update formula being:

[0110]

[0111] In the formula, For the target network parameters, For online network parameters, This is the soft update coefficient, with a value range of 0 < τ < 1.

[0112] A computer-readable storage medium having a computer program stored thereon, characterized in that, when the computer program is executed by a processor, it implements the steps of a method for a hybrid unmanned aerial vehicle energy management system based on reinforcement learning.

[0113] Compared with existing technologies, the hybrid unmanned aerial vehicle (UAV) energy management system and method based on reinforcement learning provided by this invention have the following beneficial effects:

[0114] By combining CEEMDAN decomposition with TCN-LSTM time series modeling, the real-time perception and adaptive capabilities of the energy management strategy are significantly improved. CEEMDAN decomposes the drastically fluctuating demand power into intrinsic mode functions of different scales, enabling the network to capture short-term changes and long-term trends in the load, thereby obtaining a high-quality state representation. On this basis, the SAC algorithm continuously explores online with the help of the maximum entropy mechanism, and can continuously optimize power allocation in flight without relying on an accurate system model, so that the strategy can better adapt to battery aging, drastic environmental changes and diverse flight mission profiles.

[0115] This application also constructs a two-layer mechanism of active protection and redundancy fallback. The core adopts a control barrier function to transform safety constraints such as battery SOC and temperature into linear inequalities. Through quadratic programming, actions are proactively corrected to mathematically ensure that the system state does not exceed the safety set. At the same time, physical limit hard boundaries are superimposed as a fallback, and timeout protection logic is introduced. When the reinforcement learning solves the timeout, the previous moment or rule-based safety action is directly executed, which fundamentally eliminates the risk of control interruption and effectively ensures the real-time safety of flight and system stability.

[0116] By using reinforcement learning to optimize the power distribution between the prime mover and the battery, the energy secondary conversion loss is reduced. The reward function explicitly incorporates the battery health degradation amount, enabling the policy to actively learn to avoid detrimental behaviors such as high current surges, thereby delaying battery aging. In addition, the dual-critic network architecture that takes the minimum value effectively suppresses value overestimation, prompting the agent to make more conservative and stable long-term decisions, further protecting the hardware and improving overall energy efficiency.

[0117] This application features a highly modular, scalable, and robust overall architecture. The safety filter can automatically add constraints based on the type of prime mover, and the same software framework can quickly adapt to different power platforms, from engines to fuel cells, as well as various aircraft types such as multirotors and fixed-wing aircraft. Multiple alternative solutions are provided for its core components, such as state extraction, safety filtering, and reinforcement learning algorithms, greatly facilitating technology iteration and cross-platform portability. Furthermore, CEEMDAN's signal noise immunity and TCN-LSTM's stable timing modeling ensure that the system can still output smooth and reliable action commands under sensor noise and transient disturbances, enhancing system robustness. Attached Figure Description

[0118] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.

[0119] Figure 1 This is an overall architecture diagram of the hybrid unmanned aerial vehicle energy management system based on reinforcement learning in an embodiment of the present invention;

[0120] Figure 2 This is a general flowchart of the energy management method in this embodiment of the invention;

[0121] Figure 3 This is a logic diagram of the operation of the safety filter in an embodiment of the present invention;

[0122] Figure 4 This is a structural diagram of TCN-LSTM in an embodiment of the present invention;

[0123] Figure 5 This is a flowchart of the SEEMDAN signal decomposition in an embodiment of the present invention;

[0124] Figure 6 This is a schematic diagram of the SAC reinforcement learning training framework in an embodiment of the present invention. Detailed Implementation

[0125] To enable those skilled in the art to better understand the technical solution of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings.

[0126] Example 1

[0127] like Figure 1 As shown, a hybrid unmanned aerial vehicle (UAV) energy management system based on reinforcement learning includes a battery system, a fuel power system, a state feature extraction module, a CEEMDAN signal decomposition module, a SAC intelligent decision-making module, a safety filter, an experience playback pool, and a hybrid power controller.

[0128] The fuel-powered system includes a prime mover, a generator, and a rectifier. The output shaft of the prime mover is coaxially and rigidly connected to the input shaft of the generator. The generator is electrically connected to the battery system and the motor drive controller via the rectifier. The prime mover includes, but is not limited to, an engine and a fuel cell. The state feature extraction module is used to collect data on battery SOC, battery temperature, prime mover speed, real-time power demand, and flight speed, and to calculate power differential features. The state sequence input is constructed; the CEEMDAN signal decomposition module is used to perform adaptive noise complete set empirical mode decomposition on the real-time demand power, obtain multiple intrinsic mode functions and residual terms, and reconstruct the power feature vector;

[0129] The SAC intelligent decision-making module includes a TCN-LSTM feature encoding module, an Actor network, a dual Critic network, and a minimum value comparator. The TCN-LSTM feature encoding module is used to extract multi-scale temporal features from the state sequence and output a high-dimensional temporal feature vector. The output of the dual Critic network is connected to the minimum value comparator, and the output of the minimum value comparator is connected to the Actor network. The input of the safety filter is connected to the Actor network, the battery system, and the fuel power system, respectively, and is used to constrain, verify, and correct the actions output by the Actor network. The experience replay pool is used to store the state vector, the safety-corrected actions, the immediate reward, and the state vector at the next moment. The hybrid power controller is electrically connected to the state feature extraction module, the CEEMDAN signal decomposition module, the SAC intelligent decision-making module, the safety filter, and the experience replay pool, respectively, and is used to regulate the power distribution between the battery system and the fuel power system.

[0130] Specifically, to ensure safety and system stability, the CEEMDAN signal decomposition module and the SAC intelligent decision-making module operate within a set control cycle. When the algorithm's solution time exceeds this cycle, the safety filter directly outputs the safety action from the previous moment or a rule-based safety action to ensure the real-time safety of the UAV's flight.

[0131] Specifically, the state feature extraction module normalizes each dimension of the state vector, making its values ​​range from [0,1] to [-1,1]. The constructed state sequence satisfies:

[0132]

[0133]

[0134] In the formula, Let be the state vector at time t. , For the state dimension; The length of the time window; The battery is in its state of charge. For real-time power demand; It is a power differential feature; This refers to the speed of the prime mover; Battery temperature; This refers to flight speed.

[0135] Specifically, the CEEMDAN signal decomposition module requires real-time power... Decompose it to satisfy:

[0136]

[0137] In the formula, The raw signal of the real-time power demand of the UAV at time t. Let k be the kth order eigenmode function. The term represents the residual, and K represents the total number of intrinsic mode functions. The decomposition process is as follows:

[0138] Add to the original power signal Using sub-white noise, multiple sets of noisy signals are constructed and empirical mode decomposition is performed on each signal. The ensemble average of the first-order components obtained from each decomposition is then calculated to obtain the first-order eigenmode function. ;

[0139] Calculate the first-order residual:

[0140]

[0141] Recursive calculation of the first Order residuals:

[0142]

[0143] And based on the previous residual term and the adaptive noise component, construct the first... The intrinsic mode functions (IMFs) are obtained; the decomposition terminates when the number of extreme points in the residual terms is less than 2; the CEEMDAN signal decomposition module combines the obtained IMF sequence with the residual terms to form a probability feature vector.

[0144]

[0145] And reconstruct the state vector:

[0146]

[0147] To replace the original demand power scalar input to the TCN-LSTM feature encoding module;

[0148] The TCN-LSTM feature encoding module is embedded in the front end of the Actor network, forming the core feature extraction structure of the Actor network. Its data processing process is represented as follows:

[0149]

[0150] in, For Actor network mapping functions, For network parameters, The time window length; the network performs the following operations in sequence:

[0151] The input layer receives sequence input.

[0152]

[0153] The TCN part performs dilated causal convolution on the input sequence to extract temporal features. ;

[0154] LSTM part with As input, update the hidden state step by step, and output the final hidden state. ;

[0155] Fully connected layer will Mapped to preliminary actions The specific operation process is as follows:

[0156] Probability distribution parameter prediction: The fully connected layer will hide the state. Mapped to the mean of a joint Gaussian distribution With log standard deviation ,satisfy:

[0157]

[0158]

[0159] In the formula, The weight matrix and bias terms for mean mapping are shown. The weight matrix and bias term for the log-standard deviation mapping; standard deviation This ensures that it remains positive.

[0160] Reparameterized sampling: Introducing noise sampled from the standard normal distribution. Construct a differentiable action sampling process to calculate unrestricted action samples. :

[0161]

[0162] In the formula, For element-wise multiplication;

[0163] Motion limiting: Unrestricted motion samples are mapped to the [-1,1] interval using the Tanh activation function, matching the physical execution boundaries of the prime mover and battery system to obtain preliminary motion. :

[0164]

[0165] initial actions It includes two dimensions: the target output command of the prime mover and the target battery power.

[0166] The TCN-LSTM feature encoding module consists of multi-layer dilated causal convolutional blocks, LSTM layers, and fully connected layers. The dilation coefficients of the multi-layer dilated causal convolutional blocks... It increases exponentially, satisfying , where i is the number of layers in the convolutional block;

[0167] TCN residual structure, satisfying:

[0168]

[0169] In the formula, The output features of the current TCN block, The output features of the previous TCN block It is the dilated causal convolution mapping function;

[0170] The update formula for the LSTM layer satisfies:

[0171] Input Gate:

[0172] Forgotten Gate:

[0173] Output gate:

[0174] Candidate memories:

[0175] Memory update:

[0176] Hidden state:

[0177] in, For TCN output feature sequences in Input at any time For LSTM hidden states, For memory units, For the sigmoid function, This is an element-wise product.

[0178] Specifically, the security filter employs a two-layer security correction mechanism, with the CBF linear inequality constraint as the core and physical limit constraints as a fallback, to process the initial actions output by the Actor network. The specific process is as follows:

[0179] Safety sets and control barrier function definitions: For the four safety dimensions—battery SOC, battery temperature, prime mover speed, and battery power change rate—corresponding safety sets are defined respectively. With control barrier function ,satisfy:

[0180] }

[0181] In the formula, This is the system's real-time state vector. The dimension of the state vector;

[0182] Linear inequality constraint transformation: To ensure the system always remains within the safe set, based on the system dynamics model... Construct CBF derivative constraints and transform them into standard linear inequality constraints:

[0183]

[0184] In the formula, The initial action output by the Actor network. For the constraint coefficient matrix, Both are used to constrain boundary vectors, and are calculated and generated in real time based on the current system state.

[0185] Solving the quadratic programming problem for safe actions: With the objective of minimizing the deviation between the final safe action and the initial action, a quadratic programming objective function is constructed. Combining the linear inequality constraints and the physical limit catch-all constraints mentioned above, a complete quadratic programming problem is formed:

[0186]

[0187]

[0188] In the formula, The initial action at time t is the output of the Actor network. , The physical limit constraint boundary corresponding to the prime mover and battery system;

[0189] Safety action issuance: The above problem is solved in real time using a quadratic programming solver to obtain safety actions that satisfy all constraints. The command is then sent to the prime mover controller and battery management system for execution.

[0190] If the algorithm's solution time exceeds the specified period (within the control period), the security filter will directly output the security action from the previous moment or a rule-based security action.

[0191] Specifically, the constraint buffering strength of the control barrier function is achieved by extending the K-type functions. The adjustment mechanism allows for dynamic adjustment of the stringency of safety constraints based on the operating conditions of the drone.

[0192] When the prime mover is a fuel cell, the constraints on the safety filter also include:

[0193] Fuel cell stack temperature constraints:

[0194]

[0195] Fuel cell output power change rate constraint:

[0196] In the formula, For fuel cell stack temperature, This refers to the output power of the fuel cell.

[0197] It is worth noting that:

[0198] Unmanned aerial vehicles (UAVs) include, but are not limited to, multi-rotor UAVs, fixed-wing UAVs, and compound-wing UAVs; prime movers include, but are not limited to, internal combustion engines, hydrogen fuel cell systems, micro gas turbines, turboshaft engines, methanol fuel cells, solid oxide fuel cells, and solar photovoltaic arrays.

[0199] Neural networks include, but are not limited to, backpropagation neural networks, transformer frameworks, gated recurrent units (GRUs), one-dimensional convolutional neural networks (1D-CNNs), spiking neural networks (SNNs), and liquid state machines (LSMs) / reservoir computation;

[0200] The state extractor (data preprocessing module) includes, but is not limited to, low-pass filtering, VMD and wavelet transform, Kalman filter family (EKF / UKF / CKF), short-time Fourier transform (STFT), adaptive moving average filtering, principal component analysis (PCA) or autoencoder;

[0201] Safety filters include, but are not limited to, control barrier function-based, rule-based hard limiting, safety state space-based mapping, model predictive control (MPC) as a safety safeguard, reachability analysis, Lyapunov safety reinforcement learning, and action projection method;

[0202] Reinforcement learning includes, but is not limited to, dual-delay deep deterministic policy gradient (TD3), proximal policy optimization (PPO), offline reinforcement learning, deep deterministic policy gradient (DDPG), multi-agent reinforcement learning (MARL), and parameterized deep Q network (P-DQN).

[0203] Specifically, by combining CEEMDAN decomposition with TCN-LSTM time series modeling, the real-time perception and adaptive capabilities of the energy management strategy are significantly improved. CEEMDAN decomposes the drastically fluctuating demand power into intrinsic mode functions of different scales, enabling the network to capture short-term changes and long-term trends in the load, thereby obtaining a high-quality state representation. On this basis, the SAC algorithm continuously explores online with the help of the maximum entropy mechanism, and can continuously optimize power allocation in flight without relying on an accurate system model, so that the strategy can better adapt to battery aging, drastic environmental changes and diverse flight mission profiles.

[0204] This application constructs a two-layer mechanism of active protection and redundancy fallback. The core adopts a control barrier function to transform safety constraints such as battery SOC and temperature into linear inequalities. Through quadratic programming, actions are proactively corrected to mathematically ensure that the system state does not exceed the safety set. At the same time, physical limit hard boundaries are superimposed as a fallback, and timeout protection logic is introduced. When the reinforcement learning solution times out, the previous moment or rule-based safety action is directly executed, which fundamentally eliminates the risk of control interruption and effectively ensures real-time flight safety and system stability.

[0205] Example 2

[0206] like Figures 2 to 6 As shown, a method for a hybrid unmanned aerial vehicle (UAV) energy management system based on reinforcement learning includes the following steps:

[0207] Step 1: Collect real-time drone operating status data, including battery state of charge (SOC) and power demand. power differential prime mover speed Battery temperature and flight speed The state vector is constructed, and the required power is decomposed into intrinsic mode functions and residual terms using CEEMDAN, thus reconstructing the state vector. ;

[0208] Step 2: Convert the state vector The input is fed into the SAC intelligent decision-making module, and after features are extracted by the TCN-LSTM feature encoding module, the Actor network outputs the initial action. The initial actions include the target output command of the prime mover. and target battery power ,Right now:

[0209]

[0210] Step 3: Initial actions The input safety filter constructs linear inequality constraints based on the control barrier function, combines them with physical limit catch-all constraints, and generates safe actions through quadratic programming.

[0211]

[0212] The safety actions are then sent to the prime mover controller and battery management system for execution.

[0213] Step 4: Receive immediate rewards after performing the security action and the state at the next moment ,

[0214] empirical samples Store in the experience replay pool;

[0215] Step 5: Randomly sample a small batch of experience samples from the experience replay pool, and use a dual Critic network to evaluate the value of actions through temporal difference learning. The two Critic networks output... and The minimum value is taken as the final Q value through a comparator to suppress overestimation;

[0216] Step 6: Update the online network parameters of the dual Critic network by minimizing the Bellman residual. After the two branches of the dual Critic network output the action value, the minimum value is taken as the final Q value through the minimum value comparator to suppress value overestimation; and update the TCN-LSTM feature encoding module and Actor network parameters with the goal of maximizing expected return and policy entropy.

[0217] Step 7: Repeat steps 1 to 6 until the policy converges to obtain the trained energy management policy.

[0218] Specifically, the immediate reward function in step 4 satisfies:

[0219]

[0220] In the formula, For system energy loss, For battery reference SOC, This refers to the amount of battery health degradation. These are the weighting coefficients.

[0221] System energy loss Equivalent power loss of prime mover Heat loss due to battery internal resistance Composition, namely:

[0222]

[0223] Among them, the equivalent power loss of the prime mover Calculated based on the real-time fuel consumption rate of the prime mover and the lower calorific value of the fuel:

[0224]

[0225] In the formula, The fuel consumption rate at the current moment is determined by the prime mover speed. and output torque Obtained through mapping of prime mover characteristic maps, i.e. = ( , ); It has a low calorific value as a fuel;

[0226] Battery internal resistance heat loss Calculated based on battery charging and discharging current and battery equivalent internal resistance:

[0227]

[0228] In the formula, This represents the current charging and discharging current of the battery. This represents the equivalent internal resistance at the current battery temperature and SOC.

[0229] The objective value calculation of the dual-Critic network in step 5 satisfies the Bellman equation:

[0230]

[0231] In the formula, As a discount factor, The action value output of the dual Critic network, Temperature coefficient;

[0232] The loss function of the Actor network in step 6 satisfies:

[0233]

[0234] In the formula, For temperature coefficient, For Actor network strategies, This is an experience replay pool.

[0235] Specifically, the experience replay pool adopts a priority experience replay mechanism, based on the absolute value of the timing difference error. Each empirical sample is assigned a sampling priority, and the priority is related to... Positive correlation; during sampling, samples are randomly sampled according to priority weights, with priority given to replaying samples with larger errors.

[0236] Temperature coefficient Automatic adjustment by minimizing the objective function:

[0237]

[0238] In the formula, Let be the target entropy.

[0239] Specifically, the dual-critic network consists of an online network and a target network. The target network updates its parameters using a soft update method, with the update formula being:

[0240]

[0241] In the formula, For the target network parameters, For online network parameters, This is the soft update coefficient, with a value range of 0 < τ < 1.

[0242] The system utilizes reinforcement learning to optimize the power distribution between the prime mover and the battery, reducing energy loss during secondary conversion. The reward function explicitly incorporates battery health degradation, enabling the strategy to proactively learn to avoid detrimental behaviors such as high current surges, thereby delaying battery aging. Furthermore, the dual-critic network architecture, which takes the minimum value, effectively suppresses overestimation of value, prompting the agent to make more conservative and stable long-term decisions, further protecting the hardware and improving overall energy efficiency.

[0243] This application features a highly modular, scalable, and robust overall architecture. The safety filter can automatically add constraints based on the type of prime mover, and the same software framework can quickly adapt to different power platforms, from engines to fuel cells, as well as various aircraft types such as multirotors and fixed-wing aircraft. Multiple alternative solutions are provided for its core components, such as state extraction, safety filtering, and reinforcement learning algorithms, greatly facilitating technology iteration and cross-platform portability. Furthermore, CEEMDAN's signal noise immunity and TCN-LSTM's stable timing modeling ensure that the system can still output smooth and reliable action commands under sensor noise and transient disturbances, enhancing system robustness.

[0244] Example 3

[0245] A computer-readable storage medium having a computer program stored thereon, the steps of which, when executed by a processor, implement a method for a reinforcement learning-based hybrid unmanned aerial vehicle energy management system.

[0246] The foregoing has only described certain exemplary embodiments of the present invention by way of illustration. Undoubtedly, those skilled in the art can modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the foregoing drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.

Claims

1. A hybrid unmanned aerial vehicle (UAV) energy management system based on reinforcement learning, characterized in that, It includes a battery system, a fuel power system, a state feature extraction module, a signal decomposition module, an intelligent decision-making module, a safety filter, an experience playback pool, and a hybrid power controller; The fuel-powered system includes a prime mover, a generator, and a rectifier. The output shaft of the prime mover is coaxially and rigidly connected to the input shaft of the generator. The generator is electrically connected to the battery system and the motor drive controller via the rectifier. The prime mover includes, but is not limited to, an engine and a fuel cell. The state feature extraction module is used to collect battery SOC, battery temperature, prime mover speed, real-time power demand, and flight speed, and to calculate power differential features. The state sequence input is constructed; the signal decomposition module is used to perform adaptive noise complete set empirical mode decomposition on the real-time demand power to obtain multiple intrinsic mode functions and residual terms, and reconstruct the power feature vector. The intelligent decision-making module includes a TCN-LSTM feature encoding module, an Actor network, a dual Critic network, and a minimum value comparator. The TCN-LSTM feature encoding module is used to extract multi-scale temporal features from the state sequence and output a high-dimensional temporal feature vector. The output of the dual Critic network is connected to the minimum value comparator, and the output of the minimum value comparator is connected to the Actor network. The input of the safety filter is connected to the Actor network, the battery system, and the fuel power system, respectively, and is used to perform constraint verification and correction on the actions output by the Actor network. The experience replay pool is used to store the state vector, the safety-corrected actions, the immediate reward, and the state vector at the next moment. The hybrid power controller is electrically connected to the state feature extraction module, the signal decomposition module, the intelligent decision-making module, the safety filter, and the experience replay pool, respectively, and is used to regulate the power distribution between the battery system and the fuel power system.

2. The hybrid unmanned aerial vehicle (UAV) energy management system based on reinforcement learning according to claim 1, characterized in that, The state feature extraction module normalizes each dimension of the state vector, making its values ​​range from [0,1] to [-1,1]. The constructed state sequence satisfies: In the formula, Let be the state vector at time t. , For the state dimension; The length of the time window; The battery is in its state of charge. For real-time power demand; It is a power differential feature; This refers to the speed of the prime mover; Battery temperature; This refers to flight speed.

3. The hybrid unmanned aerial vehicle (UAV) energy management system based on reinforcement learning according to claim 1, characterized in that, The signal decomposition module requires real-time power. Decompose it to satisfy: In the formula, The raw signal of the real-time power demand of the UAV at time t. Let k be the kth order eigenmode function. The term represents the residual, and K represents the total number of intrinsic mode functions. The decomposition process is as follows: Adding... Using sub-white noise, multiple sets of noisy signals are constructed and empirical mode decomposition is performed on each signal. The ensemble average of the first-order components obtained from each decomposition is then calculated to obtain the first-order eigenmode function. ; Calculate the first-order residual: Recursive calculation of the first Order residuals: And based on the previous residual term and the adaptive noise component, construct the first... The intrinsic mode functions are obtained; the decomposition terminates when the number of extreme points of the residual term is less than 2; the signal decomposition module combines the obtained intrinsic mode function sequence with the residual term to form a probability feature vector: And reconstruct the state vector: The original demand power scalar input is replaced by the TCN-LSTM feature encoding module; the TCN-LSTM feature encoding module is embedded in the front end of the Actor network, forming the core feature extraction structure of the Actor network, and its data processing process is expressed as follows: in, For Actor network mapping functions, For network parameters, The time window length; the network performs the following operations sequentially: the input layer receives the sequence input. The TCN part performs dilated causal convolution on the input sequence to extract temporal features. ;LSTM part with As input, update the hidden state step by step, and output the final hidden state. The fully connected layer will Mapped to preliminary actions The specific operation process is as follows: Probability distribution parameter prediction: The fully connected layer will hide the state. Mapped to the mean of a joint Gaussian distribution With log standard deviation ,satisfy: In the formula, The weight matrix and bias terms of the mean mapping are shown. The weight matrix and bias term for the log-standard deviation mapping; standard deviation To ensure that it is always positive; reparameterized sampling: introduce noise sampled from the standard normal distribution. Construct a differentiable action sampling process to calculate unrestricted action samples. : In the formula, Element-wise multiplication; Action limiting: Unrestricted action samples are mapped to the [-1,1] interval using the Tanh activation function, matching the physical execution boundaries of the prime mover and battery system to obtain the initial action. : The initial action It includes two dimensions: the target output command of the prime mover and the target battery power; the TCN-LSTM feature encoding module includes multi-layer dilated causal convolutional blocks, LSTM layers, and fully connected layers, and the dilation coefficient of the multi-layer dilated causal convolutional blocks... It increases exponentially, satisfying Where i is the number of layers in the convolutional block; the TCN residual structure satisfies: In the formula, The output features of the current TCN block, The output features of the previous TCN block, The LSTM layer's update formula satisfies: Input gate: Forgotten Gate: Output gate: Candidate memories: Memory update: Hidden state: in, For TCN output feature sequences in Input at any time For LSTM hidden states, For memory units, For the sigmoid function, This is an element-wise product.

4. The hybrid unmanned aerial vehicle energy management system based on reinforcement learning according to claim 1, characterized in that, The security filter employs a two-layer security correction mechanism, with the control barrier function CBF linear inequality constraint as the core and physical limit constraints as a fallback, to process the initial actions output by the Actor network. The specific process is as follows: Safety sets and control barrier function definitions: For the four safety dimensions—battery SOC, battery temperature, prime mover speed, and battery power change rate—corresponding safety sets are defined respectively. With control barrier function ,satisfy: In the formula, This is the system's real-time state vector. The dimension of the state vector; transformation of linear inequality constraints: to ensure the system always remains within the safe set, based on the system dynamics model. Construct CBF derivative constraints and transform them into standard linear inequality constraints: In the formula, The initial action output by the Actor network. For the constraint coefficient matrix, To constrain the boundary vectors, both are calculated and generated in real time based on the current system state; Solving the quadratic programming safety action: With the objective of minimizing the deviation between the final output safety action and the initial action, a quadratic programming objective function is constructed. Combining the aforementioned linear inequality constraints and physical limit catch-all constraints, a complete quadratic programming problem is formed. In the formula, The initial action at time t is the output of the Actor network. , The physical limits of the prime mover and battery system are provided as catch-all constraints; safety action issuance: the above problem is solved in real time using a quadratic programming solver to obtain the safety actions that satisfy all constraints. The algorithm is then sent to the prime mover controller and battery management system for execution. If the algorithm's solution time exceeds this cycle, the safety filter will directly output the safety action from the previous moment or a rule-based safety action.

5. The hybrid unmanned aerial vehicle (UAV) energy management system based on reinforcement learning according to claim 4, characterized in that, The constraint buffering strength of the control barrier function is achieved by extending the K-type function. The adjustment mechanism allows for dynamic adjustment of the stringency of safety constraints based on the operating conditions of the drone. When the prime mover is a fuel cell, the constraints of the safety filter also include: fuel cell stack temperature constraints. Fuel cell output power change rate constraint: In the formula, For fuel cell stack temperature, This refers to the output power of the fuel cell.

6. A method for a hybrid unmanned aerial vehicle (UAV) energy management system based on reinforcement learning, characterized in that, Includes the following steps: Step 1: Collect real-time drone operating status data, including battery state of charge (SOC) and power demand. power differential prime mover speed Battery temperature and flight speed The state vector is constructed, and the required power is decomposed into intrinsic mode functions and residual terms using CEEMDAN to reconstruct the state vector. ; Step 2: Convert the state vector The input is fed into the SAC intelligent decision-making module, and after features are extracted by the TCN-LSTM feature encoding module, the Actor network outputs the initial action. The preliminary action includes the target output command of the prime mover. and target battery power ,Right now: Step 3: Initial actions The input safety filter constructs linear inequality constraints based on the control barrier function, combines them with physical limit catch-all constraints, and generates safe actions through quadratic programming. The safety actions are then sent to the prime mover controller and battery management system for execution. Step 4: Receive immediate rewards after performing the security action and the state at the next moment , to use empirical samples Store in the experience replay pool; Step 5: Randomly sample a small batch of experience samples from the experience replay pool, and use a dual Critic network to evaluate the value of actions through temporal difference learning. The two Critic networks output... and The minimum value is taken as the final Q value through a comparator to suppress overestimation; Step 6: Update the online network parameters of the dual Critic network by minimizing the Bellman residual. After the two branches of the dual Critic network output the action value, the minimum value is taken as the final Q value through the minimum value comparator to suppress value overestimation; and update the TCN-LSTM feature encoding module and Actor network parameters with the goal of maximizing expected return and policy entropy. Step 7: Repeat steps 1 to 6 until the policy converges to obtain the trained energy management policy.

7. The method for a hybrid unmanned aerial vehicle (UAV) energy management system based on reinforcement learning according to claim 6, characterized in that, The immediate reward function in step 4 satisfies: In the formula, For system energy loss, For battery reference SOC, This refers to the amount of battery health degradation. These are weighting coefficients. The system energy loss... Equivalent power loss of prime mover Heat loss due to battery internal resistance Composition, namely: Wherein, the equivalent power loss of the prime mover Calculated based on the real-time fuel consumption rate of the prime mover and the lower calorific value of the fuel: In the formula, The fuel consumption rate at the current moment is determined by the prime mover speed. and output torque Obtained through mapping of prime mover characteristic maps, i.e. = ( , ); The fuel has a low calorific value; the battery's internal resistance causes heat loss. Calculated based on battery charging and discharging current and battery equivalent internal resistance. In the formula, This represents the current charging and discharging current of the battery. Given the equivalent internal resistance at the current battery temperature and SOC; the target value calculation of the dual Critic network in step 5 satisfies the Bellman equation: In the formula, As a discount factor, For the action value output of the dual Critic network, Temperature coefficient; The loss function of the Actor network in step 6 satisfies: In the formula, For temperature coefficient, For Actor network strategies, This is an experience replay pool.

8. The method for a hybrid unmanned aerial vehicle (UAV) energy management system based on reinforcement learning according to claim 6, characterized in that, The experience replay pool adopts a priority experience replay mechanism, based on the absolute value of the timing difference error. Each empirical sample is assigned a sampling priority, and the priority is related to... Positive correlation; sampling is performed randomly according to priority weights, with priority given to replaying samples with larger errors; the temperature coefficient Automatic adjustment by minimizing the objective function: In the formula, Let be the target entropy.

9. A method for a hybrid unmanned aerial vehicle (UAV) energy management system based on reinforcement learning according to claim 6, characterized in that, The dual-Critic network comprises an online network and a target network. The target network updates its parameters using a soft update method, with the update formula being: In the formula, For the target network parameters, For online network parameters, This is the soft update coefficient, with a value range of 0 < τ < 1.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 6 to 9.