Hydrogen energy heavy truck multi-objective energy management method based on reinforcement learning
Patent Information
- Application Number
- CN202512054059.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-31
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2045-12-31
AI Technical Summary
基于规则的控制依赖预设逻辑进行功率分配,虽然实现简单但缺乏全局优化能力;等效消耗最小策略等瞬时优化方法通过等效因子进行在线优化,但其核心参数通常难以随工况动态调整;单目标强化学习方法虽能自主学习,但普遍采用线性加权方式处理多个优化目标,导致策略灵活性不足
本发明的基于强化学习的氢能重卡多目标能量管理方法,通过构建MO-SAC算法框架并集成帕累托前沿约束机制,在训练阶段实现了对经济性、动力性、安全性等多目标的协同优化与解耦评估,生成覆盖不同权衡偏好的帕累托最优策略集合;在线应用时结合实时工况识别技术,动态匹配或调整最优控制策略,并最终通过物理约束修正确保动作的可行性与安全性,从而在复杂多变的实际行驶环境中全面提升了氢能重卡能量管理系统的自适应能力、综合性能与工程可靠性。
Smart Images

Figure CN121882573B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of energy management technology for hydrogen-powered heavy-duty trucks, specifically relating to a multi-objective energy management method for hydrogen-powered heavy-duty trucks based on reinforcement learning. Background Technology
[0002] With the deepening of the global energy transition and carbon neutrality strategy, the clean and low-carbon transformation of the transportation sector has become an inevitable trend. Hydrogen fuel cell heavy-duty trucks, also known as hydrogen-powered heavy-duty trucks, are considered a key technological path to decarbonizing medium- and long-haul logistics and heavy-duty transportation due to their advantages such as zero emissions, long range, and fast refueling. The power system of hydrogen-powered heavy-duty trucks typically adopts a hybrid architecture of fuel cells and power batteries. The core of their energy management strategy lies in the real-time and dynamic allocation of output power between the fuel cell and the power battery to simultaneously meet multiple objectives such as economy, power performance, and safety. This multi-objective optimization problem becomes extremely challenging under complex and ever-changing actual driving conditions.
[0003] Current energy management strategies for hydrogen-powered heavy-duty trucks mainly revolve around three types of technologies: rule-based control, instantaneous optimization methods, and single-objective reinforcement learning. Rule-based control relies on preset logic for power allocation, which is simple to implement but lacks global optimization capabilities. Instantaneous optimization methods, such as the equivalent consumption minimization strategy, perform online optimization through equivalent factors, but their core parameters are usually difficult to dynamically adjust according to operating conditions. Single-objective reinforcement learning methods can learn autonomously, but they generally use a linear weighting approach to handle multiple optimization objectives, resulting in insufficient strategy flexibility.
[0004] These issues collectively constrain the performance of hydrogen-powered heavy-duty truck energy management strategies in real-world, complex environments, necessitating the development of an energy management method that dynamically adapts to operating conditions and enables multi-objective collaborative optimization. Summary of the Invention
[0005] To address the aforementioned problems in existing technologies, this invention provides a multi-objective energy management method for hydrogen-powered heavy-duty trucks based on reinforcement learning. The technical problem to be solved by this invention is achieved through the following technical solution: This invention provides a multi-objective energy management method for hydrogen-powered heavy-duty trucks based on reinforcement learning, comprising: Step 1: Construct the neural network model of the MO-SAC algorithm, and define the vehicle state space, action space, and multi-objective reward function; Step 2: Based on the vehicle state space, the action space, and the multi-objective reward function, the neural network model is trained through environmental interaction to obtain the trained MO-SAC model. During the training process, Bayesian optimization parameter tuning and Pareto front constraint mechanisms are introduced. After training, the Pareto optimal policy set is searched using the NSGA-II algorithm. Step 3: Construct a comprehensive state vector consistent with the vehicle state space definition based on the vehicle data. Based on the vehicle's historical power demand data, identify the vehicle's driving conditions through adaptive mode decomposition and temporal convolutional network to obtain the driving condition type. Step 4: Based on the driving condition type, match or dynamically adjust the control strategy from the Pareto optimal strategy set to determine the control strategy, and input the integrated state vector into the Actor network in the MO-SAC model corresponding to the control strategy to obtain the preliminary power allocation; Step 5: Correct the preliminary power allocation based on the fuel cell power change rate constraint and the power battery SOC boundary limit to obtain the final power allocation command.
[0006] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention provides a multi-objective energy management method for hydrogen-powered heavy-duty trucks based on reinforcement learning. By constructing the MO-SAC algorithm framework and integrating the Pareto front constraint mechanism, it achieves collaborative optimization and decoupled evaluation of multiple objectives such as economy, power, and safety during the training phase, generating a set of Pareto optimal strategies covering different trade-off preferences. In online application, it combines real-time operating condition identification technology to dynamically match or adjust the optimal control strategy, and finally ensures the feasibility and safety of the action through physical constraint correction. Thus, it comprehensively improves the adaptive capability, overall performance, and engineering reliability of the hydrogen-powered heavy-duty truck energy management system in complex and ever-changing actual driving environments.
[0007] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described in detail below with reference to the accompanying drawings. Attached Figure Description
[0008] Figure 1 This is a flowchart of a multi-objective energy management method for hydrogen-powered heavy-duty trucks based on reinforcement learning, provided in an embodiment of the present invention. Figure 2 This is a flowchart of an adaptive mode decomposition method provided in an embodiment of the present invention; Figure 3 This is a structural block diagram of a neural network model of the MO-SAC algorithm provided in an embodiment of the present invention. Detailed Implementation
[0009] To further illustrate the technical means and effects adopted by the present invention to achieve the intended purpose, the following describes in detail a multi-objective energy management method for hydrogen-powered heavy trucks based on reinforcement learning, in conjunction with the accompanying drawings and specific embodiments.
[0010] The foregoing and other technical contents, features, and effects of the present invention will be clearly presented in the following detailed description of specific embodiments in conjunction with the accompanying drawings. Through the description of the specific embodiments, a more in-depth and concrete understanding can be gained of the technical means and effects adopted by the present invention to achieve its intended purpose. However, the accompanying drawings are for reference and illustration only and are not intended to limit the technical solutions of the present invention.
[0011] This invention provides a multi-objective energy management method for hydrogen-powered heavy-duty trucks based on reinforcement learning. It primarily generates a decision model with multi-objective optimization capabilities through offline training and enables it to adapt to complex operating conditions in online applications. The method comprises two stages: offline training and online application. The offline training stage involves training the MO-SAC model in a simulation environment and using the NSGA-II algorithm to extract the Pareto optimal policy set to obtain the decision model. The online application stage involves deploying the decision model on the vehicle to achieve real-time perception, decision-making, and execution.
[0012] Please see Figure 1 , Figure 1 This is a flowchart of a multi-objective energy management method for hydrogen-powered heavy-duty trucks based on reinforcement learning, provided by an embodiment of the present invention. Figure 1 As shown in the figure, the multi-objective energy management method for hydrogen-powered heavy trucks based on reinforcement learning in this invention may include the following steps: Step 1: Construct the neural network model of the MO-SAC algorithm, and define the vehicle state space, action space, and multi-objective reward function.
[0013] Please see Figure 3 , Figure 3 This is a structural block diagram of a neural network model of the MO-SAC algorithm provided in an embodiment of the present invention, as shown below. Figure 3 As shown, in this embodiment, the neural network model includes a shared Actor network and three independent Critic networks; wherein the shared Actor network is used to generate power allocation actions based on the input integrated state vector; and the three independent Critic networks respectively evaluate the long-term value of the economy, dynamics and safety of performing the power allocation actions.
[0014] In this embodiment, the vehicle state space includes vehicle dynamic parameters, energy system state, and environmental information; wherein, the vehicle dynamic parameters include: vehicle speed ( (km / h), longitudinal acceleration ( m / s 2 ) and the required power of the drive motor ( kW); Energy system status includes: power battery SOC, fuel cell power ( kW) and hydrogen quantity ( Environmental information includes: road slope (θ, °), drag coefficient ( ); ) and rolling resistance coefficient ( ).
[0015] In this embodiment, the input shared Actor network integrated state vector is a feature vector that includes vehicle dynamic parameters, energy system state, and environmental information.
[0016] In this embodiment, the action space includes power distribution actions, fuel cell power change rate constraints, and power battery SOC boundary limits.
[0017] Among these, the rate allocation action includes the power proportion of the fuel cell.
[0018] The constraint on the rate of change of fuel cell power is: if The power ratio of the fuel cell will then be corrected to... ;in, For the power variation of fuel cells, This refers to the rated power of the drive motor. Indicates time, The power percentage of the fuel cell, For the revised version The power percentage of the fuel cell at any given time. for The power percentage of the fuel cell at any given time. The sign function (sign mapping) represents the change in power of the fuel cell when it changes in a positive direction. +1, when the power change of the fuel cell is negative. =-1, when When it is 0 =0, for The power required by the drive motor at any given time.
[0019] In this embodiment, the power variation of the fuel cell The following formula is used to calculate: .
[0020] The SOC boundary limit for the power battery is as follows: when the SOC of the power battery is below 0.2 or above 0.8, let... In other words, when the SOC of the power battery is below 0.2 or above 0.8, it is forced... Powered by fuel cell only.
[0021] In this embodiment, the multi-objective reward function is expressed as: ; in, As the weight of economic rewards, As the weight of motivational rewards, Weighting for security rewards.
[0022] in, The economic reward function is expressed as: ; In the formula, The first constant coefficient has a value of 1. This is the second constant coefficient, with a value of 0.5. For hydrogen consumption, This represents the standard deviation of the battery current.
[0023] In this embodiment, the standard deviation of the battery current , for The battery is constantly replenished with current. To replenish current to the average battery.
[0024] in, The motivational reward function is expressed as: ; In the formula, The third constant coefficient has a value of 2. To meet the power requirements of the drive motor, This refers to the actual power of the drive motor. , The power percentage of the fuel cell, This represents the maximum power of the fuel cell. The maximum power to replenish the battery's power. This indicates that when the fuel cell power is insufficient ( ), battery replenishment power, +1; when fuel cells are in excess ( The battery absorbs excess power. It is -1.
[0025] In this embodiment, , .
[0026] in, The security reward function is expressed as: ; In the formula, It is the fourth constant coefficient, with a value of 10. This is the fifth constant coefficient, with a value of 0.1. This refers to the state of charge of the power battery. For indicator functions, For the power variation of fuel cells, This means that when the battery level is below 20% (over-discharge) or above 80% (overcharge), the function value is 1 (penalty triggered); otherwise, it is 0 (no penalty).
[0027] in, The adaptive reward function is expressed as: ; In the formula, The sixth constant coefficient has a value of 1. It is the seventh constant coefficient, with a value of 0.8. Based on hydrogen consumption, , This is the longitudinal acceleration.
[0028] Step 2: Based on the vehicle state space, action space, and multi-objective reward function, the neural network model is trained through environmental interaction to obtain the trained MO-SAC model. During the training process, Bayesian optimization parameter tuning and Pareto front constraint mechanisms are introduced. After training, the Pareto optimal policy set is searched using the NSGA-II algorithm.
[0029] In this embodiment, the process of training the neural network model through environmental interaction includes: S1: Build the simulation environment; Optionally, a simulation environment for hydrogen-powered heavy-duty trucks can be established on platforms such as MATLAB. This model should accurately reflect the fuel cell system, power battery system, vehicle dynamics, and road conditions.
[0030] S2: Obtain the current state from the simulation environment that is consistent with the vehicle state space definition; S3: Input the current state into the shared Actor network of the neural network model to obtain the corresponding power allocation action; S4: Interact with the simulation environment using the power allocation action to obtain the reward and next state corresponding to the power allocation action; S5: Combine the current state, power allocation action, reward, and next state into an experience and store it in the experience replay area; S5: Repeat S2-S4 until the experience in the experience replay area reaches the preset number; S6: Extract a preset number of experiences from the experience replay area to update the network parameters of the neural network model until the neural network model reaches the preset training cutoff condition, and obtain the trained MO-SAC model.
[0031] In this embodiment, the network parameter update includes two parts: parameter updates for three independent Critic networks and parameter updates for a shared Actor network.
[0032] For any Critic network, the power allocation action corresponding to the next state is obtained using a shared Actor network. The next state and its corresponding power allocation action are then input into the Critic network to obtain the target Q-value. Based on this target Q-value and the reward extracted from the experience, a single target value is calculated. It is important to note that when calculating a single target value, only the single reward corresponding to the Critic network is used. The current state and its corresponding power allocation action from the extracted experience are then input into the Critic network to obtain the predicted Q-value. The mean squared error loss is calculated using the predicted Q-value and the single target value, and gradient descent is performed to optimize and update the Critic network parameters.
[0033] For a shared Actor network, the current state and its corresponding power allocation action are input into three updated Critic networks to obtain three Q-values. The dynamic weights of these three Q-values are calculated using the Pareto front constraint mechanism. The weighted Q-value is obtained by weighting the three Q-values with these dynamic weights. The current policy of the shared Actor network is then calculated. entropy That is, introducing an entropy regularization term into the Actor loss function to maximize the policy entropy. Encouraging exploratory behavior and avoiding local optima caused by a single objective, while combining exploration and utilization with a dynamic balance of temperature coefficients, the shared Actor network aims to maximize expected reward and policy entropy. Its loss function is a minimization problem, which can be expressed as: Based on this loss function, gradient ascent optimization is performed to update the network parameters of the shared Actor network, where, This is the current state. The dynamic weights for the Q-value. The current state and its corresponding power allocation action are input into three updated Critic networks to obtain three Q values. For power distribution action, For network parameters, This is a temperature coefficient used to balance maximizing returns with exploration.
[0034] In this embodiment, the Pareto front constraint mechanism uses Lagrange multipliers to transform the multi-objective problem into a constrained optimization problem, and dynamically adjusts the weights to distribute the solution set along the Pareto front, ensuring coordinated optimization of economy, dynamics, and safety rather than a simple compromise.
[0035] In this embodiment, a Bayesian optimization framework is used during model training. A Gaussian process surrogate model is used to efficiently explore the hyperparameter space (such as learning rate and entropy coefficient). The evaluation point is dynamically selected in combination with the desired improvement acquisition function, thereby globally optimizing the algorithm performance and accelerating convergence within a finite number of iterations.
[0036] Specifically, a Gaussian process (GP) surrogate model is initialized, and five sets of random parameter combinations are sampled. Based on the GP prediction mean and variance, the next set of parameters is selected using the Expectation Improvement (EI) sampling function. The loss function value corresponding to the new parameters is evaluated, and the GP model is updated. The maximum number of iterations (20) is reached, or the change in the loss function is <0.01, for a total of five iterations.
[0037] Understandably, a joint loss function is defined, fusing the Bellman error and constraint terms to drive the Critic network update. The Critic loss function is: for each objective i, minimizing the Bellman error and fusing the Pareto constraint. ,in, The expectation operator represents the sample uniformly sampled from the experience replay pool D. The expected value reflects the characteristics of data distribution. To estimate the Q-value of the action in the current state using the Critic, predict the long-term cumulative reward after taking action a in state s. The immediate reward is the environmental feedback after performing action a, corresponding to the i-th goal. The discount factor is γ (0 < γ < 1), where a larger γ indicates a greater emphasis on long-term returns. To minimize the Q-value of the dual-critic objective, the smaller Q-value in the two-critic network is selected as the objective, suppressing overestimation bias and improving training stability. For Q-value estimation of the target Critic network, For the policy network j in state The generated action is used to calculate the Q-value objective for the next state. These are the Pareto constraint weighting coefficients. For the target Q-value offset, the target network soft update is: By slowly updating the target network parameters ( =0.005), to avoid training oscillations and improve convergence. For the target Critic network, a soft target network update strategy is adopted to stabilize the training process, and an experience replay mechanism is combined to break the data correlation. The parameters of the Actor network and the Critic network are simultaneously optimized through gradient backpropagation to achieve adaptive learning and multi-objective trade-offs of the global energy management strategy.
[0038] It should be noted that after training, the network parameters are fixed, and a large number of policies and their performance evaluation metrics are collected by systematically adjusting the reward weight combinations. Subsequently, the NSGA-II algorithm is used to perform non-dominated ranking and crowding calculation on these policies, ultimately selecting a Pareto-optimal set of policies that consider economy, dynamics, and safety, maintaining the diversity of the solution set distribution to avoid local optima. The neural network parameters corresponding to each policy can be stored.
[0039] Understandably, after training, the strategy performance can be verified under diverse operating conditions. The advantages of the strategy can be quantified by comparing it with benchmark methods (such as ECMS and single-objective SAC) through various evaluation indicators such as hydrogen consumption per 100 kilometers, battery life degradation, and acceleration response time.
[0040] Before practical application, the decision-making model needs to be deployed on vehicles. This can be achieved through model quantification, such as TensorRT INT8 compression, federated learning frameworks, multi-vehicle collaborative training, and hardware-in-the-loop (HIL) testing, to address issues of real-time performance, data privacy, and robustness under extreme operating conditions, thus forming a deployable, engineered energy management solution. By constructing an interactive decision-making interface, technical indicators (such as hydrogen consumption and power response deviation) are mapped to user-understandable objectives (such as total cost of ownership or transit time). Preference weights are captured through controls such as sliders, and the optimal strategy is recommended based on a comprehensive utility function. Users can also select their preferred strategy based on task requirements (such as shortest transit time or lowest TCO).
[0041] For example, assuming the user selects the "lowest TCO" preference, the system calculates the TCO value for each Pareto solution and selects the solution that minimizes the TCO (e.g., f1=0.92, f2=0.15, f3=0.28). Mapped to the hyperparameter combination: η=8.5×10 -5 β=0.12, , , By deploying corresponding strategies to vehicles, performance was achieved with a hydrogen consumption of 9.2 kg per 100 kilometers, a battery life degradation of 15%, and 28 fuel cell start-stop cycles.
[0042] Step 3: Construct a comprehensive state vector consistent with the vehicle state space definition based on vehicle data. Based on the vehicle's historical power demand data, identify the vehicle's driving conditions through adaptive mode decomposition and temporal convolutional networks to obtain the driving condition type.
[0043] In this embodiment, step 3 includes: Step 3.1: Obtain the vehicle's dynamic parameters, energy system status, and environmental information.
[0044] Among them, vehicle dynamic parameters include: vehicle speed ( (km / h), longitudinal acceleration ( m / s 2 ) and the required power of the drive motor ( kW); Energy system status includes: power battery SOC, fuel cell power ( kW) and hydrogen quantity ( Environmental information includes: road slope (θ, °), drag coefficient ( ); ) and rolling resistance coefficient ( ).
[0045] Step 3.2: Normalize the vehicle dynamic parameters, energy system state, and environmental information respectively, and then concatenate them to obtain the comprehensive state vector.
[0046] For example, for vehicle dynamic parameters, vehicle speed ( (km / h), longitudinal acceleration ( m / s 2 ) and the required power of the drive motor ( After normalization, the vectors (kW) are concatenated. ,in, , , .
[0047] For example, the SOC of the power battery is normalized for the energy system state: Mapped to the [0,1] interval, fuel cell power and hydrogen quantity Normalization: ,in, This refers to the capacity of the hydrogen storage tank.
[0048] For example, for environmental information, road slope θ and drag coefficient are obtained through map and sensor fusion. Rolling resistance coefficient After normalization, they are concatenated into a vector. Normalized to [0,1].
[0049] Step 3.2: Use adaptive variational mode decomposition to decompose the historical power demand sequence of the drive motor to obtain low-frequency trend components and high-frequency fluctuation components. Calculate the low-frequency energy ratio and high-frequency fluctuation index based on the low-frequency trend components and high-frequency fluctuation components.
[0050] Please see Figure 2 , Figure 2 This is a flowchart of an adaptive mode decomposition method provided in an embodiment of the present invention, such as... Figure 1As shown, the historical power demand sequence is decomposed into multiple scales. By constructing a variational framework and using the Alternating Directional Multiplier Method (ADMM) to iteratively solve the constrained optimization problem, the original signal is decomposed into multiple narrowband intrinsic mode functions (IMFs). Low-frequency trend components (reflecting average power demand) and high-frequency fluctuation components (capturing transient impacts) are then extracted to provide the time-frequency characteristics of power demand for the strategy. Based on the decomposition results, the low-frequency energy proportion (measuring the strength of long-term trends) and the high-frequency fluctuation index (quantifying the amplitude of transient impacts) are calculated, forming a multi-dimensional feature vector that can distinguish between urban / highway / mountainous operating conditions.
[0051] In this embodiment, the historical demand power sequence can be represented as: , The low-frequency trend component and high-frequency fluctuation component are extracted and represented as follows: , , , As a low-frequency trend component, It is a high-frequency fluctuation component. For the first One intrinsic mode function (IMF). To decompose the mode number, The number of low-frequency components.
[0052] The formula for calculating the proportion of low-frequency energy is as follows: ; In the formula, The proportion of low-frequency energy For the first The energy of each eigenmode function The number of low-frequency components, To decompose the mode number; The formula for calculating the high-frequency volatility index is: ; In the formula, Indicates a high-frequency volatility index. The time period representing historical power demand. This represents high-frequency fluctuation components.
[0053] Step 3.3: Input the low-frequency energy ratio, high-frequency fluctuation index, vehicle speed and longitudinal acceleration as feature vectors into the temporal convolutional network to obtain the driving condition type.
[0054] In this embodiment, a two-layer temporal convolutional network is constructed, taking feature vectors and vehicle states (vehicle speed, acceleration) as inputs. Temporal modeling captures the evolution of driving conditions, and the Softmax function outputs the real-time probability distribution of driving conditions, thus completing the dynamic classification of driving scenarios. Temporal convolutional networks output driving condition probabilities ,in, For urban working conditions, For high-speed operating conditions, Working conditions in mountainous areas.
[0055] Step 4: Based on the driving condition type, match or dynamically adjust the control strategy from the Pareto optimal strategy set to determine the control strategy. Input the integrated state vector into the Actor network in the MO-SAC model corresponding to the control strategy to obtain the preliminary power allocation.
[0056] In this embodiment, step 4 includes: Step 4.1: Determine the weight configuration of economic reward, power reward and safety reward in the multi-objective reward function according to the driving condition type and the preset condition-sensitive reward weight mapping rule.
[0057] For example, the working condition-sensitive reward weight mapping rule is as follows: Urban driving conditions: including frequent starts and stops (such as traffic lights), low-speed following ( The segment is <40 km / h.
[0058] High-speed operating conditions: including stable cruise control ( >80 km / h), slight slope ( <2°) segment.
[0059] Mountainous working conditions: including continuous uphill climbing ( >5°), high power requirements ( (a fragment of ).
[0060] In this embodiment, the designed condition-sensitive reward weight mapping rule can dynamically adjust the economic, dynamic, and safety weight coefficients in the multi-objective optimization function based on the classification results (such as the power weight for improving working conditions in mountainous areas) and the moving average smoothing mechanism. This guides the reinforcement learning strategy to adaptively optimize the power allocation logic in different scenarios, ensuring that the energy management strategy is accurately matched with the actual working condition requirements.
[0061] Step 4.2: Select a matching policy from the Pareto optimal policy set as the control policy based on the weight configuration; or fine-tune the network parameters of the MO-SAC model of the current policy based on the weight configuration.
[0062] For example, after a user inputs a preference weight, the policy closest to that weight is selected from the Pareto optimal policy set using distance metrics (Euclidean distance, cosine similarity).
[0063] In the MO-SAC framework, the entropy temperature parameter directly affects the exploration-balancing and fine-tuning of the strategy. The entropy temperature parameter can be dynamically adjusted according to the user's weight configuration. For example, when increasing the weight of the energy consumption target, the entropy temperature parameter can be reduced to enhance the strategy's focus on energy consumption optimization.
[0064] Step 4.3: Input the integrated state vector into the Actor network in the MO-SAC model corresponding to the control strategy or the Actor network in the fine-tuned MO-SAC model to obtain the initial power allocation.
[0065] For example, suppose the current time The demand power sequence was decomposed by VMD to obtain the low-frequency component. It accounts for 70% of the total energy, and the high-frequency component The standard deviation is 20 kW. The vehicle condition is... , After feature calculation, the following is obtained: , The input to the temporal convolutional network is [0.7, 20, 60 / 120, 2 / 3], and the output driving condition probability is... That is, the probability of working conditions in mountainous areas is 70%.
[0066] Therefore, according to the mapping rules, the weights of economic rewards, motivational rewards, and safety rewards in the multi-objective reward function are dynamically adjusted to [0.4, 0.5, 0.1].
[0067] Therefore, in the multi-objective reward function of step 1, the economic weight is reduced while the power weight is increased, prompting the reinforcement learning strategy to prioritize high-power demands (such as increasing the output ratio of fuel cells) under mountainous conditions. Simultaneously, through... Ensure the constraint of fuel cell power change rate.
[0068] If the current SOC = 0.75, the fuel cell output power =200kW, with a power battery supplement of 50kW, then: Economic incentives: (Assuming) =5A, then = 25); Motivational rewards: ; Security Rewards: (Assuming) =10kW, then = 1); Operating condition adaptability bonus: ; Therefore, the total reward is: .
[0069] Step 5: Correct the initial power allocation based on the fuel cell power change rate constraint and the power battery SOC boundary limit to obtain the final power allocation command.
[0070] Understandably, Actor networks learn under specific simulation environments and training data. When a vehicle encounters an extremely rare and uncommon extreme condition that has never appeared in the training set, the network may output an unreasonable or even dangerous action. Constraints during training are indirectly implemented through "rewards and penalties" and "state transitions," which cannot guarantee 100% safety. Therefore, in practical applications, data-driven intelligent decision-making must undergo safety verification based on physical rules. This is the last and most reliable line of defense for ensuring system safety.
[0071] The present invention provides a multi-objective energy management method for hydrogen-powered heavy-duty trucks based on reinforcement learning. By constructing the MO-SAC algorithm framework and integrating the Pareto front constraint mechanism, it achieves collaborative optimization and decoupled evaluation of multiple objectives such as economy, power, and safety during the training phase, generating a set of Pareto optimal strategies covering different trade-off preferences. In online application, it combines real-time operating condition identification technology to dynamically match or adjust the optimal control strategy, and finally ensures the feasibility and safety of the action through physical constraint correction. Thus, it comprehensively improves the adaptive capability, overall performance, and engineering reliability of the hydrogen-powered heavy-duty truck energy management system in complex and ever-changing actual driving environments.
[0072] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations are intended to cover non-exclusive inclusion, such that an article or device comprising a list of elements includes not only those elements but also other elements not expressly listed. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the article or device comprising said element. Terms such as "connected" or "linked" are not limited to physical or mechanical connections but can include electrical connections, whether direct or indirect. The orientations or positional relationships indicated by terms such as "upper," "lower," "left," and "right" are based on the orientations or positional relationships shown in the accompanying drawings and are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as limiting the invention.
[0073] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features or characteristics described may be combined in any suitable manner in one or more embodiments or examples. In addition, those skilled in the art can combine and integrate the different embodiments or examples described in this specification.
[0074] The above description, in conjunction with specific preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, various simple deductions or substitutions can be made without departing from the concept of the present invention, and all such modifications and substitutions should be considered within the scope of protection of the present invention.
Claims
1. A multi-objective energy management method for hydrogen-powered heavy-duty trucks based on reinforcement learning, characterized in that, include: Step 1: Construct the neural network model of the MO-SAC algorithm, and define the vehicle state space, action space, and multi-objective reward function; Step 2: Based on the vehicle state space, the action space, and the multi-objective reward function, the neural network model is trained through environmental interaction to obtain a trained MO-SAC model. During training, Bayesian optimization and Pareto front constraint mechanisms are introduced. After training, the NSGA-II algorithm is used to search for the Pareto optimal policy set. The process of training the neural network model through environmental interaction includes: S1: Build the simulation environment; S2: Obtain the current state from the simulation environment that is consistent with the vehicle state space definition; S3: Input the current state into the shared Actor network of the neural network model to obtain the corresponding power allocation action; S4: Interact with the simulation environment using the power allocation action to obtain the reward and next state corresponding to the power allocation action; S5: Combine the current state, the power allocation action, the reward, and the next state into an experience and store it in the experience replay area; S5: Repeat S2-S4 until the experience in the experience playback area reaches a preset number; S6: Extract a preset number of experiences from the experience replay area to update the network parameters of the neural network model until the neural network model reaches the preset training cutoff condition, and obtain the trained MO-SAC model. Step 3 includes: Step 3.1: Obtain vehicle dynamic parameters, energy system status, and environmental information; Step 3.2: Normalize the vehicle dynamic parameters, the energy system state, and the environmental information respectively, and then concatenate them to obtain a comprehensive state vector; Step 3.3: Use adaptive variational mode decomposition to decompose the historical power demand sequence of the drive motor to obtain low-frequency trend components and high-frequency fluctuation components, and calculate the low-frequency energy ratio and high-frequency fluctuation index based on the low-frequency trend components and the high-frequency fluctuation components. Step 3.4: Input the low-frequency energy ratio, the high-frequency fluctuation index, vehicle speed, and longitudinal acceleration as feature vectors into a temporal convolutional network to obtain the driving condition type; Step 4: Based on the driving condition type, match or dynamically adjust the control strategy from the Pareto optimal strategy set to determine the control strategy, and input the integrated state vector into the Actor network in the MO-SAC model corresponding to the control strategy to obtain the preliminary power allocation; Step 5: Correct the preliminary power allocation based on the fuel cell power change rate constraint and the power battery SOC boundary limit to obtain the final power allocation command.
2. The multi-objective energy management method for hydrogen-powered heavy-duty trucks based on reinforcement learning according to claim 1, characterized in that, The neural network model includes a shared Actor network and three independent Critic networks; wherein... The shared Actor network is used to generate power allocation actions based on the input integrated state vector; The three independent Critic networks respectively evaluate the long-term value of the economy, power and safety of performing the power distribution action.
3. The multi-objective energy management method for hydrogen-powered heavy-duty trucks based on reinforcement learning according to claim 1, characterized in that, The vehicle state space includes vehicle dynamic parameters, energy system status, and environmental information; wherein... The vehicle dynamic parameters include: vehicle speed, longitudinal acceleration, and drive motor power requirements; The energy system status includes: power battery SOC, fuel cell power, and hydrogen quantity; The environmental information includes: road slope, wind resistance coefficient, and rolling resistance coefficient.
4. The multi-objective energy management method for hydrogen-powered heavy-duty trucks based on reinforcement learning according to claim 1, characterized in that, The action space includes power distribution actions, the fuel cell power change rate constraint, and the power battery SOC boundary limit; wherein... The power distribution action includes the power percentage of the fuel cell; The fuel cell power change rate constraint is: if The power ratio of the fuel cell will then be corrected to... ;in, For the power variation of fuel cells, This refers to the rated power of the drive motor. Indicates time, The power percentage of the fuel cell, For the revised version The power percentage of the fuel cell at any given time. for The power percentage of the fuel cell at any given time. The sign function represents the change in power of the fuel cell when it changes in a positive direction. +1, when the power change of the fuel cell is negative. =-1, when When it is 0 =0, for The power required by the drive motor at any given time; The SOC boundary limit for the power battery is: when the SOC of the power battery is below 0.2 or above 0.8, let... .
5. The multi-objective energy management method for hydrogen-powered heavy-duty trucks based on reinforcement learning according to claim 1, characterized in that, The multi-objective reward function is expressed as follows: ; in, As the weight of economic rewards, As the weight of motivational rewards, Weighting for security rewards; The economic reward function is expressed as: ; In the formula, The first constant coefficient has a value of 1. This is the second constant coefficient, with a value of 0.
5. For hydrogen consumption, This represents the standard deviation of the battery current. The motivational reward function is expressed as: ; In the formula, The third constant coefficient has a value of 2. To meet the power requirements of the drive motor, This refers to the actual power of the drive motor. , The power percentage of the fuel cell, This represents the maximum power of the fuel cell. The maximum power to replenish the battery's power. This indicates that when the fuel cell's power is insufficient, the battery supplements the power. It is +1; when the fuel cell is in excess, the battery absorbs the excess power. =-1; The security reward function is expressed as: ; In the formula, It is the fourth constant coefficient, with a value of 10. This is the fifth constant coefficient, with a value of 0.
1. This refers to the state of charge of the power battery. For indicator functions, For the power variation of fuel cells, This means that the function value is 1 when the battery level is below 20% or above 80%, and 0 otherwise. The adaptive reward function is expressed as: ; In the formula, The sixth constant coefficient has a value of 1. It is the seventh constant coefficient, with a value of 0.
8. Based on hydrogen consumption, This is the longitudinal acceleration.
6. The multi-objective energy management method for hydrogen-powered heavy-duty trucks based on reinforcement learning according to claim 1, characterized in that, The formula for calculating the proportion of low-frequency energy is: ; In the formula, The proportion of low-frequency energy For the first The energy of each eigenmode function The number of low-frequency components, To decompose the mode number; The formula for calculating the high-frequency volatility index is as follows: ; In the formula, Indicates a high-frequency volatility index. The time period representing historical power demand. This represents high-frequency fluctuation components.
7. The multi-objective energy management method for hydrogen-powered heavy-duty trucks based on reinforcement learning according to claim 1, characterized in that, Step 4 includes: Step 4.1: Determine the weight configuration of economic reward, power reward and safety reward in the multi-objective reward function according to the driving condition type and the preset condition-sensitive reward weight mapping rule; Step 4.2: Select a matching policy from the Pareto optimal policy set as the control policy according to the weight configuration; or fine-tune the network parameters of the MO-SAC model of the current policy based on the weight configuration; Step 4.3: Input the integrated state vector into the Actor network in the MO-SAC model corresponding to the control strategy or the Actor network in the fine-tuned MO-SAC model to obtain the preliminary power allocation.
Citation Information
Patent Citations
Integrated energy system scheduling method based on enhanced exploration fallback clipping reinforcement learning
CN118485286A
Comprehensive energy system Pareto optimal scheduling method based on variable multi-target weight
CN121119787A