A range-extending vehicle energy management control method and system, and a storage medium
Patent Information
- Application Number
- CN202611183119.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-05
- Publication Date
- 2026-09-25
AI Technical Summary
部分已有研究尝试将深度强化学习应用于增程式汽车能量管理,但大多聚焦于深度确定性策略梯度算法,该算法存在探索效率低、对超参数敏感、易陷入局部最优等缺陷
[0016]本发明的有益效果:通过构建基于软演员-评论家深度强化学习算法的智能体,利用智能体动态输出发动机目标扭矩及目标扭矩区间的上下边界调整量,实现了目标扭矩区间依据实时工况的自适应调整,从而为发动机运行提供了弹性空间;根据发动机实际输出扭矩与目标扭矩区间的位置关系确定电机补偿扭矩的补偿方向和补偿量,实现了发动机与电机的深度协同,使得发动机能够在更宽广的工况范围内维持高效运行,同时通过电机精准补偿或回收扭矩偏差,进而达到降低燃油消耗、延长电池寿命、改善NVH性能与提升驾驶平顺性的综合优化效果。
Smart Images

Figure CN122808696A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of energy management and control technology for new energy vehicles, and in particular to a method, system, and storage medium for energy management and control of range-extended electric vehicles. Background Technology
[0002] Range-extended electric vehicles (REEVs) combine the low emissions of pure electric vehicles with the long range of traditional gasoline vehicles, making them an important technological route in the new energy vehicle field. Their power system consists of a range extender (combining an engine and generator) and a battery, providing dual power sources. The energy management strategy directly determines the vehicle's fuel economy, battery life, and driving quality. Most existing range-extended vehicle energy management methods employ fixed threshold control strategies. These strategies typically activate the range extender when the battery's state of charge (SOC) is below 20% and deactivate it above 90%, or use rule-based power-following strategies. These fail to adequately consider the dynamic changes in real-time vehicle operating conditions such as speed, acceleration, and road gradient, leading to frequent start-stop cycles or prolonged operation of the range extender in inefficient ranges, increasing fuel consumption and generating significant noise and vibration.
[0003] Furthermore, the torque distribution logic in existing technologies is relatively simplistic. The determination of the engine's target torque is typically based on fixed rules or simple lookup tables, lacking the ability to adaptively adjust for dynamic deviations between the engine's actual output torque and the vehicle's required torque. While compensation or recovery can be achieved through the electric motor when the engine's output torque deviates from the actual demand, existing methods still rely too heavily on coarse-grained dynamic adjustment of torque deviations. Some existing research has attempted to apply deep reinforcement learning to the energy management of range-extended electric vehicles, but most focus on deep deterministic policy gradient algorithms. These algorithms suffer from low exploration efficiency, sensitivity to hyperparameters, and a tendency to get trapped in local optima.
[0004] Therefore, there is an urgent need to develop an energy management and control method for range-extended vehicles that can achieve deep collaboration between the engine and the electric motor and dynamically adapt to real-time operating conditions. Summary of the Invention
[0005] In view of the above-mentioned technical problems in related technologies, the present invention proposes a range-extended vehicle energy management control method, system and storage medium, which can overcome the above-mentioned shortcomings of the prior art.
[0006] To achieve the above-mentioned technical objectives, the technical solution of the present invention is implemented as follows: According to a first aspect of the present invention, a method for energy management control of a range-extended electric vehicle is provided; The range-extended electric vehicle energy management control method includes: An intelligent agent is constructed based on the soft actor-critic deep reinforcement learning algorithm, and the state space, action space and reward function of the intelligent agent are defined. The intelligent agent outputs the engine target torque, motor compensation torque, lower boundary adjustment amount, and upper boundary adjustment amount for the current control cycle. A target torque range is determined based on the engine target torque, the lower boundary adjustment amount, and the upper boundary adjustment amount. The target torque range has a lower boundary and an upper boundary. The lower boundary is equal to the engine target torque minus the lower boundary adjustment amount, and the upper boundary is equal to the engine target torque plus the upper boundary adjustment amount. Obtain the actual output torque of the engine in the current control cycle; Based on the comparison between the actual output torque and the target torque range, the compensation direction and compensation amount of the motor compensation torque are determined. The engine is controlled to operate according to the engine target torque, and the motor is controlled to operate according to the motor compensation torque.
[0007] Further, determining the compensation direction and amount of the motor compensation torque based on the positional relationship between the actual output torque and the target torque range includes: When the actual output torque is less than the lower boundary, the motor compensation torque is determined to be positive. The compensation amount is the difference between the lower boundary and the actual output torque, and the compensation amount does not exceed the maximum available compensation torque of the motor under the current battery charge state. When the actual output torque is greater than the upper boundary, the motor compensation torque is determined to be negative. The compensation amount is the difference between the actual output torque and the upper boundary, and the absolute value of the compensation amount does not exceed the maximum allowable recovery torque of the motor under the current battery charge state. When the actual output torque is within the target torque range, the compensation amount of the motor compensation torque is determined to be zero.
[0008] Furthermore, the reward function includes at least one of the following: fuel consumption reward, battery health status reward, NVH performance reward, driving smoothness reward, SOC maintenance reward, and thermal status reward. The fuel consumption reward is determined based on the instantaneous fuel consumption rate, which is determined based on the engine universal characteristic fuel consumption rate mapping function, the engine's actual output torque, and the engine speed. The battery health status reward is determined based on the change in battery state of charge and battery current. The NVH performance bonus items are determined based on the engine torque change rate and / or engine speed change rate. The driving smoothness bonus is determined based on the rate of change of motor compensation torque. The SOC maintenance reward applies a penalty when the battery state of charge exceeds a preset lower limit or a preset upper limit. The thermal state reward applies a penalty when the battery temperature or engine temperature exceeds a safe temperature threshold.
[0009] Furthermore, the state space includes at least one of the following: vehicle speed, acceleration, battery state of charge, battery health status, battery temperature, engine temperature, engine speed, vehicle required torque, and deviation between the actual engine output torque and the center value of the target torque range. The motion space includes the engine target torque, the motor compensation torque, the lower boundary adjustment amount, and the upper boundary adjustment amount.
[0010] Furthermore, the intelligent agent constructed based on the soft actor-critic deep reinforcement learning algorithm includes: Based on the aforementioned soft actor-critic deep reinforcement learning algorithm, an Actor network and a Critic network are constructed. The Actor network outputs an action based on the current state, and the Critic network evaluates the state-action value. The optimization objective of the soft actor-critic deep reinforcement learning algorithm is expressed as follows: ; in, Let be the policy, representing the mapping from states to the probability distribution of actions; For policy entropy; This is a temperature coefficient used to adjust the trade-off between reward and entropy; In the state Take action below The reward value obtained; The state at time t; For the action at time t; For strategy State-action marginal distribution; Total number of control steps; Indicates by strategy Calculate the expected value of the state-action distribution.
[0011] Furthermore, the update objective of the Critic network is to minimize the soft Bellman residual, expressed as: ; in, These are the parameters of the current Critic network. The parameters of the target Critic network are... These are the parameters of the Actor network; and These represent the state-action values output by the current Critic network and the target Critic network, respectively. For Actor networks; As a discount factor, The value range is [0,1]; For experience replay pool; Let be the instantaneous reward value at time t; This represents the state at time t+1. This refers to the action at time t+1. This indicates that the expectation is calculated by randomly sampling from the experience replay pool; Indicates according to the current strategy Sample the action at time t+1 and calculate the expected value; The update objective of the Actor network is to minimize the following expression: ; in, For Actor network parameters, These are the current Critic network parameters. For experience replay pool, Let t be the state at time t. Let t be the action at time t. For Actor networks, For the Critic network, Temperature coefficient; This indicates that the expected value is calculated by sampling the state from the experience replay pool. Indicates according to the current strategy Sample the action at time t and calculate the expectation.
[0012] Furthermore, the Critic network comprises two Q-networks with identical structures and independent parameters. During training, the smaller value of the outputs of the two Q-networks is taken as the target Q-value.
[0013] Furthermore, there are offline training phases and online fine-tuning phases: During the offline training phase, the agent is trained using various standard driving cycle data until the reward stabilizes and converges. During the online fine-tuning phase, the trained and converged agent is deployed to the vehicle controller, and the network parameters are incrementally updated based on the real-time collected vehicle status data.
[0014] According to a second aspect of the present invention, a range-extended electric vehicle energy management control system is provided; The range-extended vehicle energy management control system includes a memory, a processor, and a computer program stored in the memory and executable on the processor. The feature is that when the processor executes the program, it implements a range-extended vehicle energy management control method as described above.
[0015] According to a third aspect of the present invention, a computer-readable storage medium is provided; The computer-readable storage medium, when the computer program is executed by a processor, implements the energy management control method for a range-extended vehicle as described above.
[0016] The beneficial effects of this invention are as follows: By constructing an intelligent agent based on a soft actor-critic deep reinforcement learning algorithm, the agent dynamically outputs the engine's target torque and the adjustment amount of the upper and lower boundaries of the target torque range, thereby achieving adaptive adjustment of the target torque range according to real-time operating conditions, thus providing flexible space for engine operation; the compensation direction and amount of the motor compensation torque are determined according to the positional relationship between the engine's actual output torque and the target torque range, realizing deep collaboration between the engine and the motor, enabling the engine to maintain efficient operation over a wider range of operating conditions. At the same time, the motor accurately compensates for or recovers torque deviations, thereby achieving a comprehensive optimization effect of reducing fuel consumption, extending battery life, improving NVH performance, and enhancing driving smoothness. Attached Figure Description
[0017] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0018] Figure 1 This is a flowchart of an energy management control method for a range-extended electric vehicle according to the present invention; Figure 2 This describes the main data flow of the range-extended vehicle energy management control method described in this invention. Detailed Implementation
[0019] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0020] In the description of this invention, it should be understood that the terms "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," "outer," "clockwise," and "counterclockwise," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limiting this invention.
[0021] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified. Furthermore, the terms "installed," "connected," and "linked" should be interpreted broadly; for example, they may refer to a fixed connection, a detachable connection, or an integral connection; they may refer to a mechanical connection or an electrical connection; they may refer to a direct connection or an indirect connection through an intermediate medium; and they may refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0022] like Figure 1 and Figure 2 As shown, a range-extended vehicle energy management control method according to an embodiment of the present invention includes: An intelligent agent is constructed based on the soft actor-critic deep reinforcement learning algorithm, and the state space, action space and reward function of the intelligent agent are defined. The intelligent agent outputs the engine target torque, motor compensation torque, lower boundary adjustment amount, and upper boundary adjustment amount for the current control cycle. A target torque range is determined based on the engine target torque, the lower boundary adjustment amount, and the upper boundary adjustment amount. The target torque range has a lower boundary and an upper boundary. The lower boundary is equal to the engine target torque minus the lower boundary adjustment amount, and the upper boundary is equal to the engine target torque plus the upper boundary adjustment amount. Obtain the actual output torque of the engine in the current control cycle; Based on the comparison between the actual output torque and the target torque range, the compensation direction and compensation amount of the motor compensation torque are determined. The engine is controlled to operate according to the engine target torque, and the motor is controlled to operate according to the motor compensation torque.
[0023] According to an embodiment of the present invention, an energy management control method for a range-extended electric vehicle, in a specific implementation, determining the compensation direction and amount of the motor compensation torque based on the positional relationship between the actual output torque and the target torque range includes: When the actual output torque is less than the lower boundary, the motor compensation torque is determined to be positive. The compensation amount is the difference between the lower boundary and the actual output torque, and the compensation amount does not exceed the maximum available compensation torque of the motor under the current battery charge state. When the actual output torque is greater than the upper boundary, the motor compensation torque is determined to be negative. The compensation amount is the difference between the actual output torque and the upper boundary, and the absolute value of the compensation amount does not exceed the maximum allowable recovery torque of the motor under the current battery charge state. When the actual output torque is within the target torque range, the compensation amount of the motor compensation torque is determined to be zero.
[0024] According to an embodiment of the present invention, an energy management control method for a range-extended electric vehicle, in a specific implementation, the reward function includes at least one of the following: fuel consumption reward, battery health status reward, NVH performance reward, driving smoothness reward, SOC maintenance reward, and thermal status reward. The fuel consumption reward is determined based on the instantaneous fuel consumption rate, which is determined based on the engine universal characteristic fuel consumption rate mapping function, the engine's actual output torque, and the engine speed. The battery health status reward is determined based on the change in battery state of charge and battery current. The NVH performance bonus items are determined based on the engine torque change rate and / or engine speed change rate. The driving smoothness bonus is determined based on the rate of change of motor compensation torque. The SOC maintenance reward applies a penalty when the battery state of charge exceeds a preset lower limit or a preset upper limit. The thermal state reward applies a penalty when the battery temperature or engine temperature exceeds a safe temperature threshold.
[0025] According to an embodiment of the present invention, the energy management control method for a range-extended electric vehicle includes, in a specific implementation, at least one of the following: vehicle speed, acceleration, battery state of charge, battery health status, battery temperature, engine temperature, engine speed, vehicle required torque, and deviation of the actual engine output torque from the center value of the target torque range. The motion space includes the engine target torque, the motor compensation torque, the lower boundary adjustment amount, and the upper boundary adjustment amount.
[0026] According to an embodiment of the present invention, a range-extended electric vehicle energy management control method, in a specific implementation, includes constructing an intelligent agent based on a soft actor-critic deep reinforcement learning algorithm, comprising: Based on the aforementioned soft actor-critic deep reinforcement learning algorithm, an Actor network and a Critic network are constructed. The Actor network outputs an action based on the current state, and the Critic network evaluates the state-action value. The optimization objective of the soft actor-critic deep reinforcement learning algorithm is expressed as follows: ; in, Let be the policy, representing the mapping from states to the probability distribution of actions; For policy entropy; This is a temperature coefficient used to adjust the trade-off between reward and entropy; In the state Take action below The reward value obtained; The state at time t; For the action at time t; For strategy State-action marginal distribution; Total number of control steps; Indicates by strategy Calculate the expected value of the state-action distribution.
[0027] According to an embodiment of the energy management control method for a range-extended electric vehicle, in a specific implementation, the update objective of the Critic network is to minimize the soft Bellman residual, expressed as: ; in, These are the parameters of the current Critic network. The parameters of the target Critic network are... These are the parameters of the Actor network; and These represent the state-action values output by the current Critic network and the target Critic network, respectively. For Actor networks; As a discount factor, The value range is [0,1]; For experience replay pool; Let be the instantaneous reward value at time t; This represents the state at time t+1. This refers to the action at time t+1. This indicates that the expectation is calculated by randomly sampling from the experience replay pool; Indicates according to the current strategy Sample the action at time t+1 and calculate the expected value; The update objective of the Actor network is to minimize the following expression: ; in, For Actor network parameters, These are the current Critic network parameters. For experience replay pool, Let t be the state at time t. Let t be the action at time t. For Actor networks, For the Critic network, Temperature coefficient; This indicates that the expected value is calculated by sampling the state from the experience replay pool. Indicates according to the current strategy Sample the action at time t and calculate the expectation.
[0028] According to an embodiment of the present invention, in a specific implementation of the energy management control method for a range-extended electric vehicle, the Critic network includes two Q networks with identical structures and independent parameters. During training, the smaller value of the output of the two Q networks is taken as the target Q value.
[0029] According to an embodiment of the present invention, the energy management control method for a range-extended electric vehicle further includes, in a specific implementation, an offline training phase and an online fine-tuning phase. During the offline training phase, the agent is trained using various standard driving cycle data until the reward stabilizes and converges. During the online fine-tuning phase, the trained and converged agent is deployed to the vehicle controller, and the network parameters are incrementally updated based on the real-time collected vehicle status data.
[0030] Secondly, according to an embodiment of the present invention, a range-extended vehicle energy management control system includes a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the program to implement a range-extended vehicle energy management control method as described above.
[0031] Thirdly, a computer-readable storage medium according to embodiments of the present invention; The computer-readable storage medium, when the computer program is executed by a processor, implements the energy management control method for a range-extended vehicle as described above.
[0032] To facilitate understanding of the above technical solutions of the present invention, the following detailed description of the above technical solutions of the present invention is provided through specific embodiments and working principles.
[0033] Example 1: Design of agent network structure and state / action space based on soft actor-critic (SAC) deep reinforcement learning algorithm, i.e., design of SAC agent network structure and state / action space; In this embodiment, both the Actor network and the Critic network of the SAC agent adopt a fully connected neural network structure. The input layer dimension of the Actor network is 10, corresponding to the dimension of the state space. The state space SS is defined as follows: ; in, For vehicle speed, For acceleration, The battery is in its state of charge. For battery health status, For battery temperature, For engine temperature, The current engine speed. For the torque required by the whole vehicle, This refers to the actual output torque of the engine. This represents the deviation between the actual engine torque and the center value of the target torque range. This state space encompasses vehicle motion state, battery state, engine state, and torque deviation information, providing sufficient environmental perception information for intelligent agent decision-making.
[0034] The Actor network contains three hidden layers with 256, 256, and 128 neurons respectively. The activation function is ReLU. The output layer uses the tanh activation function to map actions to the [-1, 1] interval, followed by linear scaling to obtain the actual action value. The action space A is defined as: ; in, The target torque for the engine. To compensate for the torque of the drive motor, and These are the lower and upper boundary adjustment values, respectively. The engine target torque is used to send commands to the engine controller; a positive value for the motor compensation torque indicates supplementary motor drive, and a negative value indicates energy recovery by the motor; the lower and upper boundary adjustment values are used to dynamically adjust the width of the target torque range. When the battery... Sufficient and When the condition is good, the two adjustment values are moderately increased, allowing engine torque to fluctuate within a wider range, providing greater freedom for the engine to operate within its optimal fuel economy range; when the battery... lower or When the battery deteriorates, the two adjustment values automatically narrow to ensure that the engine output torque closely follows the demand and to avoid over-discharge of the battery.
[0035] The Critic network comprises two independent Q-networks, Q1 and Q2. Each Q-network has a 14-dimensional input layer (10-dimensional state plus 4-dimensional action), three hidden layers, and 256, 256, and 128 neurons respectively. The activation function is ReLU, and the output layer uses linear activation, outputting a single Q-value. During training, the minimum of Q1 and Q2 is used as the target Q-value to suppress Q-value overestimation. The initial temperature coefficient α is set to 0.2, and the learning rate is set to 3 × 10⁻⁶. 4 The discount factor γ = 0.99 and the soft update coefficient τ = 0.005.
[0036] When the intelligent agent is running, it will collect the state at the current moment. Input to an Actor network, output action from the Actor network. This includes the engine target torque. Motor compensation torque Lower boundary adjustment amount and upper boundary adjustment amount .
[0037] Example 2: Determination of dynamic target torque range and adaptive adjustment mechanism; This embodiment details the method for determining the target torque range. Unlike traditional methods that set a single target torque value, this embodiment defines a dynamic target torque range. as follows: ; Among them, the lower boundary and upper boundary They are respectively determined as: ; ; The maximum value of the range width adjustment is dynamically limited based on the engine target torque: ; ; This means the target torque range can be dynamically adjusted within the range of -15% to +20% of the engine's target torque. This asymmetric range design fully considers the physical characteristic that the engine's positive torque adjustment capability is better than its negative adjustment capability. That is, the engine responds faster when increasing torque output, but when reducing torque, it is limited by the intake system and mechanical inertia, and the response is relatively slower. Therefore, the positive allowable range is slightly larger than the negative range, enabling the engine to operate in a more efficient and stable operating range.
[0038] During the decision-making process, if the agent detects that the engine's current operating point deviates from its optimal fuel consumption range, it will proactively adjust... and The essence of guiding engine torque toward the high-efficiency zone lies in adjusting the boundary position of the target range, thereby changing the leniency of the condition that the actual torque is within the range without changing the engine target torque setpoint, thus indirectly affecting the adjustment direction of the engine target torque in subsequent control cycles.
[0039] Actual engine output torque The relationship with the target torque range is determined as follows: Under normal operating conditions, the actual torque is within the target range; Insufficient torque requires positive motor compensation. Excess torque necessitates the use of a motor to recover energy.
[0040] Example 3: Engine-Motor Cooperative Compensation Mechanism and Multi-Objective Integrated Reward Function; This embodiment details the specific logic for determining the compensation direction and amount of the motor compensation torque based on the positional relationship between the engine's actual output torque and the target torque range.
[0041] Scenario 1: Insufficient torque: When When the actual engine output torque is lower than the lower boundary of the target range, it means the engine output cannot meet the vehicle's requirements. At this time, the electric motor outputs positive compensation torque, and the compensation amount... This is the difference between the lower boundary and the actual torque, while also being limited by the maximum available compensated torque of the motor under the current battery SOC: ; in This represents the maximum available compensated torque for the motor under the current SOC. Simultaneously, the engine target torque is adjusted upwards in the next control cycle. ; in ∈(0,1) is the compensation attenuation coefficient. By gradually increasing the engine's target torque, the engine gradually takes on more load in subsequent control cycles, avoiding sudden torque changes caused by a one-time large adjustment, thereby reducing long-term dependence on motor compensation while maintaining smoothness.
[0042] Scenario 2: Excess torque: When When the engine's actual output torque exceeds the upper boundary of the target range, it means the engine output exceeds the vehicle's requirements. At this time, the electric motor performs energy recovery, with a negative compensation torque. The absolute value of the compensation amount does not exceed the maximum allowable recovery torque under the current SOC. ; in This represents the maximum permissible regenerative torque at the current SOC. The regenerative power is: ; in This refers to the motor speed. The recycling efficiency function takes into account battery SOC and temperature.
[0043] Scenario 3: Operation within the interval: When When the actual output torque of the engine is within the target torque range, the motor compensation torque is zero, and the engine independently meets the needs of the entire vehicle.
[0044] The multi-objective comprehensive reward function in this embodiment Includes fuel consumption rewards, battery health rewards, NVH performance rewards, driving smoothness rewards, SOC maintenance rewards, and thermal state rewards: ; The coefficients for each reward item are configured as follows: .
[0045] The design principles of each reward item are as follows: Fuel Consumption Rewards ,in The instantaneous fuel consumption rate is obtained by interpolating the engine's actual output torque and speed using the engine's universal characteristic fuel consumption rate mapping function. This reward item guides the agent to tend to choose the engine operating point with the lowest fuel consumption rate.
[0046] The battery health status reward is determined based on the change in battery state of charge and battery current, guiding the agent to avoid control actions that would lead to high-current charging and discharging or deep charging and discharging of the battery.
[0047] NVH performance bonuses: Dramatic changes in engine torque and speed are the main excitation sources of noise and vibration. By penalizing large rates of change, the agent is guided to output smooth control commands.
[0048] Driving smoothness bonuses: Sudden changes in motor compensation torque can cause a step change in the vehicle's driving torque, resulting in a sense of impact. By penalizing these sudden changes, driving smoothness can be improved.
[0049] The SOC maintenance reward and thermal state reward apply penalties when the battery state of charge exceeds a preset range or when the battery temperature or engine temperature exceeds a safety threshold, respectively, to guide the agent to maintain the system within the safety boundary.
[0050] Example 4: Training and online deployment of SAC agents; This embodiment illustrates the training method and deployment process of the SAC agent. A strategy combining offline training and online fine-tuning is employed.
[0051] Offline training phase: The SAC agent was trained using composite scenarios consisting of four standard driving cycles: UDDS, HWFET, WLTC, and CLTC-P, for a total of 2000 training rounds. The Critic network was updated by minimizing the soft Bellman residuals. ; in For the current Critic network, For the target Critic network, For Actor networks, As a discount factor, This serves as an experience replay pool. The target value includes a reward for the policy entropy of the next state, ensuring that the value assessment considers not only the expected return but also the degree of policy randomness, thereby encouraging the agent to maintain a high willingness to explore among actions with similar Q values.
[0052] The Actor network updates by minimizing the following objective: ; The optimization objective of the Actor network is to maximize both the Q-value and the policy entropy, that is, to maintain action diversity while pursuing high rewards. The temperature coefficient α is automatically adjusted during training, avoiding the difficulties of manual parameter tuning.
[0053] Online fine-tuning phase: The SAC agent, after offline training convergence, is deployed to the vehicle controller. During actual driving, the network parameters are incrementally updated based on real-time vehicle status data to adapt to personalized driving habits and specific driving environments. The experience replay pool capacity is 106106, and the mini-batch sampling size is 256.
[0054] In summary, by utilizing the technical solution of this invention, an intelligent agent based on a soft actor-critic deep reinforcement learning algorithm is constructed. This agent dynamically outputs the engine's target torque and the adjustment amounts of the upper and lower boundaries of the target torque range, enabling adaptive adjustment of the target torque range according to real-time operating conditions. This provides flexibility for engine operation. Furthermore, by determining the compensation direction and amount of the motor's compensation torque based on the positional relationship between the engine's actual output torque and the target torque range, deep collaboration between the engine and motor is achieved. This allows the engine to maintain efficient operation over a wider range of operating conditions. Simultaneously, the precise compensation or recovery of torque deviations by the motor achieves a comprehensive optimization effect, reducing fuel consumption, extending battery life, improving NVH performance, and enhancing driving smoothness.
[0055] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for energy management and control of a range-extended electric vehicle, characterized in that, include: An intelligent agent is constructed based on the soft actor-critic deep reinforcement learning algorithm, and the state space, action space and reward function of the intelligent agent are defined. The intelligent agent outputs the engine target torque, motor compensation torque, lower boundary adjustment amount, and upper boundary adjustment amount for the current control cycle. A target torque range is determined based on the engine target torque, the lower boundary adjustment amount, and the upper boundary adjustment amount. The target torque range has a lower boundary and an upper boundary. The lower boundary is equal to the engine target torque minus the lower boundary adjustment amount, and the upper boundary is equal to the engine target torque plus the upper boundary adjustment amount. Obtain the actual output torque of the engine in the current control cycle; Based on the comparison between the actual output torque and the target torque range, the compensation direction and compensation amount of the motor compensation torque are determined. The engine is controlled to operate according to the engine target torque, and the motor is controlled to operate according to the motor compensation torque.
2. The energy management and control method for a range-extended electric vehicle according to claim 1, characterized in that, The step of determining the compensation direction and amount of the motor compensation torque based on the positional relationship between the actual output torque and the target torque range includes: When the actual output torque is less than the lower boundary, the motor compensation torque is determined to be positive. The compensation amount is the difference between the lower boundary and the actual output torque, and the compensation amount does not exceed the maximum available compensation torque of the motor under the current battery charge state. When the actual output torque is greater than the upper boundary, the motor compensation torque is determined to be negative. The compensation amount is the difference between the actual output torque and the upper boundary, and the absolute value of the compensation amount does not exceed the maximum allowable recovery torque of the motor under the current battery charge state. When the actual output torque is within the target torque range, the compensation amount of the motor compensation torque is determined to be zero.
3. The energy management and control method for a range-extended electric vehicle according to claim 1, characterized in that, The reward function includes at least one of the following: fuel consumption reward, battery health status reward, NVH performance reward, driving smoothness reward, SOC maintenance reward, and thermal status reward. The fuel consumption reward is determined based on the instantaneous fuel consumption rate, which is determined based on the engine universal characteristic fuel consumption rate mapping function, the engine's actual output torque, and the engine speed. The battery health status reward is determined based on the change in battery state of charge and battery current. The NVH performance bonus items are determined based on the engine torque change rate and / or engine speed change rate. The driving smoothness bonus is determined based on the rate of change of motor compensation torque. The SOC maintenance reward applies a penalty when the battery state of charge exceeds a preset lower limit or a preset upper limit. The thermal state reward applies a penalty when the battery temperature or engine temperature exceeds a safe temperature threshold.
4. The energy management and control method for a range-extended electric vehicle according to claim 1, characterized in that, The state space includes at least one of the following: vehicle speed, acceleration, battery state of charge, battery health status, battery temperature, engine temperature, engine speed, vehicle required torque, and the deviation between the actual engine output torque and the center value of the target torque range. The motion space includes the engine target torque, the motor compensation torque, the lower boundary adjustment amount, and the upper boundary adjustment amount.
5. The energy management and control method for a range-extended electric vehicle according to claim 4, characterized in that, The intelligent agent constructed based on the soft actor-critic deep reinforcement learning algorithm includes: Based on the aforementioned soft actor-critic deep reinforcement learning algorithm, an Actor network and a Critic network are constructed. The Actor network outputs an action based on the current state, and the Critic network evaluates the state-action value. The optimization objective of the soft actor-critic deep reinforcement learning algorithm is expressed as follows: ; in, Let be the policy, representing the mapping from states to the probability distribution of actions; For policy entropy; This is a temperature coefficient used to adjust the trade-off between reward and entropy; In the state Take action below The reward value obtained; The state at time t; For the action at time t; For strategy State-action marginal distribution; Total number of control steps; Indicates by strategy Calculate the expected value of the state-action distribution.
6. The energy management and control method for a range-extended electric vehicle according to claim 5, characterized in that, The update objective of the Critic network is to minimize the soft Bellman residual, expressed as: ; in, These are the parameters of the current Critic network. The parameters of the target Critic network are... These are the parameters of the Actor network; and These represent the state-action values output by the current Critic network and the target Critic network, respectively. For Actor networks; As a discount factor, The value range is [0,1]; For experience replay pool; Let be the instantaneous reward value at time t; This represents the state at time t+1. This refers to the action at time t+1. This indicates that the expectation is calculated by randomly sampling from the experience replay pool; Indicates according to the current strategy Sample the action at time t+1 and calculate the expected value; The update objective of the Actor network is to minimize the following expression: ; in, For Actor network parameters, These are the current Critic network parameters. For experience replay pool, Let t be the state at time t. Let t be the action at time t. For Actor networks, For the Critic network, Temperature coefficient; This indicates that the expected value is calculated by sampling the state from the experience replay pool. Indicates according to the current strategy Calculate the expected value of the action at time t.
7. The energy management and control method for a range-extended electric vehicle according to claim 6, characterized in that, The Critic network consists of two Q networks with identical structures and independent parameters. During training, the smaller value of the output of the two Q networks is taken as the target Q value.
8. The energy management control method for a range-extended electric vehicle according to claim 1, characterized in that, Also includes: Offline training phase and online fine-tuning phase: During the offline training phase, the agent is trained using various standard driving cycle data until the reward stabilizes and converges. During the online fine-tuning phase, the trained and converged agent is deployed to the vehicle controller, and the network parameters are incrementally updated based on the real-time collected vehicle status data.
9. A range-extended electric vehicle energy management control system, characterized in that, The system includes a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor, when executing the program, implements the method as described in any one of claims 1 to 8.
10. A computer-readable storage medium storing a computer program thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 8.