Real-time energy management method, system and equipment for hybrid electric vehicle
The hybrid electric vehicle energy management system uses LSTM and TD3 networks to predict and optimize engine power for real-time energy efficiency, addressing adaptability and performance issues in existing methods.
Patent Information
- Application Number
- CN202510698676.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-07-15
AI Technical Summary
Existing energy management methods for hybrid electric vehicles face challenges in real-time adaptability due to reliance on human experience or lack of consideration for future uncertainties, leading to suboptimal energy utilization and performance.
A hybrid electric vehicle energy management system using LSTM networks for speed prediction, TD3 networks for power management, and an optimization function to determine optimal engine output power based on historical data and real-time conditions, ensuring real-time energy efficiency.
The system achieves real-time energy management by optimizing engine output power, improving energy utilization and adaptability to changing road conditions, and enhancing overall vehicle performance.
Smart Images

Figure CN120308087A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of hybrid electric vehicles, and in particular, to a real-time energy management method, system and device for hybrid electric vehicles. Background Art
[0002] While the traditional fuel vehicle industry has been booming, the problems of energy consumption and environmental pollution have become increasingly prominent. Against this background, hybrid electric vehicles, by integrating the advantages of pure electric vehicles and fuel vehicles, have received extensive attention as an effective short-term energy alternative. However, due to the existence of multiple power sources in hybrid electric vehicles, the power distribution method between the engine and the motor, that is, the energy management method, is crucial for fully tapping the vehicle's energy-saving potential, optimizing the vehicle's power performance and reducing pollutant emissions. However, the traditional related technologies have the following deficiencies.
[0003] Currently, the adopted energy management methods are mainly divided into two categories: rule-based and optimization-based, but they each face challenges. For example: rule-based strategies rely extremely on manual experience. Once the road conditions change, rules need to be re-established, resulting in the inapplicability of the original rules and the inability to provide real-time energy management for hybrid electric vehicles. And dynamic programming (DP), as a typical optimization-based energy management method, seriously affects its real-time performance and robustness because it requires prior knowledge of the entire driving cycle information. At the same time, with the rapid development of artificial intelligence technology, methods based on reinforcement learning have been favored in the industry due to their strong self-learning ability. However, most of the traditional methods based on reinforcement learning do not consider the uncertainty of future driving cycles, which will lead to a reduction in energy utilization efficiency. Summary of the Invention
[0004] The purpose of the present application is to provide a real-time energy management method, system and device for hybrid electric vehicles, which can improve the real-time performance and energy utilization efficiency of hybrid electric vehicles.
[0005] To achieve the above object, the present application provides the following solutions.
[0006] In a first aspect, the present application provides a real-time energy management method for hybrid electric vehicles, including the following steps.
[0007] Obtain the historical speed sequence of the hybrid electric vehicle and the remaining battery power at the current moment; the historical speed sequence is the set of speeds from the historical moment to the current moment.
[0008] Input the historical speed sequence into the speed prediction model to obtain a predicted speed sequence; the speed prediction model is obtained by training an LSTM network with speed sample data.
[0009] Input the predicted speed sequence into the vehicle required power prediction model to obtain the predicted vehicle required power sequence.
[0010] Based on the predicted speed sequence, the remaining battery power at the current moment, and the predicted vehicle required power sequence, use the energy management control model to obtain the predicted remaining battery power sequence and the predicted engine output power sequence.
[0011] Construct an energy management optimization objective function with the goal of minimizing the real-time energy of the hybrid vehicle.
[0012] Based on the predicted remaining battery power sequence and the predicted engine output power sequence, use the energy management optimization objective function to determine the optimal remaining battery power and the optimal engine output power.
[0013] Control the hybrid vehicle to travel through the optimal engine output power to achieve real-time energy management of the hybrid vehicle.
[0014] In a second aspect, the present application provides a real-time energy management system for a hybrid vehicle, including: an acquisition module for acquiring the historical speed sequence of the hybrid vehicle and the remaining battery power at the current moment; the historical speed sequence is a set of speeds from the historical moment to the current moment.
[0015] A speed prediction module for inputting the historical speed sequence into a speed prediction model to obtain a predicted speed sequence; the speed prediction model is obtained by training an LSTM network with speed sample data.
[0016] A vehicle required power prediction module for inputting the predicted speed sequence into a vehicle required power prediction model to obtain a predicted vehicle required power sequence.
[0017] An energy management control module for obtaining a predicted remaining battery power sequence and a predicted engine output power sequence by using an energy management control model based on the predicted speed sequence, the remaining battery power at the current moment, and the predicted vehicle required power sequence.
[0018] An optimization objective function construction module for constructing an energy management optimization objective function with the goal of minimizing the real-time energy of the hybrid vehicle.
[0019] A determination module for determining the optimal remaining battery power and the optimal engine output power by using the energy management optimization objective function based on the predicted remaining battery power sequence and the predicted engine output power sequence.
[0020] A control module for controlling the hybrid vehicle to travel through the optimal engine output power to achieve real-time energy management of the hybrid vehicle.
[0021] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the steps of the hybrid vehicle real-time energy management method described in the first aspect above.
[0022] According to the specific embodiments provided by the present application, the present application has the following technical effects.
[0023] The present application provides a hybrid vehicle real-time energy management method, system, and device. Based on the historical speed sequence and the remaining battery power at the current moment, using a speed prediction model and an energy management control model, a predicted remaining battery power sequence and a predicted engine output power sequence are determined. Then, through the energy management optimization objective function, the optimal engine output power for controlling the driving of the hybrid vehicle is determined. Through the calculation of the method of the present application, finally, the optimal engine output power is output in real time, thereby controlling the driving of the hybrid vehicle, completing real-time prediction, and solving the problem that the rules formulated based on manual experience are no longer applicable when the road conditions change; at the same time, finally, with the goal of minimizing the real-time energy of the hybrid vehicle, the predicted engine output power sequence in the future driving cycle is calculated, and the optimal engine output power is obtained to control the driving of the hybrid vehicle, thereby improving the energy utilization efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0025] Figure 1 It is a schematic flow chart of a hybrid vehicle real-time energy management method provided by an embodiment of the present application.
[0026] Figure 2 It is a schematic flow chart of LSTM network training provided by an embodiment of the present application.
[0027] Figure 3 It is a schematic diagram of speed prediction results provided by an embodiment of the present application.
[0028] Figure 4 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0030] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0031] In an exemplary embodiment, as Figure 1 shown, a real-time energy management method for a hybrid vehicle is provided. This method is executed by a computer device, specifically, it can be executed alone by a computer device such as a terminal or a server, or jointly executed by a terminal and a server. In the embodiments of the present application, taking the application of this method to a server as an example for illustration, it includes the following steps 1 to step 6.
[0032] Step 1: Obtain the historical speed sequence of the hybrid vehicle and the remaining power at the current moment; the historical speed sequence is a set of speeds from the historical moment to the current moment.
[0033] Step 2: Input the historical speed sequence into the speed prediction model to obtain a predicted speed sequence; the speed prediction model is obtained by training an LSTM network with speed sample data.
[0034] Step 3: Input the predicted speed sequence into the vehicle required power prediction model to obtain a predicted vehicle required power sequence.
[0035] Specifically, for the vehicle required power at time t, calculate the vehicle required power at time t through the speed at time t. The expression of the vehicle required power prediction model is as follows.
[0036] P dem (t) = G′(v(t)).
[0037] Wherein, P dem (t) is the vehicle required power at time t; v(t) is the speed at time t.
[0038] Step 4: Based on the predicted speed sequence, the remaining power at the current moment, and the predicted vehicle required power sequence, use the energy management control model to obtain a predicted remaining power sequence and a predicted engine output power sequence.
[0039] Specifically, the energy management control model includes: a remaining power management control model and a power management control model; the remaining power management control model is used to determine the predicted remaining power sequence, and the power management control model is used to determine the predicted engine output power sequence. Among them, the power management control model is obtained by training the TD3 network using a reinforcement learning algorithm. The TD3 network includes a strategy network, a target strategy network, a valuation network, and a target valuation network; the valuation network and the target valuation network both contain two Q networks.
[0040] Deep reinforcement learning is a powerful artificial intelligence technology that cleverly combines the perception advantages of deep learning with the decision-making ability of reinforcement learning. The intelligent agent is the subject that can perceive the environment and make decisions. With the help of deep reinforcement learning algorithms, a strategy is carefully learned that accurately describes what actions should be taken in various different states. Based on the strategy and the current state, the intelligent agent decisively selects an action and implements it. The environment then responds to the action of the intelligent agent, jumps to a new state, and immediately gives a corresponding reward signal. The implementation of this process is closely related to the Markov process.
[0041] A remaining power management and control model is established. For the remaining power at time t+1 in the predicted remaining power sequence, the change in the battery at time t is calculated by the engine output power at time t, and then the remaining power at time t+1 is calculated by the remaining power at time t and the change in the battery at time t. The expression of the remaining power management and control model is as follows.
[0042] SOC(t+1)=SOC(t)+F′(t,P eng ,C).
[0043] Among them, SOC(t+1) is the remaining power at time t+1; SOC(t) is the remaining power at time t; F′(·) is the battery change at time t; P eng is the engine output power, and C is the battery capacity.
[0044] The energy management problem is constructed as a Markov process as follows: The state variables and action variables of the Markov process are defined, and the state variables are defined as follows.
[0045] s(t)={v(t),SOC(t),P dem (t)}.
[0046] The action variables are defined as follows.
[0047] a(t)={P eng (t)}.
[0048] Among them, P eng(t) represents the engine output power at time t.
[0049] In addition, the above variables are constrained as follows.
[0050] SOC min ≤SOC(t)≤SOC max 。
[0051] P eng,min ≤P eng (t)≤P eng,max 。
[0052] Among them, SOC min represents the lower limit value of SOC, and SOC max represents the upper limit value of SOC; P eng,min represents the lower limit value of the engine output power, and P eng,max represents the upper limit value of the engine output power.
[0053] Step 5: Construct an energy management optimization objective function with the minimum real-time energy of the hybrid vehicle as the goal.
[0054] The expression of the energy management optimization objective function is as follows.
[0055]
[0056] Among them, J t is the value of the energy management optimization objective function; N p is the preset time domain; is the instantaneous fuel consumption rate of the hybrid vehicle at time t; f SOC (t) is the penalty function between the remaining power at time t and the reference remaining power; SOC ref is the reference remaining power; SOC(t) is the remaining power at time t; ξ is the penalty coefficient.
[0057] Specifically, the vehicle speed prediction information during the actual driving process of the vehicle through the LSTM network within the N p time domain is output to the energy management control model to obtain a predicted remaining power sequence: and a predicted engine output power sequence Based on J tBased on the results, determine the instantaneous fuel consumption rate and the minimum remaining power of the minimum hybrid vehicle for vehicle control. Subsequently, take the maximum value in the predicted engine output power sequence as the upper limit value for solving the MPC problem using the classical DP algorithm, and take the minimum value in the reference action sequence as the lower limit value. Use the predicted remaining power sequence as the target tracking sequence. Solve and output the engine output power sequence of the hybrid vehicle, but only output the first control element as the control instruction, and update the vehicle state. Feed back the updated state (SOC) to the controller as the initial state value for iterative optimization at the next moment. Repeat this process to achieve real-time optimal energy management.
[0058] Step 6: Based on the predicted remaining power sequence and the predicted engine output power sequence, use the energy management optimization objective function to determine the optimal remaining power and the optimal engine output power.
[0059] Step 7: Control the hybrid vehicle to travel through the optimal engine output power to achieve real-time energy management of the hybrid vehicle.
[0060] In an exemplary embodiment, as Figure 2 shown, the training of the speed prediction model in Step 2 specifically includes the following Steps 21 to 24.
[0061] Step 21: Use a GPS device to collect the historical speed data of the vehicle to construct a training set and a prediction set.
[0062] Step 22: Establish an LSTM neural network speed prediction model, and input the sample historical speed sequence from the input end. The expression of the sample historical speed sequence is as follows.
[0063]
[0064] where V his is the sample historical speed sequence; is the sample historical speed at time t - N p ; is the sample historical speed at time t - N p +1; V t is the sample historical speed at time t; N p is the time domain of the historical speed sequence.
[0065] Step 23: Through the fully connected layer, predict the output to obtain the sample predicted speed sequence. The expression of the sample predicted speed sequence is as follows.
[0066]
[0067] where V pred is the sample predicted speed sequence; is the sample prediction speed at time t+1; is the sample prediction speed at time t+2; is at t+N p the sample prediction speed at the moment.
[0068] Step 24: Train the LSTM neural network through the sample historical speed sequence and the sample prediction speed sequence to obtain a speed prediction model; the input-output function relationship of the LSTM neural network can be expressed as follows.
[0069]
[0070] where, is the speed prediction model.
[0071] After obtaining the speed prediction model, use the speed prediction model for prediction to obtain prediction data. The results of the prediction data and the actual data are as Figure 3 shown, and the prediction data and the actual data are basically the same.
[0072] In an exemplary embodiment, the TD3 network is trained using a reinforcement learning algorithm in step 4, which specifically includes the following steps 41 to 47.
[0073] Step 41: Input the state variables at the current moment into the policy network to obtain the action variables at the current moment; the state variables include speed, remaining battery power, and vehicle required power; the action variables represent the engine output power.
[0074] Step 42: Add noise to the action variables at the current moment to obtain the action variables at the current moment after adding noise.
[0075] Step 43: Determine the corresponding reward value at the current moment and the state variables at the next moment based on the action variables at the current moment after adding noise, and use the target policy network to determine the action variables at the next moment.
[0076] Further, after determining the action variables at the next moment, noise can be added to the action variables at the next moment.
[0077] Specifically, define the feedback reward of the Markov process, with the goal of minimizing the fuel economy cost and the SOC fluctuation. The calculation formula of the reward value at the current moment is as follows.
[0078]
[0079] where, r(t) is the sample reward at time t; β1 and β2 are the weight coefficients of the two reward terms; is the instantaneous fuel consumption rate of the hybrid vehicle at time t; SOC(t) is the remaining battery power of the hybrid vehicle at time t; SOCref is the remaining battery power; P eng is the engine output power of the hybrid vehicle.
[0080] The agent continuously optimizes its own strategy through repeated attempts and corrections. The ultimate goal is to find an optimal strategy that maximizes the expected cumulative reward starting from any initial state. Deep reinforcement learning enables the agent to make optimal decisions in a Markov process by learning the optimal strategy and value function, thus striving to maximize the long-term cumulative reward. The formula is as follows.
[0081]
[0082] where r t is the reward obtained at each step, λ ∈ (0, 1) is the discount factor, representing the discount coefficient of the reward, and R(s) is the set of all rewards.
[0083] Step 44: Input the state variable at the current moment and the action variable at the current moment into the evaluation network to obtain two value functions, and use the minimum value function to update the policy network as the policy loss function.
[0084] Step 45: Input the state variable at the next moment and the action variable at the next moment into the target evaluation network to obtain two target value functions, and calculate the evaluation loss function based on the minimum target value function and the reward value at the current moment to update the evaluation network.
[0085] Specifically, the calculation formula of the evaluation loss function is as follows.
[0086]
[0087] where Loss[·] is the evaluation loss function; θ i is the parameter of the i-th evaluation network; y is the target value function value; s(t) is the state variable at time t; a(t) is the action variable at time t; is the minimum value function in the evaluation network; r(t) is the sample reward at the current moment; γ is the discount factor; is the minimum value function in the target evaluation network; s(t + 1) is the state variable at the next moment; is the action variable at the next moment after adding noise.
[0088] Step 46: Update the parameters in the target policy network according to the parameters in the updated policy network.
[0089] Step 47: Update the parameters in the target evaluation network according to the parameters in the updated evaluation network to complete the training of the TD3 network. The trained policy network is used as the power management control model.
[0090] Regarding the training of the TD3 network using the reinforcement learning algorithm in step 3, a specific embodiment is provided, taking time t as an example. The TD3 network includes two types of networks: a policy network (actor network) and a valuation network (critic network); the policy network represents the mapping from the state variable s(t) to the action variable a(t), and the valuation network further evaluates the former, so it is also called the value function Q, which can be expressed as follows.
[0091]
[0092] The above formula reflects that at time t, the value function Q corresponding to {s(t), a(t)} is obtained, and the policy network (critic) is continuously updated in the direction of increasing Q, so as to approximate the optimal policy. In addition, the error between the value function Q and the true value is continuously iteratively corrected through the time difference error based on the Q function.
[0093] When Q is continuously iterated in the increasing direction, it often leads to the estimated value of Q being much larger than the true value, which will prolong the algorithm convergence time. Therefore, the valuation network is decomposed into two channels to output two value functions and Each time, the minimum value between the two is used as the output result of the valuation network.
[0094] Then, a target valuation network and a target policy network are constructed. The two network structures are the same as those of the policy network and the valuation network, and the parameters are passed from the original policy network and valuation network to the target network through soft updates to control the target network update rate; in addition, the TD3 network stores the quadruple experience [s(t), a(t), r(t), s(t + 1)] obtained during the interaction process until the amount of experience in the experience pool reaches a certain quantity, and then randomly extracts the quadruple to update the network parameters.
[0095] The specific network parameter update process is as follows.
[0096] (1), Initialize the parameters ψ of the policy network, the parameters ψ′ of the target policy network, the parameters θ1 and θ2 of the valuation network, and the parameters θ1′ and θ2′ of the target valuation network respectively.
[0097] (2), Take the state variable s(t) feedback by the environment in real time as the input, and the policy network outputs the corresponding action variable a(t).
[0098] (3), To ensure the robustness of the policy, the agent is allowed to explore the environment to a certain extent. Therefore, a certain degree of noise is usually added to the action variable a(t), and the formula is as follows.
[0099]
[0100] Among them, is the action variable at time t after adding noise, a(t) is the action variable at time t, ρ is the noise attenuation coefficient, ε ∼ N(0, σ) indicates that the noise parameter ε follows a normal distribution with a mean of 0 and a variance of σ, N t is the current training batch, N total is the total number of training batches. As the training process progresses, the noise gradually weakens, that is, the action variable is gradually determined.
[0101] (4), The agent takes the action variable After that, the state variable s(t + 1) and the reward r(t) at the next moment are obtained, and the obtained four-tuple experience is stored. When the number of experiences in the experience pool reaches a certain amount, an experience array [s′(t), a′(t), r′(t), s′(t + 1)] is randomly selected from the experience pool for network update.
[0102] (5), Further obtain the action variable a′(t + 1) of the state variable s′(t + 1) through the target policy network, and add a certain amount of noise to the action variable according to the operation in (3) to obtain the action variable for random search The formula is as follows.
[0103]
[0104] Similarly, the action variable for random search of s′(t) can be obtained, and the formula is as follows.
[0105]
[0106] (6), Obtain the value function values corresponding to the actions of the two evaluation networks under the state variable s′(t), which are respectively denoted as: and
[0107] (7), Obtain the target value function corresponding to the random action under the state variable s′(t + 1) according to the target evaluation network and take the smaller value function of the two, which is denoted as: At the same time, calculate the target value function value according to the Bellman equation, and the formula is as follows.
[0108]
[0109] Among them, y is the target value function value; r′(t) is the state variable for random search at time t; γ is the discount factor used to adjust the weight of future rewards, and its value range is [0, 1], which determines the degree of importance that the agent attaches to immediate rewards and future rewards; is the minimum value function in the target evaluation network; s′(t + 1) is the state variable for random search at time t + 1; is the action variable for random search at time t+1.
[0110] (8) Update the parameters of the value network by minimizing the loss function, and the formula is as follows.
[0111]
[0112] where Loss[·] is the value loss function; θ i is the parameter of the i-th value network; is the minimum value function in the value network.
[0113] (9) Update the parameters of the policy network by minimizing the loss function, and the formula is as follows.
[0114]
[0115] (10) Update the parameters ψ′ of the target policy network and the parameters θ i ′ in the target value network by soft update, and the formula is as follows.
[0116] ψ″ = (1 - τ)·ψ′ + τ·ψ.
[0117] θ i ″ = (1 - τ)·θ i ′ + τ·θ i .
[0118] where ψ″ is the parameter of the updated target policy network; θ i ″ is the parameter in the updated target value network; τ ∈ (0, 1) is the soft update coefficient. When τ is closer to 1, the update speed of the parameters of the policy network and the value network towards the target network parameters is faster. This means that the parameters ψ′ of the target policy network and the parameters θ i ′ will approach the parameters ψ, θ of the main network i , but not completely replace them, but smoothly transition with τ as the weight. This can ensure that the parameter changes of the target network are more stable and reduce the fluctuations in the learning process.
[0119] The beneficial effects of the real-time energy management method for hybrid electric vehicles proposed in this application are mainly manifested in the following three aspects.
[0120] (1) A new energy management method combining model predictive control and deep reinforcement learning is proposed, thus solving the problems that the existing technology cannot adapt to the complex and changeable working conditions in actual driving and the real-time ability is insufficient.
[0121] (2) By screening the predicted remaining power sequence and the predicted engine output power sequence, the optimal engine output power and the optimal remaining power are obtained, and the real-time prediction is completed, solving the problem that the rules formulated based on manual experience are no longer applicable when the road conditions change.
[0122] (3) With the goal of minimizing the real-time energy of a hybrid vehicle, the energy management optimization objective function is optimized, and the predicted engine output power sequence within the future driving cycle is calculated to obtain the optimal engine output power for controlling the driving of the hybrid vehicle, thereby improving the energy utilization efficiency.
[0123] Based on the same inventive concept, the embodiments of the present application also provide a hybrid vehicle real-time energy management system for implementing the above-mentioned hybrid vehicle real-time energy management system. The implementation solutions provided by the device to solve problems are similar to the implementation solutions described in the above method. Therefore, the specific limitations in one or more embodiments of the hybrid vehicle real-time energy management system provided below can refer to the limitations on the hybrid vehicle real-time energy management method in the above text, and will not be repeated here.
[0124] In an exemplary embodiment, a hybrid vehicle real-time energy management system is provided, which includes the following modules.
[0125] An acquisition module for acquiring the historical speed sequence of the hybrid vehicle and the remaining power at the current moment; the historical speed sequence is a set of speeds from the historical moment to the current moment.
[0126] A speed prediction module for inputting the historical speed sequence into a speed prediction model to obtain a predicted speed sequence; the speed prediction model is trained by speed sample data for an LSTM network.
[0127] A vehicle required power prediction module for inputting the predicted speed sequence into a vehicle required power prediction model to obtain a predicted vehicle required power sequence.
[0128] An energy management control module for obtaining a predicted remaining power sequence and a predicted engine output power sequence by using an energy management control model based on the predicted speed sequence, the remaining power at the current moment, and the predicted vehicle required power sequence.
[0129] An optimization objective function construction module for constructing an energy management optimization objective function with the goal of minimizing the real-time energy of the hybrid vehicle.
[0130] A determination module for determining the optimal remaining power and the optimal engine output power by using the energy management optimization objective function based on the predicted remaining power sequence and the predicted engine output power sequence.
[0131] A control module is used to control the driving of a hybrid vehicle through the optimal engine output power to achieve real-time energy management of the hybrid vehicle.
[0132] In an exemplary embodiment, a computer device is provided. The computer device can be a server or a terminal, and its internal structure diagram can be as Figure 4 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the optimal remaining battery power and the optimal engine output power. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a method for real-time energy management of a hybrid vehicle.
[0133] Those skilled in the art can understand that Figure 4 the structure shown in
[0134] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout. In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0135] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tapes, floppy disks, flash memories, optical memories, high-density embedded non-volatile memories, resistive random access memories (ReRAM), magnetoresistive random access memories (MRAM), ferroelectric random access memories (FRAM), phase change memories (PCM), graphene memories, etc. Volatile memories can include random access memory (RAM) or external cache memories, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0136] The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logics, data processing logics based on quantum computing, etc., without limitation.
[0137] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.
[0138] In this article, specific examples are used to elaborate on the principles and implementation manners of the present application. The descriptions of the above embodiments are only used to help understand the methods and core ideas of the present application; at the same time, for those of ordinary skill in the art, according to the ideas of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A real-time energy management method for a hybrid vehicle, characterized in that, The real-time energy management method for the hybrid vehicle includes: Obtaining the historical speed sequence of the hybrid vehicle and the remaining battery power at the current moment; the historical speed sequence is a set of speeds from the historical moment to the current moment; Inputting the historical speed sequence into the speed prediction model to obtain a predicted speed sequence; the speed prediction model is obtained by training an LSTM network with speed sample data; Inputting the predicted speed sequence into the vehicle required power prediction model to obtain a predicted vehicle required power sequence; Based on the predicted speed sequence, the remaining battery power at the current moment, and the predicted vehicle required power sequence, using the energy management control model, obtaining a predicted remaining battery power sequence and a predicted engine output power sequence; Constructing an energy management optimization objective function with the minimum real-time energy of the hybrid vehicle as the target; Based on the predicted remaining battery power sequence and the predicted engine output power sequence, using the energy management optimization objective function to determine the optimal remaining battery power and the optimal engine output power; Controlling the driving of the hybrid vehicle through the optimal engine output power to achieve real-time energy management of the hybrid vehicle.
2. The real-time energy management method for a hybrid vehicle according to claim 1, wherein, The energy management control model includes: a remaining battery power management control model and a power management control model; the remaining battery power management control model is used to determine the predicted remaining battery power sequence, and the power management control model is used to determine the predicted engine output power sequence.
3. The real-time energy management method for a hybrid vehicle according to claim 2, characterized in that The power management control model is obtained by training a TD3 network using a reinforcement learning algorithm.
4. The real-time energy management method for a hybrid electric vehicle according to claim 3, wherein The TD3 network includes a policy network, a target policy network, a value network, and a target value network; both the value network and the target value network contain two Q networks; Training the TD3 network using a reinforcement learning algorithm specifically includes: Inputting the state variables at the current moment into the policy network to obtain the action variables at the current moment; the state variables include speed, remaining battery power, and vehicle required power; the action variables represent the engine output power; Adding noise to the action variables at the current moment to obtain the action variables at the current moment after adding noise; Determining the corresponding reward value at the current moment and the state variables at the next moment based on the action variables at the current moment after adding noise, and using the target policy network to determine the action variables at the next moment; Inputting the state variables at the current moment and the action variables at the current moment into the value network to obtain two value functions, and using the minimum value function as the policy loss function to update the policy network; Inputting the state variables at the next moment and the action variables at the next moment into the target value network to obtain two target value functions, and calculating the value loss function based on the minimum target value function and the reward value at the current moment to update the value network; Updating the parameters in the target policy network according to the parameters in the updated policy network; Updating the parameters in the target value network according to the parameters in the updated value network to complete the training of the TD3 network, and the trained policy network is used as the power management control model.
5. The real-time energy management method for a hybrid vehicle according to claim 4, characterized in that After using the target policy network to determine the action variables at the next moment, it further includes: adding noise to the action variables at the next moment.
6. The real-time energy management method for a hybrid vehicle according to claim 4, characterized in that, The calculation formula for the reward value at the current moment is as follows: Among them, r(t) is the sample reward at time t; β1 and β2 are the weight coefficients of the two reward terms; is the instantaneous fuel consumption rate of the hybrid vehicle at time t; SOC(t) is the remaining power of the hybrid vehicle at time t; SOC ref is the reference remaining power; P eng is the engine output power of the hybrid vehicle.
7. The real-time energy management method for a hybrid vehicle according to claim 4, wherein The calculation formula for the valuation loss function is as follows: where Loss[·] is the valuation loss function; θ i is the parameter of the i-th valuation network; y is the target value function value; s(t) is the state variable at time t; a(t) is the action variable at time t; is the minimum value function in the valuation network; r(t) is the sample reward at the current time; γ is the discount factor; is the minimum value function in the target valuation network; s(t + 1) is the state variable at the next time; is the action variable at the next time after adding noise.
8. The real-time energy management method for a hybrid vehicle according to claim 1, wherein, The expression of the energy management optimization objective function is as follows: Among them, J t is the value of the energy management optimization objective function; N p is the preset time domain; is the instantaneous fuel consumption rate of the hybrid vehicle at time t; f SOC (t) is the penalty function between the remaining battery charge at time t and the reference remaining battery charge; SOC ref is the reference remaining battery charge; SOC(t) is the remaining battery charge at time t; ξ is the penalty coefficient.
9. A real-time energy management system for a hybrid vehicle, characterized in that, Applied to the real-time energy management method for a hybrid vehicle described in any one of claims 1-8, the real-time energy management system for a hybrid vehicle includes: An acquisition module, configured to acquire the historical speed sequence of the hybrid vehicle and the remaining power at the current moment; the historical speed sequence is a set of speeds from the historical moment to the current moment; A speed prediction module, configured to input the historical speed sequence into a speed prediction model to obtain a predicted speed sequence; the speed prediction model is obtained by training an LSTM network with speed sample data; A predicted vehicle power demand module, configured to input the predicted speed sequence into a predicted vehicle power demand model to obtain a predicted vehicle power demand sequence; An energy management control module, configured to, based on the predicted speed sequence, the remaining power at the current moment, and the predicted vehicle power demand sequence, adopt an energy management control model to obtain a predicted remaining power sequence and a predicted engine output power sequence; An optimization objective function construction module, configured to construct an energy management optimization objective function with the minimum real-time energy of the hybrid vehicle as the objective; A determination module, configured to, based on the predicted remaining power sequence and the predicted engine output power sequence, adopt the energy management optimization objective function to determine the optimal remaining power and the optimal engine output power; A control module, configured to control the running of the hybrid vehicle through the optimal engine output power to achieve real-time energy management of the hybrid vehicle.
10. A computer device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the real-time energy management method for a hybrid vehicle described in any one of claims 1-8.
Citation Information
Cited By
Hybrid vehicle energy management method based on traffic flow density perception and reinforcement learning
CN121034086A