Vehicle multi-power coordination control method, equipment, medium and product
The vehicle multi-power coordination control method addresses stability and efficiency issues in hybrid vehicles by using real-time vehicle speed and battery state as variables and a trained offline controller to optimize engine power, enhancing dynamic stability and energy efficiency.
Patent Information
- Application Number
- CN202510687346.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-07-15
AI Technical Summary
Traditional vehicle multi-power coordination control methods have low stability and energy utilization in dynamic environments, especially in heavy-duty hybrid vehicles, which fail to effectively consider the response speed and dynamic characteristics of system components to control signals.
Define the real-time speed of the vehicle and the remaining battery power as the state quantity, combine the engine output power as the action variable, and build an offline energy management controller through deep reinforcement learning and gray wolf optimization algorithm to optimize the engine output power to achieve dynamic coordinated control.
It improves the control quality and energy utilization rate of the dynamic process of the vehicle operation, and improves the stability and energy utilization efficiency of the vehicle during driving.
Smart Images

Figure SMS_1 
Figure SMS_4 
Figure SMS_5
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of deep learning, and particularly to a method, device, medium and product for coordinated control of multiple vehicle power sources. Background Art
[0002] While the automotive industry dominated by traditional fuel vehicles has been booming, the problem of energy consumption has become increasingly prominent. Against this backdrop, hybrid electric vehicles, which integrate the advantages of pure electric vehicles and fuel vehicles, have received extensive attention as an effective short-term energy alternative and have been widely promoted. Due to the presence of multiple power sources, the power distribution method between the engine and the motor, that is, the energy management method, is crucial for fully exploring the vehicle's energy-saving control potential, optimizing the vehicle's power performance, and reducing pollutant emissions.
[0003] The current traditional related technologies have the following deficiencies. For example, rule-based strategies are extremely dependent on human experience. When the road conditions change, new rules need to be formulated, resulting in huge costs. Dynamic Programming (DP), as a typical optimization-based energy management method, requires prior knowledge of the entire driving cycle information, seriously affecting real-time performance and robustness. In addition, most current energy management strategies focus on the steady-state conditions of the vehicle and optimize the system under a long time scale, without considering the dynamic characteristics such as the response speed of system components to control signals, especially in heavy hybrid vehicles.
[0004] With the development of artificial intelligence technology, methods based on reinforcement learning have become popular in the industry due to their strong self-learning ability. However, in the research of traditional reinforcement learning-based methods, most do not consider the uncertainties during driving caused by dynamic characteristics such as the response speed of system components to control signals in future driving cycles. Therefore, the control quality of the vehicle's dynamic operation process cannot be improved, resulting in low stability and low energy utilization rate during vehicle driving. Summary of the Invention
[0005] The purpose of the present application is to provide a method, device, medium and product for coordinated control of multiple vehicle power sources, which can solve the problems of low stability and low energy utilization rate during vehicle driving in traditional related technologies.
[0006] To achieve the above purpose, the present application provides the following solutions.
[0007] In a first aspect, the present application provides a method for coordinated control of multiple vehicle power sources, including: defining the real-time vehicle speed and the remaining battery power as state variables, and defining the engine output power as an action variable; the vehicle being a heavy hybrid vehicle; determining the reward for each moment based on the instantaneous fuel consumption rate and the remaining power fluctuation degree of the vehicle at each moment; taking the state variables at each moment, the action variables at each moment, and the rewards at each moment as the dynamic characteristics during the dynamic driving process of the vehicle, and taking the dynamic characteristics as the real-time state of the vehicle; during the driving process of the vehicle in a preset time period, inputting the real-time state of the vehicle within the preset time period into a trained offline energy management controller to determine the reference engine output power at each moment within the preset time period; the trained offline energy management controller being constructed based on the historical real-time state of the vehicle; determining the reference engine speed at each moment based on the reference engine output power at each moment within the preset time period; determining the target engine output power at each moment according to the reference engine speed at each moment and the grey wolf optimization algorithm; taking the target engine output power at each moment as the optimal control action, and controlling the vehicle to drive within the preset time period based on the optimal control action.
[0008] In a second aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the method for coordinated control of multiple vehicle power sources described above.
[0009] In a third aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the method for coordinated control of multiple vehicle power sources described above.
[0010] In a fourth aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the method for coordinated control of multiple vehicle power sources described above.
[0011] According to the specific embodiments provided by the present application, the following technical effects are disclosed:
[0012] The present application provides a method, device, medium and product for coordinated control of multiple vehicle power sources. The present application first defines the real-time vehicle speed and the remaining battery power as state variables, and the engine output power as an action variable, and determines the reward for each moment based on the instantaneous fuel consumption rate and the degree of remaining power fluctuation at each moment. Then, the state variables, action variables and rewards at each moment are used as the dynamic characteristics during the dynamic driving process of the vehicle, and are input into the trained offline energy management controller to determine the reference engine output power at each moment. That is, the present application combines the dynamic characteristics such as the output power of the engine, the instantaneous fuel consumption rate of the motor and the remaining battery power in the system components of the vehicle around the dynamic driving process of the vehicle, so as to better realize the coordinated control of the heavy hybrid vehicle using two power sources, namely fuel and battery, in the future. Considering the dynamic characteristics of the vehicle during the dynamic driving process, a more accurate engine parameter output power that is more in line with the actual driving process of the vehicle is obtained. Finally, based on the engine parameter output power and the Grey Wolf Optimization Algorithm within a preset time period, the target engine output power at each moment within the preset time period is obtained, and the vehicle driving is controlled by the target engine output power at each moment, improving the control quality of the vehicle's dynamic operation process and achieving high stability during the vehicle driving process. At the same time, in the long-term scale optimization control of the vehicle, the coordinated control of the two power sources, fuel and battery, is combined to improve the energy utilization rate of the vehicle. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0014] Figure 1 It is a schematic flowchart of a method for coordinated control of multiple vehicle power sources provided in an embodiment of the present application.
[0015] Figure 2 It is a schematic diagram of the principle of a method for coordinated control of multiple vehicle power sources provided in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0016] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0017] To make the above objects, features, and advantages of the present application more apparent and understandable, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0018] As Figure 1 shown, the present application provides a method for coordinated control of multiple vehicle power sources, including steps 101 to 107.
[0019] Step 101: Define the real-time vehicle speed and the remaining battery power at the current moment as state variables, and define the engine output power as an action variable; the vehicle is a heavy hybrid vehicle.
[0020] In practical applications, a hybrid vehicle model is established. For the energy management problem, a Markov process is constructed. The state variables are defined as: s(t) = {v(t), soc(t)}, including the real-time vehicle speed v(t) and the remaining battery power soc(t) at the current moment; the action variable is a(t) = {p eng (t)}, including the engine output power p eng (t) at each moment. At the same time, the upper and lower limits of the above variables are constrained: respectively: soc min ≤soc(t)≤soc max , p eng,min ≤p eng (t)≤p eng,max . soc max , soc min represent the upper and lower limits of soc respectively, and p eng,max , p eng,min represent the upper and lower limits of the engine output power respectively; soc ∈ (0, 1). The above-mentioned upper and lower limits are all determined based on the minimum and maximum values recorded in the historical state variables and historical action variables.
[0021] Step 102: Determine the reward at each moment based on the instantaneous fuel consumption rate and the degree of remaining power fluctuation of the vehicle at each moment.
[0022] In practical applications, vehicle historical state information is collected according to the vehicle GPS system, load monitoring system, and battery management system. It includes the real-time vehicle speed v(t) and the real-time remaining battery power soc(t).
[0023] In practical applications, in the Markov decision process of defining the energy management problem of a hybrid vehicle, the goal of the feedback reward function is to minimize the fuel economy cost and the SOC fluctuation, which is defined as:
[0024]
[0025] Among them, is the instantaneous fuel consumption rate of the vehicle; SOCref is the reference value of the remaining battery power; SOC(t) is the remaining battery power at the current moment; r(t) is the reward at the current moment; α1 and α2 are both weight coefficients, which are used to balance the weights between the two optimization objectives of fuel consumption and SOC fluctuation respectively.
[0026] Among them, the fuel consumption rate is determined by the engine map. Specifically, the engine map details the fuel consumption of the engine under different working conditions. By querying this map, the instantaneous fuel consumption rate of the vehicle when the engine is in the current operating state can be obtained. SOC ref ∈(0, 1).
[0027] Specifically, the reward function includes the instantaneous fuel consumption rate of the vehicle as a negative reward term to prompt the agent to reduce fuel consumption; at the same time, the degree of SOC fluctuation is also used as a negative reward term to encourage the agent to maintain the stability of SOC. The agent continuously tries and corrects errors, and continuously optimizes its own energy management strategy to pursue the maximization of the long-term cumulative reward, that is:
[0028] Among them, the reward obtained at each moment is r t , λ ∈ (0, 1) is the discount factor, representing the discount coefficient of the reward, and R(s) is the set of all rewards.
[0029] Step 103: Take the state variables at each moment, the action variables at each moment, and the rewards at each moment as the dynamic characteristics during the dynamic driving process of the vehicle, and take the dynamic characteristics as the real-time state of the vehicle.
[0030] In some embodiments, the real-time state of the vehicle further includes the engine output speed at each moment; according to the engine output speed at each moment and the second-order inertia model, the engine output power at each moment is determined; the expression of the second-order inertia model is:
[0031] J e ω″(t) + Bω′(t) = T input (t) - T load (t).
[0032] Among them, J e is the engine moment of inertia, B is the damping coefficient, ω′(t) is the first derivative of the engine output speed at the current moment; ω″(t) is the second derivative of the engine output speed at the current moment; T input (t) is the engine input torque; T load (t) is the engine load torque.
[0033] Specifically, most energy management strategies do not fully consider the inertia and friction characteristics of components. Power coordination takes into account the differences in responsiveness between components inside heavy vehicles, combines the response characteristics of each component, and focuses on establishing a typical component model. The second-order inertia model is a classic dynamic system model, so the second-order inertia model is used to characterize the engine response characteristics. The second-order inertia model considers the inertia and damping characteristics of the engine and can more accurately describe the dynamic behavior of the engine. The expression of the second-order inertia model is as follows:
[0034] J e ω″(t)+Bω′(t)=T input (t)-T load (t).
[0035] T load (t)=T transmission (t)+T rolling (t).
[0036] Among them, T transmission (t) and T rolling (t) are the resistance torque of the vehicle's transmission system and the rolling resistance torque of the wheels respectively.
[0037] Due to the dynamic characteristics of the engine, the actual output torque will be affected by the input torque and the load torque.
[0038] Among them, the real-time state of the vehicle can also include the engine output speed at each moment.
[0039] Among them, the PID control method is used to achieve precise control of the engine speed value. To achieve closed-loop control of the engine speed and ensure that the output of the engine speed can quickly and accurately track the set value. Through the PID controller, according to the engine speed error e(t) = ω ref (t)-ω(t) to adjust the engine input torque T input (t) to maintain the set target speed ω ref (t); where ω(t) is the engine output speed at the current moment, and the adjustment of the input torque by the PID controller can be expressed as:
[0040]
[0041] Among them, K p is the proportional coefficient, which determines the intensity of the immediate reaction to the error; K i is the integral coefficient, which is used to eliminate the steady-state error; K d is the differential coefficient, which is used to predict the change trend of the error, enhance the response speed of the system and the ability to suppress overshoot.
[0042] T out (t)=T input(t) - T load (t).
[0043] Among them, T out (t) represents the actual output torque.
[0044] Calculate the actual output power P of the engine at the current moment out of (t) as follows:
[0045]
[0046] Take the actual output power of the engine at each moment as the output power of the engine at each moment.
[0047] In some embodiments, before step 104, steps 201 - 203 are further included.
[0048] Step 201: Based on the deep deterministic policy gradient algorithm, construct an offline energy management controller; the offline energy management controller includes a policy network, a value network, a target policy network, and a target value network.
[0049] Step 202: Based on the historical vehicle real - time state and the offline energy management controller, determine multiple experience sample data and store them in an experience pool; each experience sample data includes the historical state quantity at the current moment, the historical action variable at the current moment, the historical reward at the current moment, and the historical state quantity at the next moment; the historical state quantity includes the historical vehicle real - time speed and the historical real - time remaining battery power; the historical action variable includes the historical engine output power.
[0050] Step 203: Train the offline energy management controller according to the multiple experience sample data to determine the trained offline energy management controller.
[0051] In some embodiments, step 202 specifically includes: input the historical state quantity at the current moment into the policy network to obtain the historical action variable at the current moment; take the historical state quantity at the current moment, the historical action variable at the current moment, the historical state quantity at the next moment, and the historical reward at the current moment as a quadruple at the current moment; the quadruple is the experience sample data; input the historical state quantity at the next moment into the target policy network to obtain the historical action variable at the next moment; take the historical state quantity at the next moment as the historical state quantity at the current moment, and return to the step "take the historical state quantity at the current moment, the historical action variable at the current moment, the historical state quantity at the next moment, and the reward at the current moment as a quadruple" to determine the quadruple at the next moment; take the quadruples at each moment as multiple experience sample data and store them in the experience pool.
[0052] Specifically, during the training process, two dual experience replay buffers D1 and D2 are introduced to improve the training efficiency and ensure the diversity and quality of the experience sample data. They are used to store all historical sample data and high-feedback reward historical experience sample data respectively. [s(t), a(t), r(t), s(t+1)] represents the experience sample data, where s(t) represents the vehicle state at time t, a(t) represents the action variable at time t, r(t) represents the reward obtained after taking action a(t) in state s(t), and s(t+1) represents the vehicle state at time t+1. Here, time t is the current time, and time t+1 is the next time.
[0053] In traditional related technologies, only a single experience replay buffer can be used, where all experience data is stored in one buffer. High correlations may occur between samples during random sampling, especially during consecutive sampling. By storing experience data in two different buffers, the correlations between samples can be further reduced. This design can ensure that the data sources are more diverse during each sampling, thereby improving the stability and efficiency of training.
[0054] During sampling in this application, experience sample data is drawn from the two experience replay buffers respectively, as Figure 2 shown, for deep reinforcement learning network training. If the data in the experience replay buffer reaches the upper limit, the old data is deleted in chronological order and new data is then stored. Otherwise, the new data can be directly stored.
[0055] In some embodiments, step 203 specifically includes: Based on the valuation network, evaluating the historical state quantity at the current moment and the historical action variable at the current moment to determine the action state estimation value at the current moment; Based on the target value network, evaluating the historical state quantity at the next moment and the action variable at the next moment to determine the target valuation of the action state at the next moment; Based on the historical reward at the current moment and the target valuation of the action state at the next moment, determining the updated target valuation of the action state at the next moment, and using the updated target valuation of the action state at the next moment as the target Q value; Determining the difference between the target Q value and the action state estimation value at the current moment, and minimizing the difference to train the parameters of the valuation network to determine the trained valuation network; Based on the gradient of the historical action variable at the current moment of the valuation network and the parameters of the policy network at the current moment, determining the gradient of the policy network at the current moment, and updating the parameters of the policy network with the goal of minimizing the gradient of the policy network at the current moment to determine the trained policy network; According to the parameters of the valuation network, the parameters of the policy network, and the method of soft update, updating the parameters of the target valuation network and the parameters of the target policy network, and determining the trained target valuation network and the target policy network; Using the trained policy network, the trained valuation network, the trained target policy network, and the trained target valuation network as the trained offline energy management controller.
[0056] In some embodiments, the calculation formula for the gradient of the policy network at the current moment is:
[0057]
[0058] Where, is the gradient of the policy network at the current moment; θ π (t) are the parameters of the policy network; a′(t) is the historical action variable at the current moment; s′(t) is the historical state quantity at the current moment; is the gradient of the valuation network for a′(t) under the policy network π[s, θ π (t)] at the current moment; θ Q (t) are the parameters of the valuation network at the current moment; Q is the valuation network; π is the policy network; is the gradient of the policy network π[s, θ π (t)] with respect to θ π (t).
[0059] Specifically, the introduction of the Deep Deterministic Policy Gradient (DDPG) algorithm is an algorithm that combines deep learning and reinforcement learning, used to solve Markov decision problems in continuous action spaces. The DDPG network model consists of two major networks, an online network (Actor-Critic network) and a target network (Target network), working together. The online network includes a policy network (Actor network) and a value network (Critic network), and the target network includes a target policy network and a target value network.
[0060] The Critic network is responsible for evaluating the quality of the actions generated by the Actor network. It takes the state (i.e., state variables) and actions (i.e., action variables) as inputs and outputs a value function, representing the expected reward of the action in the current state. The target network is used to provide a stable target to help the online network learn more stably. The structure of the target network is the same as that of the online network, but its parameters are updated more slowly to maintain stability. Through continuous learning and updating, the decision-making strategy of the agent is optimized to maximize the long-term cumulative reward. The specific networks can be divided into: policy network π, target policy network π′, value network Q, and target value network Q′.
[0061] The specific steps are as follows.
[0062] Step 1: First, initialize the parameters of the policy network, target policy network, value network, and target value network.
[0063] Step 2: Take the historical state of the real-time feedback s(t) of the environment as the input, and the policy network π with network parameters θ π outputs the corresponding historical action a(t); to better simulate the actual driving environment of the vehicle, introduce random interference λ t to the executed action, so the new action can be expressed as
[0064] Step 3: Then, after the agent takes the action , according to the Markov process defined above, the reward r(t) can be obtained, and at the same time, the state s(t + 1) at the next moment is observed. The quadruple experience is stored in two experience pools. When the experience arrays in the two experience pools reach a certain number, alternately extract the experience arrays [s′(t), a′(t), r′(t), s′(t + 1)] from the two experience pools, and use s′(t + 1) as s(t + 1).
[0065] Step 4. The target policy network π′ selects the action a′(t + 1) at the next moment according to the sampled next state s′(t + 1). Calculate the target Q-value of the action a′(t + 1) in the next state s′(t + 1) according to the target value network Q′, denoted as Q′. Obtain the target Q-value, denoted as y(t):
[0066] y(t) = r(t) + β·Q′[s′(t + 1), a′(t + 1), θ Q′ .
[0067] Among them, r(t) represents the historical reward at the current moment, and β ∈ (0, 1) represents the discount factor, which measures the importance of future rewards.
[0068] 5. Update the parameters of the value network Q by minimizing the loss function. The loss function update function L can be expressed as:
[0069]
[0070] y(t) = r(t) + β·Q′[s′(t + 1), a′(t + 1), θ Q′ .
[0071] Among them, N represents the empirical sample data sampled from the experience pool, that is, the number of samples drawn from the experience pool each time; θ Q and θ Q′ are the value network Q and the target value network Q′ respectively.
[0072] Among them, the loss function well measures the difference between the currently estimated Q-value of the value network and the target Q-value, and this update method is more conducive to the learning process tending to converge.
[0073] The update process of the policy network π is as follows.
[0074] Update by minimizing the policy gradient to maximize the value network Q-value:
[0075]
[0076] Among them, is the gradient of the policy network at time t, indicating the gradient update direction of the policy network parameter θ π (t). represents the gradient of the Q-value function with respect to the action a′(t) under the policy π(s′(t), θ π (t)) at time t, which represents the rate of change of the Q-value when the action changes slightly in the state s′(t). represents the output action π[s, θ π (t)] of the policy network with respect to θ π(t), which represents the rate of change of the output action when the parameters of the policy network change slightly.
[0077] Specifically, the previous gradient formula represents the gradient of the Q-value function with respect to the action under the current policy network. By adjusting the action a′(t), the Q-value (i.e., Q[s,a,θ Q (t)]) can be increased. Then, the latter gradient formula represents the gradient of the action output by the policy network with respect to the parameters of the policy network. By adjusting the parameters of the policy network, the output action can be made more optimal. Multiplying these two gradients gives a gradient with respect to the parameters of the policy network, which indicates how to adjust the parameters of the policy network so that the action taken in the current state can maximize the Q-value.
[0078] According to the policy gradient, the parameters of the policy network are updated with the aim of maximizing the Q-value function. The update formula is:
[0079]
[0080] where θ π (t + 1) are the updated parameters of the policy network, and δ represents the learning rate of the policy network; represents the gradient update direction for the parameters θ π (t) of the policy network.
[0081] The target policy network uses a soft update mechanism to slowly approach the policy network, as follows.
[0082] θ π′ (t + 1) = τ·θ π (t)+(1 - τ)·θ π′ (t).
[0083] where τ is the update coefficient; θ π′ (t) are the parameters of the target policy network at the current time; θ π′ (t + 1) are the parameters of the target policy network at the next time.
[0084] To reduce oscillations during the learning process, the target evaluation network also adopts a soft update mechanism, as follows.
[0085] θ Q′ (t + 1) = τ·θ Q (t)+(1 - τ)·θ Q′ (t).
[0086] where τ is the update coefficient; θ Q′ (t) are the parameters of the target value network at the current time; θ Q′ (t + 1) are the parameters of the target value network at the next time.
[0087] The update of the policy network parameters is achieved through gradient ascent, that is, on the basis of the current parameters, the parameters are updated along the direction of the gradient. This gradient is obtained by calculating the product of the gradient of the Q-value function with respect to the action and the gradient of the output of the policy network with respect to the parameters. The learning rate δ controls the step size of the parameter update to ensure the stability of the update. In this way, the policy network can learn how to output the optimal action to maximize the Q-value function and thus maximize the expected return.
[0088] Repeat the above training process in each training episode until offline training converges or reaches the preset number of training times. Thus, an effective trained offline energy management controller is generated.
[0089] Step 104: During the driving of the vehicle in a preset time period, input the real-time state of the vehicle within the preset time period into the trained offline energy management controller to determine the engine reference output power at each moment within the preset time period; the trained offline energy management controller is constructed based on the historical real-time state of the vehicle.
[0090] Step 105: Based on the engine reference output power at each moment within the preset time period, determine the engine reference speed at each moment.
[0091] Step 106: According to the engine reference speed at each moment and the grey wolf optimization algorithm, determine the engine target output power at each moment.
[0092] In some embodiments, step 106 specifically includes steps 301 - 303.
[0093] Step 301: Set the constraint range of the engine reference speed at each moment.
[0094] Step 302: According to the constraint range and the cost function of the grey wolf optimization algorithm, use the grey wolf optimization algorithm to process the engine reference speed at each moment to determine the engine reference output speed at each moment. The cost function J(t) of the grey wolf optimization algorithm is:
[0095] J(t) = min{ω1|SOC(t) - SOC ref | + ω2|P eng,ref (t) - P eng (t)|}.
[0096] Where, both ω2 and ω1 are cost coefficients; SOC(t) is the real-time remaining battery power at the current moment; SOC ref is the preset reference remaining power reference output; P eng,ref (t) is the engine output reference power at the current moment; P eng(t) is the actual output power of the engine at the current moment; || is the absolute value.
[0097] Step 303: Determine the target output power of the engine at each moment according to the engine reference output speed and power calculation formula at each moment.
[0098] As Figure 2 shown, integrate the trained offline energy management controller and the second-order inertia model into the traditional vehicle model to form a new dynamic model. Based on the new dynamic model, design a coordinated control algorithm and the coordinated control optimization steps are as follows: The Grey Wolf Optimizer (GWO) is an efficient swarm intelligence optimization algorithm suitable for solving various optimization problems. Introduce the classical grey wolf algorithm to ensure that while achieving the energy management goal, the SOC fluctuation is reduced. Therefore, set the cost function of the grey wolf algorithm as:
[0099] J(t) = min{ω1|SOC(t) - SOC ref | + ω2|P eng,ref (t) - P eng (t)|}.
[0100] Among them, ω2 and ω1 represent the cost coefficients. ω1, ω2 ∈ (0, 1), and different values can be set according to different scenarios. For example, in the urban driving condition, frequent starts and stops and low-speed driving may cause large SOC fluctuations, so the value of ω1 can be increased. In the high-speed driving condition, fuel economy is more important, so the value of ω2 can be increased.
[0101] First, define the parameters of the grey wolf optimization algorithm: the population size N w , N w ∈ (30, 100); the maximum number of iterations N max_iter ; N max_iter ∈ (100, 2000); according to the real-time state of the vehicle (the real-time engine speed n eng (t), the real-time vehicle speed, and the real-time remaining battery power SOC), combine the offline EMS controller to obtain the reference output power of the engine, denoted as: P eng,ref (t), and further obtain the reference output speed n eng,ref (t) of the engine according to the optimal economic function of engine fuel consumption, that is:
[0102] n eng,ref (t) = f(P eng,ref (t)).
[0103] Determine the range constraints of the state quantity and the speed input quantity:
[0104] SOC min ≤ SOC(t) ≤ SOCmax .
[0105] n eng,ref (t) - c ≤ n eng (t) ≤ n eng,ref (t) + c, c ∈ (0, n eng,max )
[0106] Where SOC min and SOC max represent the maximum and minimum values of the remaining battery power respectively.
[0107] Next, optimize and solve the values within the engine speed constraint range at each moment using the Grey Wolf Optimization algorithm as follows.
[0108] The engine output speed can be understood as the position of a grey wolf. Evaluate the cost function of each grey wolf according to the dynamic model and cost function, and save the optimal, sub - optimal, and third - optimal grey wolf position solution information (α pos , β pos , δ pos ).
[0109] D α [i] = |C1·α pos - wolves[i]|
[0110] D β [i] = |C2·β pos - wolves[i]|
[0111] D δ [i] = |C3·δ pos - wolves[i]|
[0112] Where C1, C2, C3 are random variables, whose values are between [0, 1], and wolves[i] represents the current grey wolf position solution information, representing a candidate solution. Calculate the new position vector according to the distance variable:
[0113] X1[i] = α pos - A1·D α [i]
[0114] X2[i] = β pos - A2·D β [i]
[0115] X3[i] = δ pos - A3·D δ [i]
[0116] Among them, A1, A2, and A3 are coefficient vectors, and their values are between [-a, a], where a is a parameter that decreases with time, linearly decreasing from an initial value of 2 to 0. As the number of iterations increases, the convergence factor a linearly decreases from 2 to 0.
[0117]
[0118] Among them, t is the current number of iterations, and N max_iter is the maximum number of iterations.
[0119] Update the positions of other grey wolves according to the new position vector:
[0120]
[0121] Repeat the above iteration. After updating the positions of all grey wolves, recalculate the fitness and check whether the fitness value of any grey wolf is better than the current α pos , β pos , δ pos . When the cost function reaches the threshold or the number of iterations reaches a certain value, the current α pos is the optimal solution. Therefore, the reference output speed n eng (t) of the engine at this moment is obtained. The engine output power P out (t) is obtained through the power calculation formula. Repeat the above process until the vehicle stops, and the coordinated control of the two power sources of fuel and battery can be achieved.
[0122] Step 107: Use the engine target output power at each moment as the optimal control action, and control the vehicle to travel within the preset time period based on the optimal control action.
[0123] In response to the challenge of the curse of dimensionality faced by traditional reinforcement learning methods, the present application can better handle high-dimensional state and action spaces in complex environments by introducing a deep neural network, thereby improving energy utilization efficiency. In addition, most current energy management strategies focus on the steady-state conditions of the vehicle and optimize the system control over a long time scale, without considering the dynamic characteristics such as the response speed of each component to control signals, especially in heavy hybrid vehicles. Therefore, based on the dynamic characteristics of each component, a dynamic coordination control algorithm is designed to improve the control quality of the system's dynamic process and ensure the stability of the vehicle during driving.
[0124] In an exemplary embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the above method is implemented.
[0125] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, which implements the above method when executed by a processor.
[0126] In an exemplary embodiment, a computer program product is provided, including a computer program, which implements the above method when executed by a processor.
[0127] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0128] Those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include Read-Only Memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRdM), magnetoresistive random access memory (MRdM), ferroelectric random access memory (FRdM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (Random access Memory, RdM) or external cache memory, etc. By way of illustration and not limitation, RdM can be in various forms, such as static random access memory (Static Random access Memory, SRdM) or dynamic random access memory (Dynamic Random access Memory, DRdM), etc.
[0129] In each of the embodiments provided in the present application, the databases involved may include at least one of a relational database and a non-relational database. The non-relational database may include, but is not limited to, a distributed database based on blockchain, etc. In each of the embodiments provided in the present application, the processor may be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without limitation.
[0130] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0131] Specific examples are used in this article to elaborate on the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A vehicle multi-power coordinated control method, characterized in that, Including: Defining the real-time vehicle speed and the remaining battery power as state variables, and defining the engine output power as an action variable; The vehicle is a heavy-duty hybrid vehicle; Based on the instantaneous fuel consumption rate and the degree of remaining battery power fluctuation at each moment of the vehicle, determining the reward at each moment; Regarding the state variables at each moment, the action variables at each moment, and the rewards at each moment as the dynamic characteristics during the dynamic driving process of the vehicle, and regarding the dynamic characteristics as the real-time state of the vehicle; During the driving process of the vehicle in a preset time period, inputting the real-time state of the vehicle within the preset time period into a trained offline energy management controller to determine the reference engine output power at each moment within the preset time period; the trained offline energy management controller is constructed based on the historical real-time state of the vehicle; Based on the reference engine output power at each moment within the preset time period, determining the reference engine speed at each moment; According to the reference engine speed at each moment and the Grey Wolf Optimization Algorithm, determining the target engine output power at each moment; Regarding the target engine output power at each moment as the optimal control action, and controlling the vehicle to drive within the preset time period based on the optimal control action.
2. The vehicle multi-power coordinated control method according to claim 1, wherein Before defining the engine output power as an action variable, it further includes: According to the engine output speed and the second-order inertia model at each moment, determining the engine output power at each moment; The expression of the second-order inertia model is: J e ω″(t) + Bω′(t) = T input (t) - T load (t); Among them, J e is the engine moment of inertia, B is the damping coefficient, ω′(t) is the first derivative of the engine output speed at the current moment; ω″(t) is the second derivative of the engine output speed at the current moment; T input (t) is the engine input torque; T load (t) is the engine load torque.
3. The vehicle multi-power coordinated control method according to claim 1, characterized in that Before inputting the real-time state of the vehicle within the preset time period into a trained offline energy management controller to determine the reference engine output power at each moment within the preset time period during the driving process of the vehicle in a preset time period, it further includes: Constructing an offline energy management controller based on the Deep Deterministic Policy Gradient Algorithm; the offline energy management controller includes a policy network, a value network, a target policy network, and a target value network; Based on the historical real-time state of the vehicle and the offline energy management controller, determining multiple empirical sample data and storing them in an experience pool; each empirical sample data includes the historical state variables at the current moment, the historical action variables at the current moment, the historical reward at the current moment, and the historical state variables at the next moment; the historical state variables include the historical real-time vehicle speed and the historical remaining battery power; the historical action variables include the historical engine output power; Training the offline energy management controller according to multiple empirical sample data to determine the trained offline energy management controller.
4. The vehicle multi-power coordinated control method according to claim 3, wherein Based on the historical real-time state of the vehicle and the offline energy management controller, determining multiple empirical sample data and storing them in an experience pool, specifically including: Inputting the historical state variables at the current moment into the policy network to obtain the historical action variables at the current moment; Regarding the historical state variables at the current moment, the historical action variables at the current moment, the historical state variables at the next moment, and the historical reward at the current moment as a quadruple at the current moment; the quadruple is the empirical sample data; Inputting the historical state variables at the next moment into the target policy network to obtain the historical action variables at the next moment; Take the historical state quantity at the next moment as the historical state quantity at the current moment, return to the step "Take the historical state quantity at the current moment, the historical action variable at the current moment, the historical state quantity at the next moment, and the reward at the current moment as a quadruple", and determine the quadruple at the next moment; Take the quadruples at each moment as multiple experience sample data and store them in the experience pool.
5. The vehicle multi-power coordinated control method according to claim 3, wherein Train the offline energy management controller according to the multiple experience sample data, and determine the trained offline energy management controller, specifically including: Based on the value network, evaluate the historical state quantity at the current moment and the historical action variable at the current moment, and determine the action state estimation value at the current moment; Based on the target value network, evaluate the historical state quantity at the next moment and the action variable at the next moment, and determine the target valuation of the action state at the next moment; Based on the historical reward at the current moment and the target valuation of the action state at the next moment, determine the updated target valuation of the action state at the next moment, and use the updated target valuation of the action state at the next moment as the target Q value; Determine the difference between the target Q value and the action state estimation value at the current moment, and minimize the difference to train the parameters of the value network, and determine the trained value network; Based on the gradient of the historical action variable at the current moment of the value network and the parameters of the policy network at the current moment, determine the gradient of the policy network at the current moment, and update the parameters of the policy network with the goal of minimizing the gradient of the policy network at the current moment, and determine the trained policy network; According to the parameters of the value network, the parameters of the policy network, and the soft update mechanism, update the parameters of the target value network and the parameters of the target policy network, and determine the trained target value network and target policy network; Take the trained policy network, the trained value network, the trained target policy network, and the trained target value network as the trained offline energy management controller.
6. The vehicle multi-power coordinated control method according to claim 5, characterized in that The calculation formula for the gradient of the policy network at the current moment is: Among them, is the gradient of the policy network at the current moment; θ π (t) are the parameters of the policy network; a′(t) and a are both historical action variables at the current moment; s′(t) and s are both historical state variables at the current moment; is the gradient of the value network for a′(t) under the policy network π[s, θ π (t)] at the current moment; θ Q (t) are the parameters of the value network at the current moment; Q is the value network; π is the policy network; is the gradient of the policy network π[s, θ π (t)] with respect to θ π (t) at the current moment; N is the number of empirical sample data, and t is the current moment.
7. The vehicle multi-power coordinated control method according to claim 1, wherein According to the engine reference speed at each moment and the grey wolf optimization algorithm, determine the engine target output power at each moment, specifically including: Set the constraint range of the engine reference speed at each moment; According to the constraint range and the cost function of the grey wolf optimization algorithm, use the grey wolf optimization algorithm to process the engine reference speed at each moment, and determine the engine reference output speed at each moment; According to the engine reference output speed at each moment and the power calculation formula, determine the engine target output power at each moment; The cost function J(t) of the grey wolf optimization algorithm is: J(t) = min{ω1|SOC(t) - SOC ref | + ω2|P eng,ref (t) - P eng (t)|}; Among them, both ω2 and ω1 are cost coefficients; SOC(t) is the real-time remaining battery power at the current moment; SOC ref is the preset reference remaining power reference output power; P eng,ref (t) is the reference output power of the engine at the current moment; P eng (t) is the actual output power of the engine at the current moment; || is the absolute value.
8. A computer device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the vehicle multi-power coordinated control method according to any one of claims 1-7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the vehicle multi-power coordinated control method according to any one of claims 1-7.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the vehicle multi-power coordinated control method described in any one of claims 1-7.
Citation Information
Cited By
Vehicle energy management method, device, equipment, medium and product
CN121268807A