Energy management strategy making method of hybrid power integrated vehicle and related device
The enhanced soft actor-critic algorithm with online hierarchical clustering and multi-dimensional priority metrics addresses the stability and efficiency challenges in deep reinforcement learning for hybrid electric vehicles, enhancing fuel economy by ensuring diverse and relevant training data.
Patent Information
- Application Number
- CN202510431486.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-15
AI Technical Summary
The existing hybrid vehicle energy management strategies are difficult to adapt to in complex and changeable driving environments. Deep reinforcement learning lacks sample diversity and learning efficiency during training, resulting in poor fuel economy.
The initial experience pool is clustered online by a comprehensive hierarchical clustering algorithm, and the sample importance and diversity are evaluated through multi-dimensional priority indicators, and combined with an enhanced soft behavior-evaluation algorithm to improve the learning efficiency and convergence speed of deep reinforcement learning models.
It improves the generalization ability of hybrid vehicles in dynamic environments, significantly improving the fuel economy and the robustness of learning models.
Smart Images

Figure CN120308084A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of energy management, and particularly to a method for formulating an energy management strategy for a hybrid integrated vehicle and related devices. Background Art
[0002] Hybrid vehicles have become an important development direction in the automotive industry because they can combine the advantages of internal combustion engines and electric motors. One of the core challenges of hybrid vehicles is how to reasonably allocate energy between different power sources to achieve the best fuel economy and emission performance. The energy management strategy (EMS) plays a crucial role in this process and directly affects the overall performance and user experience of the vehicle.
[0003] Currently, the energy management strategies of hybrid vehicles are mainly divided into three categories: rule-based, optimization-based, and reinforcement learning-based strategies. Each strategy has its unique advantages and limitations, as follows:
[0004] Rule-based energy management strategies: These strategies rely on predefined rules and expert knowledge, and have good real-time performance and operability. However, their limitation is that it is difficult to cope with complex and changeable driving environments, and they rely heavily on manually designed rules, lacking adaptability.
[0005] Optimization-based energy management strategies: These strategies formulate energy allocation schemes through mathematical modeling and optimization algorithms, and can theoretically provide more precise energy management. However, optimization methods usually require a large amount of computing resources and have high requirements for the vehicle's state space and constraint conditions, resulting in limitations in their real-time performance and operability in practical applications.
[0006] Energy management strategies based on deep reinforcement learning (DRL): In recent years, methods based on deep reinforcement learning have gradually become a research hotspot. Deep reinforcement learning learns the optimal energy allocation strategy through interaction with the environment and can achieve automatic adaptation under dynamic working conditions. Compared with traditional optimization methods, deep reinforcement learning can optimize the strategy through continuous exploration and feedback, and has better flexibility and adaptability. However, deep reinforcement learning usually requires a large amount of sampling and computing resources during the training process, and how to ensure the stability and efficiency of learning is still a challenge when facing complex multi-source power systems. Summary of the Invention
[0007] The purpose of the present application is to provide a method and device for formulating an energy management strategy for a hybrid integrated vehicle, which can effectively improve the learning efficiency and convergence speed of the deep reinforcement learning model, and significantly improve the fuel economy of hybrid vehicles.
[0008] To achieve the above object, the present application provides the following solutions:
[0009] In a first aspect, the present application provides a method for formulating an energy management strategy for a hybrid integrated vehicle, including:
[0010] Obtain the state data of the hybrid vehicle at each moment; the state data includes vehicle speed, acceleration, remaining battery power, and vehicle demand power;
[0011] Input the state data of the hybrid vehicle at each moment into a deep reinforcement learning model to obtain an energy management strategy for the hybrid vehicle; the energy management strategy includes the actions corresponding to the state data at each moment; the action is the throttle opening of the engine in the hybrid vehicle; the deep reinforcement learning model is determined based on an enhanced soft behavior-evaluation algorithm;
[0012] The determination process of the deep reinforcement learning model is as follows:
[0013] Input the state data of the hybrid vehicle at the current moment into the policy network to obtain the action at the current moment, and let the hybrid vehicle execute the action at the current moment to obtain the state data at the next moment;
[0014] Calculate the reward signal at the current moment according to the state data at the current moment, the state data at the next moment, and the reward function;
[0015] Construct an initial experience pool; the initial experience pool includes multiple initial experiences, and the initial experience includes the state data at the current moment, the state data at the next moment, the action at the current moment, and the reward signal at the current moment;
[0016] Use the comprehensive hierarchical clustering algorithm to perform online clustering on the initial experiences in the initial experience pool to obtain a first experience pool;
[0017] Perform sampling processing on the initial experience pool and the first experience pool to obtain a second experience pool; the second experience pool includes multiple experiences;
[0018] Calculate the multi-dimensional priority index values corresponding to each experience in the second experience pool;
[0019] Based on the second experience pool and the multi-dimensional priority index values, update the weights of the policy network and the weights of the evaluation network to obtain a trained policy network and a trained evaluation network, and use the trained policy network and the trained evaluation network as the deep reinforcement learning model.
[0020] In a second aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the method for formulating an energy management strategy for a hybrid integrated vehicle described above.
[0021] In a third aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the method for formulating an energy management strategy for a hybrid integrated vehicle described above.
[0022] In a fourth aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the method for formulating an energy management strategy for a hybrid integrated vehicle described above.
[0023] According to the specific embodiments provided by the present application, the following technical effects are disclosed:
[0024] The present application provides a method for formulating an energy management strategy for a hybrid integrated vehicle and related devices. By introducing a comprehensive hierarchical clustering algorithm to perform online clustering on the experience data in the initial experience pool, dynamically merging or splitting newly sampled data into existing clusters, and extracting samples from different clusters, this method effectively improves the sample diversity and avoids the problem of over-reliance on certain high-frequency samples in traditional experience replay methods. Moreover, the low time complexity of the comprehensive hierarchical clustering algorithm ensures the efficiency of online operations, enabling experience management to adapt to the large-scale data requirements in complex environments, thereby enhancing the generalization ability of the reinforcement learning model under variable working conditions. In addition, the present application comprehensively evaluates the learning priorities of samples by calculating multi-dimensional priority indicators, which can ensure that both the importance and diversity of samples are considered during experience sampling, effectively improving the learning efficiency and convergence speed of the deep reinforcement learning model, and making the reinforcement learning model more robust in dynamic environments. Therefore, by inputting the state data of the hybrid vehicle at each moment into the deep reinforcement learning model, a more optimized energy management strategy can be obtained, thus significantly improving the fuel economy of the hybrid vehicle. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present application or in the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0026] Figure 1 It is an application environment diagram of a method for formulating an energy management strategy for a hybrid integrated vehicle in an embodiment of the present application;
[0027] Figure 2 A flowchart showing a method for formulating an energy management strategy for a hybrid integrated vehicle provided in an embodiment of the present application;
[0028] Figure 3 A schematic structural diagram of a deep reinforcement learning model provided in another embodiment of the present application;
[0029] Figure 4 A flowchart showing the training process of a deep reinforcement learning model provided in another embodiment of the present application;
[0030] Figure 5 A schematic structural diagram of a policy network and an evaluation network provided in another embodiment of the present application;
[0031] Figure 6 A schematic structural diagram of a computer device provided in an embodiment of the present application. Detailed implementation manners
[0032] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0033] When applying the deep reinforcement learning algorithm to formulate the energy management strategy, the deep reinforcement learning algorithm still faces two major challenges in terms of experience sampling. First, traditional sampling methods often ignore the importance of sample diversity, which limits the exploration of the agent among different states, actions, and rewards, resulting in its over-reliance on repeated experiences; this lack of diversity not only hinders the generalization ability of the agent in diverse scenarios, especially in complex real-world environments, but is also crucial for expanding the learning horizon of the agent, improving its adaptability, and enhancing the robustness of the strategy. Second, sampling methods that rely on a single evaluation metric (such as the Temporal Difference (TD) error), such as Prioritized Experience Replay (PER), have limitations in determining the importance of samples. Although the TD error is crucial for evaluating the direct learning value of experiences, it does not comprehensively reflect the contribution of each experience to the entire learning process. Therefore, experiences with low TD error but high potential value in specific transitions may be ignored due to the lack of "salience", which limits the efficiency of sampling.
[0034] Therefore, to address these challenges, the present application proposes an Enhanced Soft Actor-Critic (ESAC) algorithm. By improving the experience sampling strategy, this algorithm aims to overcome the limitations of traditional methods and enhance the learning effect of agents in complex environments.
[0035] To make the above objects, features, and advantages of the present application more apparent and understandable, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0036] The method for formulating an energy management strategy for a hybrid integrated vehicle provided by an embodiment of the present application can be applied to an application environment as Figure 1 shown. Among them, the terminal 102 communicates with the server 104 through a network. The data storage system can store the data that the server 104 needs to process. The data storage system can be set up separately, integrated on the server 104, placed on the cloud, or on other servers. The terminal 102 can send the state data of the hybrid vehicle at each moment to the server 104. After receiving the state data, the server 104 inputs the state data of the hybrid vehicle at each moment into the deep reinforcement learning model to obtain the energy management strategy of the hybrid vehicle. The server 104 can feedback the obtained energy management strategy to the terminal 102. In addition, in some embodiments, the method for formulating an energy management strategy can also be implemented separately by the server 104 or the terminal 102. For example, the terminal 102 can directly process the state data at each moment, or the server 104 can obtain the state data at each moment from the data storage system and process it.
[0037] Among them, the terminal 102 can be, but is not limited to, various desktop computers, laptop computers, and Internet of Things devices. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers, and can also be a cloud server.
[0038] In an exemplary embodiment, as Figure 2 shown, a method for formulating an energy management strategy for a hybrid integrated vehicle is provided. This method is executed by a computer device, and can be specifically executed alone by a computer device such as a terminal or a server, or jointly executed by a terminal and a server. In the embodiment of the present application, taking this method applied to Figure 1 the server 104 in it as an example for illustration, it includes the following steps 201 to step 202. Among them:
[0039] Step 201, obtaining the state data of the hybrid vehicle at each moment; the state data includes vehicle speed, acceleration, remaining battery power, and vehicle demand power.
[0040] Step 202: Input the state data of the hybrid vehicle at each moment into the deep reinforcement learning model to obtain the energy management strategy of the hybrid vehicle; the energy management strategy includes the actions corresponding to the state data at each moment; the action is the throttle opening of the engine in the hybrid vehicle; the deep reinforcement learning model is determined based on the enhanced soft behavior-evaluation algorithm.
[0041] Among them, the determination process of the deep reinforcement learning model is as follows:
[0042] Step 301: Input the state data of the hybrid vehicle at the current moment into the policy network to obtain the action at the current moment, and let the hybrid vehicle execute the action at the current moment to obtain the state data at the next moment.
[0043] Step 302: Calculate the reward signal at the current moment according to the state data at the current moment, the state data at the next moment, and the reward function.
[0044] Step 303: Construct an initial experience pool; the initial experience pool includes multiple initial experiences, and the initial experience includes the state data at the current moment, the state data at the next moment, the action at the current moment, and the reward signal at the current moment.
[0045] Step 304: Use the comprehensive hierarchical clustering algorithm to perform online clustering on the initial experiences in the initial experience pool to obtain the first experience pool.
[0046] Step 305: Perform sampling processing on the initial experience pool and the first experience pool to obtain the second experience pool; the second experience pool includes multiple experiences.
[0047] Step 306: Calculate the multi-dimensional priority index values corresponding to each experience in the second experience pool.
[0048] Step 307: Based on the second experience pool and the multi-dimensional priority index values, update the weights of the policy network and the weights of the evaluation network to obtain the trained policy network and the trained evaluation network, and use the trained policy network and the trained evaluation network as the deep reinforcement learning model.
[0049] Furthermore, the calculation formula of the reward function is:
[0050] r = α eng ·m f +β bat ·(SOC - SOC ref ) 2 ;
[0051] Among them, r is the reward signal, m f is the fuel consumption, SOC ref is the reference value of the remaining battery charge, SOC is the remaining battery charge, αeng is the first pre - coefficient, β bat is the second pre - coefficient.
[0052] Furthermore, the network structure of the policy network includes a first fully - connected layer, a first activation function, a branch layer, and an output layer connected in sequence; the branch layer includes a mean branch and a standard - deviation branch connected in parallel; both the mean branch and the standard - deviation branch include a second fully - connected layer, a second activation function, and a third fully - connected layer connected in sequence. The network structure of the evaluation network includes two independent Q - networks; the Q - network includes a fourth fully - connected layer, a third activation function, a fifth fully - connected layer, a fourth activation function, and a sixth fully - connected layer connected in sequence. Among them, the first activation function, the second activation function, the third activation function, and the fourth activation function are all ReLU activation functions.
[0053] Furthermore, in step 305, sampling is performed on the initial experience pool and the first experience pool to obtain a second experience pool, which specifically includes:
[0054] Randomly sample 10% of the data from the initial experience pool and randomly sample 90% of the data from the first experience pool, and fuse them to obtain the second experience pool.
[0055] Furthermore, the calculation formula for the multi - dimensional priority index is:
[0056]
[0057] Among them, TD srror is the time - difference error of the current experience, and each experience in the second experience pool is taken as the current experience in turn; Educlidean distance is the average Euclidean distance between the current experience and other experiences; Time correlation is the time correlation, which measures the time difference between the current experience and the latest experience; α TD 、β Edu and γ T are the sensitivities of the first index, the second index, and the third index respectively. The first index is the experience importance, the second index is the sample diversity, and the third index is the time correlation; ω1, ω2, and ω3 are the pre - coefficients of the first index, the second index, and the third index respectively.
[0058] Furthermore, in step 307, based on the second experience pool and the multi - dimensional priority index values, the weights of the policy network and the weights of the evaluation network are updated to obtain the trained policy network and the trained evaluation network, which specifically includes:
[0059] Based on the second experience pool and the multi-dimensional priority metric values, update the weights of the policy network according to the loss function of the policy network, update the weights of the evaluation network according to the loss function of the evaluation network, and determine whether the current training round has reached the preset number of training rounds.
[0060] If the current training round reaches the preset number of training rounds, obtain the trained policy network and the trained evaluation network; if the current training round does not reach the preset number of training rounds, return to the step "Input the state data of the hybrid vehicle at the current moment into the policy network to obtain the action at the current moment."
[0061] Furthermore, the loss function of the policy network is specifically:
[0062]
[0063] where Lactor is the policy network loss value; Es ~D ,a ~ π is the policy network expectation, s is the state data at the current moment, a is the action at the current moment, D is the second experience pool, π is the policy; αentropylogπ(a|s) is the first entropy term, which increases the exploration of the policy, αentropy is the weight for controlling the entropy term, and logπ(a|s) is the probability of executing action a at state s according to policy π; is the smaller Q value among the two Q networks in the evaluation network;
[0064] Furthermore, the loss function of the evaluation network is specifically:
[0065] Lcritic = E(s,a,r,s ′ ) ~D [(Q(s,a) - y) 2 ;
[0066] where Lcritic is the evaluation network loss value; E( s ,a,r,s′) ~D is the evaluation network expectation, s is the state data at the current moment, a is the action at the current moment, r is the reward signal at the current moment, s' is the state data at the next moment, D is the second experience pool; Q(s,a) is the Q value calculated by the evaluation network; y is the target Q value.
[0067] With the above deep reinforcement learning model, the present application performs online clustering on the experience data in the initial experience pool by introducing a comprehensive hierarchical clustering algorithm, dynamically merges or splits newly sampled data into existing clusters, and samples according to the empirical distribution ratio of each cluster. This method effectively improves sample diversity and avoids the problem of over-reliance on certain high-frequency samples in traditional experience replay methods. Moreover, the low time complexity of the comprehensive hierarchical clustering algorithm ensures the efficiency of online operations, enabling experience management to adapt to the large-scale data requirements in complex environments, thereby enhancing the generalization ability of the reinforcement learning model under variable working conditions. In addition, the present application comprehensively evaluates the learning priority of samples by integrating TD error, sample distance, and time correlation, which can ensure that both the importance and diversity of samples are considered during experience sampling, effectively improving the learning efficiency and convergence speed of the deep reinforcement learning model, and making the reinforcement learning model more robust in dynamic environments.
[0068] In another exemplary embodiment of the present application, for a hybrid tracked vehicle combining an engine-generator set with a single battery pack, the determination process of the deep reinforcement learning model in the method for formulating an energy management strategy is specifically and detailedly described. The structure of its deep reinforcement learning model is as Figure 3 shown, and the training process of the deep reinforcement learning model is as Figure 4 shown. In the energy management strategy based on deep reinforcement learning, the hybrid tracked vehicle model itself constitutes the environment.
[0069] Step 1: Set the state space, action space, and reward function.
[0070] The state space refers to the vehicle speed v, acceleration a acc , battery SOC, vehicle demand power P dem , etc., which are quantities used to describe the current environment, and are expressed as:
[0071] S = [v, a acc , SOC, P dem (1).
[0072] The action space refers to the decisions made. The throttle openings of the two engines are used as actions, and are expressed as:
[0073] A = [Thr1, Thr2](2).
[0074] Among them, Thr1 and Thr2 are the throttle openings of engine 1 and engine 2 respectively.
[0075] The reward signal refers to the feedback value given by the environment to the agent after the agent executes an action and acts on the environment. In this embodiment, it is considered that while the vehicle achieves fuel economy, the battery SOC is maintained in the efficient range, and is expressed as:
[0076] r = α eng ·m f + β bat ·(SOC - SOC ref ) 2 (3);
[0077] Wherein, r is the reward signal; m f is the fuel consumption; SOC ref is the reference value of the remaining power, which is set to 0.7 in this embodiment; SOC is the remaining power; α eng is the first pre - coefficient; β bat is the second pre - coefficient. Step two: Initialize the parameters of the policy network (Actor network) and the evaluation network (Critic network).
[0078] The network structure of the SAC algorithm includes an Actor network and a Critic network, and its structure is as Figure 5 shown. Among them, the Actor network optimizes the action policy according to the value evaluation provided by the Critic network, while the Critic network calculates the state - action value through the actions generated by the Actor network. The two cooperate with each other to jointly achieve the optimization of the policy and the improvement of performance.
[0079] The Actor network is used for policy learning and generates actions based on state information (i.e., state data). Its structure receives state information through the input layer, passes through a first fully - connected layer with 256 neurons and uses the first activation function (i.e., ReLU activation function), and then is divided into a mean (Mean) branch and a standard deviation (Std) branch. Each branch contains two fully - connected layers and an activation function (i.e., the second fully - connected layer, the second activation function, and the third fully - connected layer), which respectively output the mean and standard deviation of the Gaussian distribution, and finally sample the actions through the Gaussian distribution to output the actions.
[0080] The Critic network is used to evaluate the value of the current policy and contains two independent Q - networks. By introducing a dual - network structure, the problem of over - estimation of Q - values is alleviated. Each Q - network takes state information as input, passes through three fully - connected layers with 256 neurons (i.e., the fourth fully - connected layer, the fifth fully - connected layer, and the sixth fully - connected layer), and uses the ReLU activation function after the fourth fully - connected layer and the fifth fully - connected layer, and finally outputs the corresponding Q - value. It should be noted that each Critic network also has a corresponding target Critic network, which is used as a reference value to guide the learning of the main Critic network, aiming to reduce the fluctuation between the target value and the predicted value and improve the stability of training.
[0081] Since the Actor network and the Critic network contain a large number of parameters that need to be optimized through training iterations, it is necessary to initialize them reasonably first, laying a foundation for the subsequent update of network parameters and the training process.
[0082] Step 3: Select an action according to the environmental state.
[0083] The agent inputs the state data s at the current moment into the Actor network. The state data is a specific instance in the state space, representing the specific state of the environment observed by the agent at a certain moment. The Actor network will generate the mean and standard deviation parameters of the action, sample through the Gaussian distribution, and after being processed by a series of fully connected layers and non-linear activation functions, output a continuous action a suitable for the current environment.
[0084] Step 4: The environment generates a new state and a reward signal
[0085] The environment receives the action a executed by the agent, dynamically updates according to the environment and generates new state data (i.e., the state data at the next moment) s', and at the same time calculates and outputs the reward signal r through the reward function. The reward signal r reflects the contribution of the current action a to the agent's achievement of the goal, and is used as feedback information to guide policy optimization, helping the agent learn better action selection.
[0086] Step 5: Construct an initial experience pool.
[0087] The state data s at the current moment, the state data s' at the next moment, the action a, and the reward signal r at the current moment are jointly formed into a data pair (s, s', a, r) and stored in the experience pool. The experience pool is used to store the historical data in the process of the agent's interaction with the environment. Subsequently, by randomly sampling the samples in the experience pool, the time correlation between the data is broken, the stability of training and the data utilization efficiency are improved, and the parameters of the Actor and Critic networks are further optimized.
[0088] Step 6: Online cluster the initial experiences in the initial experience pool using the BIRCH algorithm.
[0089] For each newly generated experience, in this embodiment, the Balanced Iterative Reducing and Clustering using Hierarchies (BIRCH) algorithm is first used to perform online clustering on the data in the experience pool based on the central features of the clusters and the distances between clusters, and dynamically merge new data into existing clusters or split existing clusters when necessary. The BIRCH algorithm is specifically designed to solve the clustering problem of large-scale data sets. Traditional clustering algorithms, such as K-Means and hierarchical clustering, often encounter problems related to computational complexity and memory consumption when dealing with large-scale data sets. The BIRCH algorithm aims to efficiently process large data sets by gradually constructing a hierarchical clustering structure, achieving a balance between clustering accuracy and computational efficiency. After online clustering, experiences with similar states, actions, and rewards are grouped into one category, and the first experience pool is formed.
[0090] When the experience pool is full, the BIRCH algorithm only retains the data within a fixed time window, and the outdated data will be logically deleted. Therefore, subsequent sampling will not include this data, ensuring the real-time relevance of the data set. The reason for using the logical method is that the BIRCH algorithm can complete the merging and splitting of data with a relatively low time complexity (O(N)) when dealing with large-scale data. However, the process of deleting outdated data requires frequent reconstruction of the BIRCH tree, which will significantly increase the time overhead.
[0091] Step 7: Sample to form a new experience pool.
[0092] Sample from each cluster according to the proportion of the data in each cluster to the total data volume to form a new experience pool (i.e., the second experience pool). The new experience pool contains N * batchSize data (batch_size refers to the number of data updated in batches when the neural network is updated, and N is a parameter specified by oneself. In this embodiment, it is set to 2). Among them, 90% of the data comes from the clusters in the first experience pool, and the remaining 10% of the data is randomly sampled from the initial experience pool, which ensures the diversity and randomness of the data in the new experience pool and prevents the model from converging to a local optimum.
[0093] Step 8: Calculate the multi-dimensional priorities for the experiences in the new experience pool.
[0094] Calculate a multi-dimensional priority index for each data in the new experience pool, and sample batchSize data from the new experience pool according to the multi-dimensional priority index. The multi-dimensional priority index is defined as follows:
[0095]
[0096] Among them, TD srroris the temporal difference error of the current experience, and each experience in the second experience pool is taken as the current experience in turn; Educlidean distance is the average Euclidean distance between the current experience and other experiences; Time correlation is the time correlation, which measures the time difference between the current experience and the latest experience; α TD 、β Edu and γ T are the sensitivities of the first index, the second index, and the third index respectively. The first index is the experience importance, the second index is the sample diversity, and the third index is the time correlation; ω1, ω2, and ω3 are the weighting coefficients of the first index, the second index, and the third index respectively.
[0097] This method fully considers the importance of experience, the diversity of samples, and the time correlation. At the same time, this method adopts the way of non-linear function and adjustable hyperparameters. Compared with the traditional multi-dimensional index linear combination method, it effectively improves the adaptability of the model to different types of experience, so that the reinforcement learning model performs more robustly and efficiently in complex environments.
[0098] Therefore, the sampling probability P of the i-th data point i is defined as follows:
[0099]
[0100] where N is the total number of experiences in the new experience pool; p i is the priority of the i-th data point; p k is the priority of any experience in the new experience pool.
[0101] Step Nine: Update the weights of the Actor network and the Critic network.
[0102] The agent samples batchSize experiences from the new experience pool according to the multi-dimensional priority to update the weights of the Actor network and the Critic network of the SAC algorithm.
[0103] The Critic network updates the network weights by minimizing the mean square error loss, and the target Q value is:
[0104]
[0105] where r is the reward signal; γ is the discount factor; Qtarget(s′,a′) is the Q value calculated by the target network; αentropy·logπ(a′|s′) is the second entropy term, which increases the exploration of the policy. αentropy is the weight for controlling the entropy term, and logπ(a′|s′) is the probability of executing action a′ at state s′ according to the policy π; d represents the termination signal; The smaller Q-value from the two target Critic networks.
[0106] The loss function of the Critic network is:
[0107] Lcritic = E(s,a,r,s ′ ) ~D [(Q(s,a) - y) 2 (7);
[0108] Where, Lcritic is the evaluation network loss value; E( s ,a,r,s′) ~D is the evaluation network expectation, s is the state data at the current moment, a is the action at the current moment, r is the reward signal at the current moment, s' is the state data at the next moment, D is the new experience pool (i.e., the second experience pool); Q(s,a) is the Q-value calculated by the evaluation network; y is the target Q-value.
[0109] By minimizing this loss function, the weights of the Critic network are updated. The target Critic network updates the network parameters in a soft update manner every few rounds.
[0110]
[0111] Where, θ1 and θ2 are the parameters of the two Critic networks respectively; and are the parameters of the two target Critic networks respectively; τ is the soft update factor, which is set to 0.05 in this embodiment.
[0112] The goal of the Actor network is to maximize the Q-value and the entropy of the policy, so as to balance exploration and policy optimization. The specific update is based on the following loss function of the Actor network:
[0113]
[0114] Where, L actor is the policy network loss value; E s~D,a~π is the policy network expectation, s is the state data at the current moment, a is the action at the current moment, D is the new experience pool (i.e., the second experience pool), π is the policy; α entropy logπ(a|s) is the first entropy term, which increases the exploration of the policy, and logπ(a|s) is the probability of performing action a at state s according to policy π; is the smaller Q-value of the two Q networks in the evaluation network.
[0115] By minimizing the above loss function, the Actor network learns a policy that balances high Q-values and high entropy. Finally, the parameters of the Actor network are updated through gradient descent.
[0116] Step Ten: Determine whether the end condition is reached.
[0117] During the training process of the Actor network and the Critic network, it is necessary to determine whether the current training round has reached the preset maximum number of training rounds. If not, continue to execute the training loop from Step Three to Step Ten; if it reaches the preset maximum number of training rounds, end the training and save the model parameters.
[0118] In an exemplary embodiment, a computer device is provided. The computer device can be a server or a terminal, and its internal structure diagram can be as Figure 6 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store processed data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a method for formulating an energy management strategy for a hybrid integrated vehicle.
[0119] Those skilled in the art can understand that Figure 6 the structure shown in
[0120] merely represents a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0121] In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, it implements the steps in the above method embodiments.
[0122] In an exemplary embodiment, a computer program product is provided, including a computer program which, when executed by a processor, implements the steps in the above method embodiments.
[0123] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include Read-Only Memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0124] The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0125] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.
[0126] In this article, specific examples are used to illustrate the principles and implementation manners of the present application. The descriptions of the above embodiments are only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A method for formulating an energy management strategy for a hybrid integrated vehicle, characterized in that, The method for formulating the energy management strategy of the hybrid integrated vehicle includes: Obtaining the state data of the hybrid vehicle at each moment; the state data includes vehicle speed, acceleration, remaining battery power, and vehicle demand power; Inputting the state data of the hybrid vehicle at each moment into the deep reinforcement learning model to obtain the energy management strategy of the hybrid vehicle; the energy management strategy includes the actions corresponding to the state data at each moment; the action is the throttle opening of the engine in the hybrid vehicle; the deep reinforcement learning model is determined based on the enhanced soft behavior-evaluation algorithm; The determination process of the deep reinforcement learning model is: Inputting the state data of the hybrid vehicle at the current moment into the policy network to obtain the action at the current moment, and having the hybrid vehicle execute the action at the current moment to obtain the state data at the next moment; Calculating the reward signal at the current moment according to the state data at the current moment, the state data at the next moment, and the reward function; Constructing an initial experience pool; the initial experience pool includes multiple initial experiences, and the initial experience includes the state data at the current moment, the state data at the next moment, the action at the current moment, and the reward signal at the current moment; Using the comprehensive hierarchical clustering algorithm to perform online clustering on the initial experiences in the initial experience pool to obtain the first experience pool; Sampling the initial experience pool and the first experience pool to obtain the second experience pool; the second experience pool includes multiple experiences; Calculating the multi-dimensional priority index values corresponding to each experience in the second experience pool; Based on the second experience pool and the multi-dimensional priority index values, updating the weights of the policy network and the weights of the evaluation network to obtain the trained policy network and the trained evaluation network, and using the trained policy network and the trained evaluation network as the deep reinforcement learning model.
2. The method for formulating an energy management strategy for a hybrid integrated vehicle according to claim 1, characterized in that The calculation formula of the reward function is: r = α eng ·m f + β bat ·(SOC - SOC ref ) 2 ; Among them, r is the reward signal, m f is the fuel consumption, SOC ref is the reference value of the remaining power, SOC is the remaining power, α eng is the first pre-factor, β bat is the second pre-factor.
3. The method for formulating the energy management strategy of the hybrid integrated vehicle according to claim 1, wherein The policy network includes a first fully connected layer, a first activation function, a branch layer, and an output layer connected in sequence; the branch layer includes a mean branch and a standard deviation branch connected in parallel; both the mean branch and the standard deviation branch include a second fully connected layer, a second activation function, and a third fully connected layer connected in sequence; The evaluation network includes two independent Q networks; the Q network includes a fourth fully connected layer, a third activation function, a fifth fully connected layer, a fourth activation function, and a sixth fully connected layer connected in sequence.
4. The method for formulating an energy management strategy of a hybrid integrated vehicle according to claim 1, wherein Sampling the initial experience pool and the first experience pool to obtain the second experience pool, specifically including: Randomly sampling 10% of the data from the initial experience pool and randomly sampling 90% of the data from the first experience pool, and fusing them to obtain the second experience pool.
5. The method for formulating the energy management strategy of the hybrid integrated vehicle according to claim 1, characterized in that, The calculation formula of the multi-dimensional priority index is: Among them, TD srror is the time difference error of the current experience, and each experience in the second experience pool is used as the current experience in turn; Educlidean distance is the average Euclidean distance between the current experience and other experiences; Time correlation is the time correlation, which measures the time difference between the current experience and the latest experience; α TD , β Edu and γ T are the sensitivities of the first index, the second index, and the third index respectively. The first index is the experience importance, the second index is the sample diversity, and the third index is the time correlation; ω1, ω2, and ω3 are the weighting coefficients of the first index, the second index, and the third index respectively.
6. The method for formulating an energy management strategy for a hybrid integrated vehicle according to claim 1, wherein Based on the second experience pool and the multi-dimensional priority index values, updating the weights of the policy network and the weights of the evaluation network to obtain the trained policy network and the trained evaluation network, specifically including: Based on the second experience pool and the multi-dimensional priority metric value, update the weights of the policy network according to the loss function of the policy network, update the weights of the evaluation network according to the loss function of the evaluation network, and determine whether the current training round has reached the preset number of training rounds; If the current training round has reached the preset number of training rounds, obtain the trained policy network and the trained evaluation network; if the current training round has not reached the preset number of training rounds, return to the step "input the state data of the hybrid vehicle at the current moment into the policy network to obtain the action at the current moment".
7. The method for formulating an energy management strategy for a hybrid integrated vehicle according to claim 1, characterized in that, The loss function of the policy network is specifically: Among them, Lactor is the loss value of the policy network; Es ~D ,a ~ π is the policy network expectation, s is the state data at the current moment, a is the action at the current moment, D is the second experience pool, and π is the policy; α entropy logπ(a|s) is the first entropy term, which increases the exploration of the policy, α entropy is the weight for controlling the entropy term, and logπ(a|s) is the probability of executing action a at state s according to policy π; is the smaller Q value among the two Q networks in the evaluation network; The loss function of the evaluation network is specifically: Lcritic = E(s,a,r,s ′ ) ~D [(Q(s,a)-y) 2 ; Among them, Lcritic is the evaluation network loss value; E( s , a, r, s′) ~D is the evaluation network expectation, s is the state data at the current moment, a is the action at the current moment, r is the reward signal at the current moment, s' is the state data at the next moment, D is the second experience pool; Q(s, a) is the Q value calculated by the evaluation network; y is the target Q value.
8. A computer device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the energy management strategy formulation method for the hybrid integrated vehicle according to any one of claims 1-7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the energy management strategy formulation method for the hybrid integrated vehicle according to any one of claims 1-7.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the energy management strategy formulation method for the hybrid integrated vehicle according to any one of claims 1-7.