A battery fast charging control method based on PPO algorithm and considering charging electricity cost

Through multi-objective optimization based on PPO algorithm, combined with the electric and thermal coupling model of lithium-ion batteries and the charging and electricity bill optimization model, the conflict between fast charging and battery health and cost is solved, and a safe, fast and economical charging strategy is achieved.

CN115447431BActive Publication Date: 2025-08-26NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202211100593.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-09
Publication Date
2025-08-26
Estimated Expiration
2042-09-09

AI Technical Summary

Technical Problem

The existing charging strategies cannot effectively take into account the fast charging needs of lithium batteries, the health of the battery and the charging cost, and the calculation complexity is high, which poses safety risks.

Method used

Using a multi-objective optimization method based on PPO algorithm, a lithium-ion battery electric and thermal coupling model and charging electricity bill optimization model are constructed, and a charging policy network and policy evaluation network are established through offline training to achieve rapid charging while optimizing electricity bills and safety constraints.

Benefits of technology

Fast charging is achieved while meeting battery safety and health constraints, reducing the computational complexity of online charging decisions and reducing charging costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115447431B_ABST
    Figure CN115447431B_ABST
Patent Text Reader

Abstract

The present invention discloses a battery fast charging control method based on the PPO algorithm and taking charging electricity costs into consideration. The method constructs a lithium-ion battery electrothermal coupling model and a charging electricity cost optimization model, determines their key state variables, normalizes them and places them into the reinforcement learning state space, and defines the action space and reward function. The constructed charging strategy network and strategy evaluation network are trained based on the proximal strategy optimization algorithm until the charging strategy network and strategy evaluation network converge, and the charging strategy network is derived as the battery fast charging strategy. Real-time data is collected and input into the trained charging strategy network to determine the optimal charging action at the current moment. After each charging cycle, the state variables are recollected and the charging current is determined until charging is completed. The present invention can achieve fast charging with active awareness of safety and health and low charging cost. The complex calculations caused by multi-constraint and multi-objective optimization solutions are transferred to the offline training link, significantly reducing the computational complexity of online charging decision-making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of battery fast charging, and in particular relates to a battery fast charging control method based on a PPO algorithm and taking charging electricity charges into consideration. Background Art

[0002] To alleviate the increasingly severe energy shortages and environmental pollution, traditional fuel-powered vehicles are gradually being phased out. However, electric vehicles, with their zero emissions and high energy efficiency, are an ideal alternative and a hot topic of research. Lithium batteries, the power source of electric vehicles, suffer from long charging times due to their inherent electrochemical limitations, which can easily lead to range anxiety among users and severely restrict the development of the electric vehicle industry. Furthermore, charging with excessive current poses safety risks, potentially causing fires, explosions, and other accidents. Therefore, research on safe and efficient charging control strategies is urgently needed.

[0003] Among traditional charging strategies, the most widely used are constant current-constant voltage (CCCV) and multi-stage constant current methods. The selection of charging parameters for these methods depends on the designer's experience and knowledge, resulting in a mismatch between the charging mode and the battery. Furthermore, with increasing user demands, the battery charging process requires an increasing number of optimization objectives, and traditional strategies are unable to meet these growing needs. Prior art methods utilize electrothermal and aging models, along with multi-objective optimization algorithms based on biogeography, to optimize the charging process from three perspectives: battery health, charging time, and energy conversion efficiency. This achieves a reasonable trade-off between these objectives. However, these methods require solving multi-constraint, multi-objective optimization problems involving high dimensions, strong coupling, and nonlinearities, resulting in high computational complexity and challenging online applications. Chinese patent publication number CN112018465A utilizes the DDPG algorithm to train a neural network for battery charging current decisions, ensuring that the battery meets charging physical constraints while completing the charging task. This method reduces computational complexity in online applications by pre-training the neural network. However, considering only the battery itself, it fails to meet user demands for lower charging costs. Summary of the Invention

[0004] Purpose of the invention: The purpose of the present invention is to overcome the shortcomings of the existing technology and propose a battery fast charging control method based on the PPO algorithm and taking into account the charging electricity cost. By establishing a multi-objective optimization problem and using the PPO algorithm to solve it, fast charging that complies with the battery voltage and temperature constraints is achieved; this method migrates the complex calculations caused by multi-constraint and multi-objective optimization solutions to the offline training link, ensuring the real-time performance of the algorithm.

[0005] Technical solution: The present invention provides a battery fast charging control method based on the PPO algorithm and taking charging electricity costs into consideration, which specifically includes the following steps:

[0006] (1) Construct a lithium-ion battery electrothermal coupling model and a charging electricity cost optimization model, and establish an offline training scenario based on the two constructed models to determine their key state variables;

[0007] (2) Normalize the key state variables determined in step (1) and place them in the reinforcement learning state space, defining the action space and reward function;

[0008] (3) Training the constructed charging strategy network and strategy evaluation network based on the proximal strategy optimization algorithm; the charging strategy network generates a charging action according to the acquired state variables, updates the battery state according to the lithium-ion battery electrothermal coupling model in step (1), and records the charging action, battery state, and reward value in the experience pool, and synchronously updates the charging strategy network and strategy evaluation network through the experience pool information;

[0009] (4) Execute step (3) repeatedly until the charging strategy network and the strategy evaluation network converge, and derive the charging strategy network as the battery fast charging strategy;

[0010] (5) Real-time collection of the battery's current power, terminal voltage, ambient temperature, battery surface temperature, and current electricity price, and normalization of these data, which are then input into the charging strategy network trained in step (4) to determine the optimal charging action at the current moment;

[0011] (6) After each charging cycle, the state quantity is collected again and the charging current is determined until charging is completed.

[0012] Furthermore, the key state variables in step (1) include battery charge SOC, battery voltage V B , average battery temperature T a and electricity price p.

[0013] Furthermore, the process of constructing the lithium-ion battery electrothermal coupling model in step (1) is as follows:

[0014] Voltage source V OC The resistors R0 and R1 are used to simulate the battery's energy storage and charge / discharge energy loss, respectively. The RC networks (R1, C1) and (R2, C2) characterize the battery's short-term and long-term transient responses. According to Kirchhoff's current and voltage laws, the battery's dynamic characteristics are described as follows:

[0015]

[0016] Where, SOC(k), C n , I B (k), VB (k) represents the battery's SOC state, nominal capacity, charging current, and voltage; the battery's open circuit voltage V OC (k) is a nonlinear function of SOC(k): V OC (k) = g(SOC(k));

[0017] V1(k) and V2(k) are the voltages across capacitors C1 and C2, respectively; R0 is a constant resistance; T C 、T S represents the battery core temperature and surface temperature, which are calculated according to the principle of conservation of energy as:

[0018]

[0019]

[0020] Where, T amb is the ambient temperature of the battery; R C 、R u Represents thermal conduction resistance and convection resistance respectively; C C 、C S Represent the internal heat capacity and surface heat capacity of the battery respectively; the temperature of the battery is defined as T S and T C The average value of:

[0021]

[0022] Furthermore, the charging electricity fee optimization model construction process in step (1) is as follows:

[0023] Minimize the time it takes for the battery to charge from any initial SOC (0) to the desired SOC d The time spent and the objective function corresponding to the charging speed are:

[0024] min J1=NT (5)

[0025] Where T represents the sampling period, N is SOC(N)=SOC d The corresponding number of sampling steps;

[0026] The charging cost is affected by the battery's charging current and the current electricity price. The objective function for optimizing the charging cost is:

[0027]

[0028] Where p(k) is the time-of-use electricity price at charging sampling period k; J2 is the total electricity cost during the battery charging process;

[0029] The charging safety constraints are:

[0030] 0≤I B (k)≤I max (7)

[0031] Where, I max is the maximum allowable charging current of the battery; prevents the battery's SOC, voltage, and temperature from exceeding their allowable limits:

[0032]

[0033] Where, SOC max 、V max and T max Represent the upper limits of battery SOC, voltage and temperature respectively.

[0034] Furthermore, the implementation process of step (2) is as follows:

[0035] The charging speed reward function is set as:

[0036] r s (k)=-k·|SOC(k)-SOC d | (9)

[0037] In the formula, k is the current charging step number, SOC(k) is the SOC state of the battery at charging step k, and SOC d is the target SOC, i.e. the desired SOC when charging is complete;

[0038] The electricity cost reward function is:

[0039] r p (k)=(p max -p(k))·(SOC(k)-SOC(k-1)) (10)

[0040] Where p(k) is the time-of-use electricity price in kilowatt-hours at sampling step k, and p max p(k) represents the highest time-of-use electricity price in the past day;

[0041] During the battery charging process, the battery's charging current, terminal voltage, and internal temperature must be controlled below a specific threshold. The charging current is the control variable of the charge controller and can be constrained by hard constraints. The current constraints are as follows:

[0042] 0 B (k) max (11)

[0043] The battery terminal voltage and internal temperature are affected by the battery model. Even if the charging current meets the constraints, there is a high probability of violating the constraints. The following soft penalties are used:

[0044] ​​

[0045]

[0046]

[0047] When the battery voltage, temperature, and SOC do not exceed their respective maximum allowable thresholds, the battery will not be penalized. When the thresholds are exceeded, the battery will be penalized, thereby guiding the battery to ensure battery safety during charging;

[0048] Assign weights to the above charging goals according to their importance and obtain the total reward function R for each step:

[0049] R(k)=ω1r s (k)+ω2r p (k)+ω3r v (k)+ω4r t (k)+ω5r s' (k) (15)

[0050] Where, ω i (1≤i≤5) are weight coefficients describing different objectives respectively; the optimization objective is transformed into:

[0051]

[0052] Where N is the maximum number of charging steps.

[0053] Furthermore, the charging strategy network and strategy evaluation network implementation process in step (3) are as follows:

[0054] The hidden layers of the charging strategy network and the strategy evaluation network are both fully connected layers. The first layer of the charging strategy network is the ReLU function, and the second layer contains the expectation and variance of the battery charging strategy distribution. The activation function of the expectation part is the Tanh function, and the activation function of the variance part is the SoftPlus function. The strategy evaluation network consists of two layers. The activation function of the first layer is the ReLU function; the activation function of the second layer is the Tanh function. The charging strategy network is:

[0055]

[0056] The policy evaluation network is:

[0057]

[0058] Where w a1 、w a2 、w a3 、w c1 、w c2 are the weight coefficients in the neural network, ba1 、b a2 、b a3 、b c1 、b c2 is the bias in the neural network, the charging strategy network parameters are collectively referred to as θ, and the strategy evaluation network parameters are collectively referred to as δ. θ and δ are continuously updated during the training process of the neural network.

[0059] Furthermore, the interactive training process of the lithium battery fast charging strategy environment and algorithm based on the PPO algorithm in step (3) is as follows:

[0060] The training phase uses an action-evaluation training framework, which consists of a charging strategy network and a strategy evaluation network. The charging strategy network receives the state space information s of the lithium battery charging environment. k =[k,SOC(k),V B (k),T a (k),p(k)] T , output charging current action a k The strategy evaluation network receives the experience information in the experience pool and evaluates the state value function V corresponding to the current strategy network θ. θ (s k ), used to evaluate the pros and cons of the current network charging strategy;

[0061] In the charging strategy network update section, the battery charging target is first defined:

[0062]

[0063] Where, represents the mean value of the interval [0, k], θ represents the network parameter of the current round action, θ old Represents the network parameters of the charging strategy in the last update round; π θ Represents the current round battery charging strategy, Represents the battery charging strategy of the previous round, To estimate the battery charging current a k In battery status k The advantage function under , is calculated by the state value function estimated by the policy evaluation network;

[0064] In the action-evaluation network training framework, the policy evaluation network outputs the state value function V θ (s k ), and calculate the advantage function by the following formula:

[0065]

[0066] Where, represents the charging environment state s, r of the kth charging stage in the nth round randomly drawn from the experience pool k represents the charging reward obtained in the kth charging cycle, γ is the reward discount factor; N is the total number of steps in a charging round;

[0067] The PPO algorithm is used to tailor the battery charging target:

[0068]

[0069] Where:

[0070]

[0071] Among them, L CLIP (θ) implements a trust region correction method compatible with stochastic gradient descent, simplifies the algorithm and reduces the need for adaptive correction by eliminating the KL loss; the charging strategy network updates its own network parameters θ by achieving this goal; in the strategy evaluation network, the traditional time difference error TD-error method is used to update the network parameters δ.

[0072] Beneficial effects: Compared with the existing technology, the beneficial effects of the present invention are: the present invention can achieve comprehensive optimization of several conflicting objectives such as charging speed, battery charging physical constraints and battery charging electricity costs, realize fast charging with active awareness of safety and health and low charging cost, and migrate the complex calculations caused by multi-constraint and multi-objective optimization solutions to the offline training link, significantly reducing the computational complexity of online charging decisions. BRIEF DESCRIPTION OF THE DRAWINGS

[0073] Figure 1 Schematic diagram of the coupled electrothermal model structure of a cylindrical lithium battery;

[0074] Figure 2 This is a schematic diagram of the charging strategy network structure;

[0075] Figure 3 Flowchart for training the battery charging control network;

[0076] Figure 4 Comparison of charging curves between the charging strategy and the traditional CCCV charging algorithm. DETAILED DESCRIPTION

[0077] The present invention will be further described in detail below with reference to the accompanying drawings.

[0078] This invention addresses the problem of battery charging, comprehensively considering multiple objectives such as fast charging requirements, economic costs, and related safety constraints. It provides a battery fast charging control method based on the PPO algorithm and taking into account charging electricity costs. First, the multi-objective charging is converted into a reward function, and a charging strategy neural network and a strategy evaluation neural network are trained to maximize the reward function. Then, the trained action neural network is used to intelligently adjust the charging current according to the current electricity price and battery SOC. This method can meet the battery charging safety constraints while meeting the fast charging requirements and reducing charging costs. Specifically, it includes the following steps:

[0079] Step 1: Establish a lithium-ion battery electrothermal coupling model and a charging electricity cost optimization model, and establish an offline training scenario based on the constructed model to determine the battery capacity SOC, battery voltage V B , average battery temperature T a and the key state variables of electricity price p.

[0080] (1.1) Battery model

[0081] The present invention adopts Figure 1 The electrothermal coupling model shown is used to describe the characteristics of lithium-ion batteries. Figure 1 In the example, the voltage source V OC The resistors R0 and R1 are used to simulate the energy storage and charge and discharge energy loss of the battery, respectively. The RC networks (R1, C1) and (R2, C2) characterize the short-term and long-term transient responses of the battery.

[0082] According to Kirchhoff's current and voltage laws, the dynamic characteristics of the battery can be described as:

[0083]

[0084] Where, SOC(k), C n , I B (k), V B (k) represents the SOC state, nominal capacity, charging current and voltage of the battery at time k; T represents an operating cycle; the battery open circuit voltage V OC (k) is a nonlinear function of SOC(k) and can be expressed as V OC (k) = g(SOC(k)); V1(k) and V2(k) represent the voltages across capacitors C1 and C2, respectively; R0 is a constant resistance. C 、T S represents the battery core temperature and surface temperature, which can be calculated according to the principle of energy conservation as [17-18]

[0085]

[0086]

[0087] Where T amb is the ambient temperature of the battery; R C 、R u Represents thermal conduction resistance and convection resistance respectively; C C 、C S Represent the internal heat capacity and surface heat capacity of the battery respectively. The temperature of the battery is defined as T S and T C The average value of:

[0088]

[0089] (1.2) Charging targets and safety constraints

[0090] In charging control, it is necessary to consider the charging time, the optimization target of the electricity cost required for charging, and the charging safety constraints.

[0091] 1) Fast charging

[0092] Since long charging time will make lithium batteries inconvenient to use, improving charging speed is one of the most important goals to be considered in charging control. Improving the charging speed of lithium batteries means minimizing the time it takes for the battery to charge from any initial charge SOC (0) to the desired charge SOC. d The objective function corresponding to the charging speed can be expressed as

[0093] min J1=NT (5)

[0094] Where T represents the sampling period, N is SOC(N)=SOC d The corresponding number of sampling steps.

[0095] 2) Electricity cost optimization goal

[0096] Optimizing the electricity cost during battery charging is another important objective to consider. Electricity prices around the world typically fluctuate with peaks and valleys in daily electricity consumption, defined as time-of-use pricing. In practical applications, optimizing the charging current during different electricity price periods to reduce electricity costs is valuable and necessary. The charging cost is affected by the battery's charging current and the current electricity price. The objective function for optimizing the charging cost can be expressed as:

[0097]

[0098] Where p(k) is the time-of-use electricity price at charging sampling period k; J2 is the total electricity cost during the battery charging process.

[0099] 3) Charging safety constraints

[0100] Since high charging current can cause irreversible damage to the battery, it should be kept within an appropriate range, which can be expressed as the following input constraints:

[0101] 0≤I B (k)≤I max (7)

[0102] Where, I max is the maximum allowable charging current of the battery. In addition, overcharging, undercharging, and excessive battery temperature can lead to accelerated degradation of battery capacity and even safety issues. To avoid these situations, the battery's SOC, voltage, and temperature must be prevented from exceeding their allowable limits:

[0103]

[0104] Where, SOC max 、V max and T max Represent the upper limits of battery SOC, voltage and temperature respectively.

[0105] Step 2: Normalize the key state variables determined in step 1 and place them in the reinforcement learning state space, defining the action space and reward function.

[0106] 1) Charging speed reward function

[0107] In terms of battery charging speed, the objective function in formula (5) is converted into the difference between the current SOC of the battery and the expected SOC. This means that when the battery is charged with a low charging current or not charged, it will receive a smaller penalty, and charging with a high charging current will receive more rewards. This can guide the battery to pursue a high charging current during training, thereby shortening the charging time. The reward function for each step is set as:

[0108] r s (k)=-k·|SOC(k)-SOC d | (9)

[0109] In the formula, k is the current charging step number, SOC(k) is the SOC state of the battery at charging step k, and SOC d is the target SOC, the desired SOC upon completion of charging. The battery's objective is the expected total reward over future time steps. As charging time increases, not charging causes the reward of the charging strategy to decrease in a gradient, thus encouraging the charging strategy to charge quickly.

[0110] 2) Electricity cost reward function

[0111] To minimize the total electricity cost of battery charging, it is necessary to encourage batteries to charge at high current during valley hours. When electricity prices are at their peak, battery charging will not receive any rewards. When electricity prices are at other times (i.e., valley hours and between peak and valley hours), battery charging will receive certain rewards, thereby guiding batteries to charge during low electricity prices. Based on this, the reward function is set as:

[0112] r p (k)=(p max -p(k))·(SOC(k)-SOC(k-1)) (10)

[0113] Where p(k) is the time-of-use electricity price in kilowatt-hours at sampling step k, and p max p(k) represents the highest time-of-use electricity price in the past day.

[0114] 3) Charging safety constraints

[0115] During battery charging, the battery's charging current, terminal voltage, and internal temperature must be controlled below a specific threshold. The charging current is a control variable of the charge controller and can be constrained using hard constraints. The current constraints are as follows:

[0116] 0 B (k) max (11)

[0117] The battery terminal voltage and internal temperature are affected by the battery model. Even if the charging current meets the constraints, there is a high possibility of violating the constraints. Directly using the hard constraints in formula (8) may destroy the exploration process of reinforcement learning. Therefore, the following soft penalty is adopted:

[0118]

[0119]

[0120]

[0121] When the battery voltage, temperature, and SOC do not exceed their respective maximum allowable thresholds, the battery will not be penalized. When they exceed the thresholds, the battery will be penalized, thereby guiding the battery to ensure battery safety during charging. Since battery safety is extremely important in practical applications, a large weight coefficient should be used to ensure that the battery does not overcharge and the voltage and temperature do not exceed the upper limit.

[0122] 3) Total Reward Function

[0123] Assign weights to the above charging goals according to their importance and obtain the total reward function R for each step:

[0124] ​​R(k)=ω1r s (k)+ω2r p (k)+ω3r v (k)+ω4r t (k)+ω5r s' (k) (15)

[0125] Where, ω i (1≤i≤5) are weight coefficients describing different objectives respectively; the optimization objective is transformed into:

[0126]

[0127] Where N is the maximum number of charging steps. The present invention maximizes the sum of the reward functions of all steps, thereby achieving multi-objective optimization of fast charging and saving electricity costs.

[0128] Step 3: Build the charging strategy network and the strategy evaluation network, and train them using the proximal policy optimization (PPO) algorithm. The charging strategy network generates charging actions based on the acquired state variables, updates the battery state according to the battery model from step 1, and records the charging actions, battery state, and reward value in the experience pool. This information is then used to synchronously update the charging strategy network and the strategy evaluation network.

[0129] The hidden layers of the charging strategy network and the strategy evaluation network are both fully connected layers. The first layer of the charging strategy network is the relu function, and the second layer contains the expectation and variance of the battery charging strategy distribution. The activation function of the expectation part is the tanh function, and the activation function of the variance part is the softplus function. The strategy evaluation network contains two layers of networks. The activation function of the first layer is the relu function; the activation function of the second layer is the tanh function. The schematic diagram of the charging strategy network is as follows Figure 2 As shown, the expression is:

[0130]

[0131] The policy evaluation network can be expressed as:

[0132]

[0133] Where w a1 、w a2 、w a3 、w c1 、w c2 are the weight coefficients in the neural network, b a1 、b a2 、b a3 、b c1 、b c2is the bias in the neural network, the charging strategy network parameters are collectively referred to as θ, and the strategy evaluation network parameters are collectively referred to as δ. θ and δ are continuously updated during the training process of the neural network.

[0134] The interactive training process of the lithium battery fast charging strategy environment and algorithm based on the PPO algorithm is as follows Figure 3 As shown in the figure, the training phase adopts the action-evaluation training framework, which consists of two parts: the charging strategy network and the strategy evaluation network. The charging strategy network receives the state space information s of the lithium battery charging environment. k =[k,SOC(k),V B (k),T a (k),p(k)] T , output charging current action a k The strategy evaluation network receives the experience information in the experience pool and evaluates the state value function V corresponding to the current strategy network θ θ (s k ), which is used to evaluate the pros and cons of the current network charging strategy.

[0135] In the charging strategy network update section, the battery charging target is first defined:

[0136]

[0137] Where, represents the mean value of the interval [0, k], θ represents the network parameter of the current round action, θ old Represents the network parameters of the charging strategy in the last update round. θ Represents the current round battery charging strategy, Represents the battery charging strategy of the previous round, To estimate the battery charging current a k In battery status k The advantage function under is calculated by the state value function estimated by the policy evaluation network.

[0138] In the action-evaluation network training framework, the policy evaluation network outputs the state value function V θ (s k ), and calculate the advantage function by the following formula:

[0139]

[0140] Where, represents the charging environment state s, r of the kth charging stage in the nth round randomly drawn from the experience pool k represents the charging reward obtained in the kth charging cycle, γ is the reward discount factor, and N is the total number of steps in a charging round.

[0141] In order to reduce the complexity of the calculation and speed up the training convergence, the PPO algorithm tailors the above battery charging target:

[0142]

[0143] Where,

[0144]

[0145] Among them, L CLIP (θ) implements a trust region correction method compatible with stochastic gradient descent, eliminating the KL loss to simplify the algorithm and reduce the need for adaptive corrections. The charging policy network updates its own network parameters θ by achieving this goal. The policy evaluation network uses the traditional temporal difference error (TD-error) method to update the network parameters δ.

[0146] Step 4: Loop through step 3 until the policy network and policy evaluation network converge, and derive the charging policy network as the battery fast charging strategy.

[0147] Step 5: The battery's current charge level, terminal voltage, ambient temperature, battery surface temperature, and current electricity price are collected in real time, normalized, and fed into the charging strategy network trained in Step 4 to determine the optimal charging action at the current moment. After each charging cycle, the state variables are collected again and the charging current is determined until charging is complete.

[0148] The algorithm flow of the charging strategy network execution phase is as follows:

[0149] 1) Initialize battery status;

[0150] 2) Get the current battery status s k =[k,SOC(k),V B (k),T a (k),p(k)] T ;

[0151] 3) The charging strategy network receives status information and outputs (μ k ,σ k )=F(s k ), charging strategy distribution The charging current a is sampled and output in this normal distribution k ;

[0152] 4) If SOC(k+1)=1, charging is completed; otherwise, k=k+1 and return to step 2).

[0153] In this embodiment, the battery fast charging control method considering the charging electricity cost is verified and compared with the widely used CCCV method (including CCCV-1C, CCCV-2C, CCCV-3C). The comparison results are as follows: Figure 4 As shown in Table 1, Figure 4 (a) is a battery capacity comparison chart; Figure 4 (b) is a comparison diagram of charging current response; Figure 4 (c) is a comparison diagram of the battery terminal voltage response; Figure 4 Figure (d) shows the average battery temperature response. Using a PPO-based charge control algorithm, charging is faster than using a CC-CV algorithm with a 1C charge current. While charging time is slightly longer compared to the 2C and 3C CCCV algorithms, the electricity cost is 0.0090 yuan less than using a 2C CCCV algorithm and 0.0092 yuan less than using a 3C CC-CV algorithm. These results indicate that the CCCV algorithm focuses solely on charging time and fails to optimize battery electricity costs.

[0154] Table 1 Comparison of battery charging algorithm effects

[0155]

[0156] The results show that compared with the traditional CC-CV algorithm, the present invention can effectively reduce the charging cost while improving the charging speed, thereby alleviating the user's charging economic pressure.

[0157] The above description is merely a preferred embodiment of the present invention and does not constitute any form of limitation to the present invention. Those skilled in the art may make various equivalent changes and improvements based on the above embodiment. All equivalent changes and modifications made within the scope of the claims shall fall within the scope of protection of the present invention.

Claims

1. A battery fast charging control method based on the PPO algorithm and taking into account the charging electricity cost, characterized in that: The following steps are involved: (1) Construct a lithium-ion battery electrothermal coupling model and a charging electricity cost optimization model, and establish an offline training scenario based on the two constructed models to determine their key state variables; (2) Normalize the key state variables determined in step (1) and place them in the reinforcement learning state space, defining the action space and reward function; (3) Training the constructed charging strategy network and strategy evaluation network based on the proximal strategy optimization algorithm; the charging strategy network generates a charging action according to the acquired state variables, updates the battery state according to the lithium-ion battery electrothermal coupling model in step (1), and records the charging action, battery state, and reward value in the experience pool, and synchronously updates the charging strategy network and strategy evaluation network through the experience pool information; (4) Execute step (3) repeatedly until the charging strategy network and the strategy evaluation network converge, and derive the charging strategy network as the battery fast charging strategy; (5) Real-time collection of the battery's current power, terminal voltage, ambient temperature, battery surface temperature, and current electricity price, and normalization of these data, which are then input into the charging strategy network trained in step (4) to determine the optimal charging action at the current moment; (6) After each charging cycle, re-collect the state quantity and decide the charging current until charging is completed; The process of constructing the lithium-ion battery electrothermal coupling model in step (1) is as follows: Voltage source V OC The resistors R0 and R1 are used to simulate the battery's energy storage and charge / discharge energy loss, respectively. The RC networks (R1, C1) and (R2, C2) characterize the battery's short-term and long-term transient responses. According to Kirchhoff's current and voltage laws, the battery's dynamic characteristics are described as follows: Where, SOC(k), C n , I B (k), V B (k) represents the battery's SOC state, nominal capacity, charging current, and voltage; the battery's open circuit voltage V OC (k) is a nonlinear function of SOC(k): V OC (k) = g(SOC(k)); V1(k) and V2(k) represent the voltages across capacitors C1 and C2, respectively; R0 is a constant resistance; T C 、T S represents the battery core temperature and surface temperature, which are calculated according to the principle of conservation of energy as: Where, T amb is the ambient temperature of the battery; R C 、R u Represents thermal conduction resistance and convection resistance respectively; C C 、C S Represent the internal heat capacity and surface heat capacity of the battery respectively; the temperature of the battery is defined as T S and T C The average value of: The process of constructing the charging electricity fee optimization model in step (1) is as follows: Minimize the time it takes for the battery to charge from any initial SOC (0) to the desired SOC d The time spent and the objective function corresponding to the charging speed are: minJ1=NT(5) Where T represents the sampling period, N is SOC(N)=SOC d The corresponding number of sampling steps; The charging cost is affected by the battery's charging current and the current electricity price. The objective function for optimizing the charging cost is: Where p(k) is the time-of-use electricity price at charging sampling period k; J2 is the total electricity cost during the battery charging process; The charging safety constraints are: 0≤I B (k)≤I max (7) Where, I max is the maximum allowable charging current of the battery; prevents the battery's SOC, voltage, and temperature from exceeding their allowable limits: Where, SOC max 、V max and T max Represent the upper limits of battery SOC, voltage and temperature respectively; The implementation process of the charging strategy network and strategy evaluation network in step (3) is as follows: The hidden layers of the charging strategy network and the strategy evaluation network are both fully connected layers. The first layer of the charging strategy network is the ReLU function, and the second layer contains the expectation and variance of the battery charging strategy distribution. The activation function of the expectation part is the Tanh function, and the activation function of the variance part is the SoftPlus function. The strategy evaluation network consists of two layers. The activation function of the first layer is the ReLU function; the activation function of the second layer is the Tanh function. The charging strategy network is: The policy evaluation network is: Where w a1 、w a2 、w a3 、w c1 、w c2 are the weight coefficients in the neural network, b a1 、b a2 、b a3 、b c1 、b c2 is the bias in the neural network, the charging strategy network parameters are collectively referred to as θ, and the strategy evaluation network parameters are collectively referred to as δ. θ and δ are continuously updated during the training process of the neural network; The interactive training process of the lithium battery fast charging strategy environment and algorithm based on the PPO algorithm in step (3) is as follows: The training phase uses an action-evaluation training framework, which consists of a charging strategy network and a strategy evaluation network. The charging strategy network receives the state space information s of the lithium battery charging environment. k =[k,SOC(k),V B (k),T a (k),p(k)] T , output charging current action a k The strategy evaluation network receives the experience information in the experience pool and evaluates the state value function V corresponding to the current strategy network θ. θ (s k ), used to evaluate the pros and cons of the current network charging strategy; In the charging strategy network update section, the battery charging target is first defined: Where, represents the mean value of the interval [0, k], θ represents the network parameter of the current round action, θ old Represents the network parameters of the charging strategy in the last update round; π θ Represents the current round battery charging strategy, Represents the battery charging strategy of the previous round, To estimate the battery charging current a k In battery status k The advantage function under , is calculated by the state value function estimated by the policy evaluation network; In the action-evaluation network training framework, the policy evaluation network outputs the state value function V θ (s k ), and calculate the advantage function by the following formula: Where, represents the charging environment state s, r of the kth charging stage in the nth round randomly drawn from the experience pool k represents the charging reward obtained in the kth charging cycle, γ is the reward discount factor; N is the total number of steps in a charging round; The PPO algorithm is used to tailor the battery charging target: Where: Among them, L CLIP (θ) implements a trust region correction method compatible with stochastic gradient descent, simplifies the algorithm and reduces the need for adaptive correction by eliminating the KL loss; the charging strategy network updates its own network parameters θ by achieving this goal; in the strategy evaluation network, the traditional time difference error TD-error method is used to update the network parameters δ.

2. A battery fast charging control method based on the PPO algorithm and considering charging electricity costs according to claim 1, characterized in that: The key state variables in step (1) include battery charge SOC, battery voltage V B , average battery temperature T a and electricity price p.

3. The battery fast charging control method based on the PPO algorithm and considering charging electricity costs according to claim 1, characterized in that: The implementation process of step (2) is as follows: The charging speed reward function is set as: r s (k)=-k·|SOC(k)-SOC d |(9) In the formula, k is the current charging step number, SOC(k) is the SOC state of the battery at charging step k, and SOC d is the target SOC, i.e. the desired SOC when charging is complete; The electricity cost reward function is: r p (k)=(p max -p(k))·(SOC(k)-SOC(k-1))(10) Where p(k) is the time-of-use electricity price in kilowatt-hours at sampling step k, and p max p(k) represents the highest time-of-use electricity price in the past day; During the battery charging process, the battery's charging current, terminal voltage, and internal temperature must be controlled below a specific threshold. The charging current is the control variable of the charge controller and can be constrained by hard constraints. The current constraints are as follows: 0<I B (k)<I max (11) The battery terminal voltage and internal temperature are affected by the battery model. Even if the charging current meets the constraints, there is a high probability of violating the constraints. The following soft penalties are used: When the battery voltage, temperature, and SOC do not exceed their respective maximum allowable thresholds, the battery will not be penalized. When the thresholds are exceeded, the battery will be penalized, thereby guiding the battery to ensure battery safety during charging; Assign weights to the above charging goals according to their importance and obtain the total reward function R for each step: R(k)=ω1r s (k)+ω2r p (k)+ω3r v (k)+ω4r t (k)+ω5r s′ (k) (15) Where, ω i (1≤i≤5) are weight coefficients describing different objectives respectively; the optimization objective is transformed into: Where N is the maximum number of charging steps.

Citation Information

Patent Citations

  • Unmanned aerial vehicle target tracking control method based on reinforcement learning PPO algorithm

    CN111580544A

  • Multi-physical field constrained intelligent quick charging method for lithium ion battery

    CN112018465A

  • Multi-objective optimization charging control method for lithium battery

    CN112886674A

  • Electric vehicle layered charging strategy method based on probability transfer matrix

    CN113919229A

  • Electric vehicle charging and discharging strategy optimization method based on interior point strategy optimization

    CN114997935A