A battery pack equalization method based on reinforcement learning

By employing a reinforcement learning-based battery pack balancing method, and utilizing the Actor-Critic architecture and TD3 algorithm to optimize the battery pack balancing strategy, the problem of output oscillation and energy waste caused by reliance on experience in existing technologies is solved, and more efficient battery pack balancing control is achieved.

CN116674431BActive Publication Date: 2026-04-24FUZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
FUZHOU UNIV
Filing Date
2023-05-27
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing rule-based battery pack equalization control methods rely on subjective experience, which may lead to output oscillations and over-equalization, reducing equalization efficiency, and causing energy waste when battery pack consistency decreases.

Method used

A battery pack balancing method based on reinforcement learning is adopted. The Actor-Critic architecture and the dual-delay deep deterministic policy gradient algorithm (TD3) are used to optimize the balancing strategy through a deep learning network and design a simple reward function to achieve self-learning and stable control.

Benefits of technology

It shortens the battery pack equalization time, reduces energy waste, avoids over-equalization, and improves equalization efficiency and battery pack lifespan.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116674431B_ABST
    Figure CN116674431B_ABST
Patent Text Reader

Abstract

The application relates to a battery pack equalization method based on reinforcement learning, which comprises the following steps: determining equalization targets and constraint conditions of a battery pack equalization process according to rated capacities of single batteries in the battery pack and equalization topological parameters in an equalization system; establishing an action space of an equalization system agent by using equalization current control amounts of a battery pack equalizer, and establishing a state space of the equalization system agent by using inconsistency state information of the battery pack and equalization current control amounts generated by the agent under the state information; establishing a deep learning network of an Actor-Critic architecture, and constructing a deep reinforcement learning equalization strategy based on a double-delay deep deterministic policy gradient algorithm; designing a battery equalization system reward function, training the deep reinforcement learning equalization strategy, and randomly initializing SOC states of the single batteries in each training round; and using the trained reinforcement learning equalization strategy to perform battery pack equalization control. The method is beneficial to shortening a battery pack equalization time and reducing energy waste in a battery pack equalization process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of battery pack equalization control technology, and specifically to a battery pack equalization method based on reinforcement learning. Background Technology

[0002] With the development of the new energy vehicle industry, the demand for and scrapping of battery packs are rapidly increasing. To meet people's travel needs, the capacity of battery packs is constantly increasing, which in turn requires an increase in the number of batteries connected in parallel. However, the performance of battery packs will degrade with the increase of charging cycles. When the consistency of the battery pack decreases, it is easy for a certain cell in the battery pack to be overcharged or over-discharged, resulting in energy waste and seriously affecting the service life of the battery pack.

[0003] Therefore, battery pack equalization control is necessary to eliminate various inconsistencies arising from the battery pack itself and during use. Currently, most battery pack equalization control methods are rule-based active equalization control. The principle involves compiling operator or expert experience into fuzzy rules, then fuzzifying real-time signals from sensors. The fuzzified signals are used as input to the fuzzy rules to complete fuzzy inference, and the resulting output is added to the actuator. Therefore, the formulation of fuzzy rules relies on subjective experience, and improper rule design may lead to output oscillations, reduced equalization efficiency, or even over-equalization. Therefore, an equalization strategy is needed to address these issues. Summary of the Invention

[0004] The purpose of this invention is to provide a battery pack equalization method based on reinforcement learning, which helps to shorten the battery pack equalization time and reduce energy waste during the battery pack equalization process.

[0005] To achieve the above objectives, the technical solution adopted by this invention is: a battery pack equalization method based on reinforcement learning, comprising the following steps:

[0006] Step 1: Determine the balancing objective and constraints of the battery pack balancing process based on the rated capacity of the individual cells in the battery pack and the balancing topology parameters in the balancing system.

[0007] Step 2: Establish the action space of the equalization system agent using the equalization current control quantity of the battery pack equalizer, and establish the state space of the equalization system agent using the inconsistent state information of the battery pack and the equalization current control quantity generated by the agent under the state information.

[0008] Step 3: Build a deep learning network with an Actor-Critic architecture, and based on this, construct a deep reinforcement learning equilibrium strategy based on a dual-delay deep deterministic policy gradient algorithm;

[0009] Step 4: Design the reward function of the battery equalization system, initialize the training parameters of the deep reinforcement learning equalization strategy, train the deep reinforcement learning equalization strategy, and randomly initialize the SOC state of a single battery in each training round.

[0010] Step 5: Use the trained reinforcement learning balancing strategy to perform battery pack balancing control.

[0011] Furthermore, in step 1, the rated capacity of a single battery cell is determined by diag[C1,C2,...,C]. n ] T Indicated by C, which represents the rated capacity of the corresponding single cell, the bidirectional adjacent equalization topology parameters are expressed as follows:

[0012]

[0013] Where m = 2n, η d For battery charging coulomb efficiency, η c The battery discharge coulomb efficiency;

[0014] The equilibrium objective of the equilibrium process is:

[0015]

[0016] Where ε is the maximum allowable error for battery pack consistency, ΔSOC = SOC max -SOC min Indicates the range of battery SOC;

[0017] The constraints of the equilibrium process are:

[0018]

[0019] Among them, I ch,i For the charging current, I ch,max I is the maximum rechargeable current of the battery. dis,i I is the discharge current. dis,max I is the maximum discharge current of the battery. eq,i For ICE current, I eq,max This is the maximum equalizing current for the battery equalizer.

[0020] Furthermore, in step 2, the equalization current control quantity is expressed as u = [u1, u2, ..., u...]. i ,…,u N ] represents, where u∈[-1,1], and the positive or negative sign indicates the charging or discharging state, u i This represents the current equalization current control value of the current equalization battery, u i Further represented as u i1 +u i2 In a bidirectional adjacent topology, ui1 u i2 =0, when u i When u ≤ 0 i1 =0, u i2 =|u i |, when u i When >0, u i1 =|u i |,u i2 =0;

[0021] The action space of the agent is a = [u 11 +u 12 ,…,u i1 +u i2 ,…,u N1 +u N2 ], converting the action space into a standard equalization control vector u = [u 11 ,u 12 ,u i1 ,u i2 ,…,u N1 ,u N2 ] T m×1 , where m=2N, and after conversion u∈[0,1] is the equivalent equalization current coefficient or the duty cycle of the MOSFET;

[0022] Battery pack inconsistency status information includes the SOC difference between individual cells. diff Given the range of individual battery cells ΔSOC, the state space of the agent is s = [SOC]. diff [,a,ΔSOC], where SOC diff =[SOC diff1 SOC diff2 ,…,SOC diffj ,...,SOC diffN SOC diffj Let SOC be the difference between the j-th and (j+1)-th individual cells. diffN This represents the SOC difference between the first and last individual cells.

[0023] Furthermore, the SOC difference between individual cells diff It can be obtained from the following formula:

[0024] SOC diff =T ref x(t)

[0025] Where x is the SOC state matrix of a single cell in the balanced battery pack, and T ref Represented as:

[0026]

[0027] Furthermore, in step 3, the deep learning network based on the deep reinforcement learning equalization strategy using the dual-delay deep deterministic policy gradient algorithm consists of one Actor network and two Critic networks. The Critic network comprises two input layers, four fully connected layers, and one output layer, with ReLU activation. The Actor network comprises one input layer, three fully connected layers, and one output layer, with ReLU activation. The Actor network parameters are updated with a delay through the dual Critic network.

[0028] After the state space and action space of the intelligent agent in the battery pack balancing system are determined, the parameters of the Critic network and Actor network are initialized.

[0029] Furthermore, in step 4, the reward function of the battery balancing system is:

[0030] R = CJ

[0031] Where C is a constant, the function J is expressed as:

[0032]

[0033] The coefficients of function J are K = [k1 k2…k n ] 1×N The specific value of K is obtained by the following formula:

[0034] When (SOC) i -SOC i+1 )×(SOC i _ init -SOC i+1 _ init When k ≥ 0, i =a1;

[0035] When (SOC) n -SOC1)×(SOC n _ init -SOC1_ init When k ≥ 0, n =a1;

[0036] When (SOC) i -SOC i+1 )×(SOC i _ init -SOC i+1 _ init When k < 0, i =a1×a2;

[0037] When (SOC) n -SOC1)×(SOCn _ init -SOC1_ init When k < 0, n = a1 × a2, where a1 ≥ 1, a2 > 1;

[0038] In the above relationship, i = 1, 2, ..., n-1; SOC i The current SOC value of the battery. i_init The initial value of the battery's SOC is used to determine whether the battery pack has experienced over-balancing during the balancing process. a1 and a2 are penalty coefficients. a1 is used to penalize inconsistencies to quickly improve consistency, and a2 is used to penalize balancing processes that have experienced over-balancing.

[0039] Compared with existing technologies, this invention has the following advantages: Compared with traditional rule-based equilibrium strategy formulation, this method can enable the equilibrium controller to continuously improve its equilibrium control decisions through an exploration process, gradually approaching the optimal equilibrium target within the control domain, thus achieving self-learning design of the equilibrium management strategy; This method enables the agent to automatically explore the optimal equilibrium current within constraints, while effectively shortening the time required for battery pack equilibrium and reducing energy waste during the equilibrium process; Furthermore, this invention employs the TD3 algorithm based on reinforcement learning, using a dual-Critic network battery pack equilibrium training framework, using the minimum value between the two Critics to suppress overestimation of the Q value, adding perturbations to the action of the next state when calculating the target value, thereby making the value assessment more accurate, and updating the Actor network after multiple updates to the Critic network, thus ensuring more stable training of the Actor network; This invention also proposes a novel and simple reward function, which enables the agent to explore in the direction of optimal equilibrium effect within constraints. Attached Figure Description

[0040] Figure 1 This is a flowchart illustrating the method implementation of an embodiment of the present invention;

[0041] Figure 2 This is a schematic diagram of a deep learning network in an embodiment of the present invention;

[0042] Figure 3 This is a balanced topology diagram in an embodiment of the present invention;

[0043] Figure 4 This is a diagram illustrating the balancing effect of a rule-based battery pack balancing strategy in an embodiment of the present invention.

[0044] Figure 5 This is a diagram showing the active balancing effect of the battery pack after training in an embodiment of the present invention.

[0045] Figure 6 The graph shows the average SOC variation curves of the rule-based battery pack equalization strategy and the deep reinforcement learning equalization strategy of this method in the embodiments of the present invention. Detailed Implementation

[0046] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0047] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of this application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.

[0048] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments according to this application. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0049] In this embodiment, the battery pack consists of five 18650 individual cells, model SONY-VTC6, with parameters shown in Table 1. The inductance and capacitance parameters of the equalizer main circuit components are 220uH and 220uF, respectively.

[0050] Table 118650 Lithium Battery Parameters

[0051]

[0052] like Figure 1 This embodiment provides a battery pack equalization method based on reinforcement learning, including the following steps:

[0053] Step 1: Using the Buck-boost converter as the equalizer, determine the equalization target and constraints of the battery pack equalization process based on the rated capacity of the individual cells in the battery pack and the equalization topology parameters in the equalization system.

[0054] Step 2: Establish the action space of the equalization system agent using the equalization current control quantity of the battery pack equalizer, and establish the state space of the equalization system agent using the inconsistent state information of the battery pack and the equalization current control quantity generated by the agent under the state information.

[0055] Step 3: Build a deep learning network with an Actor-Critic architecture, and based on this, construct a deep reinforcement learning equilibrium strategy based on a dual-delay deep deterministic policy gradient algorithm.

[0056] Step 4: Considering factors such as battery pack balancing path optimization, balancing current and balancing time constraints, design the battery balancing system reward function based on the initial inconsistent state information and current inconsistent state information of the battery pack and the current balancing action of the balancing system agent, initialize the training parameters of the deep reinforcement learning balancing strategy, then train the deep reinforcement learning balancing strategy, and randomly initialize the SOC state of individual batteries in each training round.

[0057] Step 5: Use the trained reinforcement learning balancing strategy to perform battery pack balancing control.

[0058] In step 1, the battery pack balancing topology is as follows: Figure 3 As shown, the rated capacity of a single battery cell is defined by diag[C1,C2,...,C]. n ] T Indicated by C, which represents the rated capacity of the corresponding single cell, the bidirectional adjacent equalization topology parameters are expressed as follows:

[0059]

[0060] Where m = 2n, η d For battery charging coulomb efficiency, η c The discharge coulombic efficiency of the battery.

[0061] The equilibrium objective of the equilibrium process is:

[0062]

[0063] Where ε is the maximum allowable error for battery pack consistency, ΔSOC = SOC max -SOC min This indicates the range of battery SOC.

[0064] The constraint condition for the equalization process is the charging current I. ch,i and discharge current I dis,i and ICE current I eq,i The size relationship between them should satisfy:

[0065]

[0066] Among them, I ch,i For the charging current, I ch,max I is the maximum rechargeable current of the battery. dis,i I is the discharge current. dis,max I is the maximum discharge current of the battery. eq,i For ICE current, I eq,max This is the maximum equalizing current for the battery equalizer.

[0067] Assuming the I of each battery and ICE circuit module ch,max Idis,max and I eq,max If they are the same, then I eq,i Satisfying the relation:

[0068]

[0069] In step 2, the equalization current control quantity is expressed as u = [u1, u2, ..., u i ,…,u N ] represents, where u∈[-1,1], and the positive or negative sign indicates the charging or discharging state, u i This represents the current equalization current control value of the current equalization battery, u i Further represented as u i1 +u i2 In a bidirectional adjacent topology, u i1 u i2 =0, when u i When u ≤ 0 i1 =0, u i2 =|u i |, when u i When >0, u i1 =|u i |,u i2 =0.

[0070] The action space of the agent is a = [u 11 +u 12 ,…,u i1 +u i2 ,…,u N1 +u N2 ], converting the action space into a standard equalization control vector u = [u 11 ,u 12 ,u i1 ,u i2 ,…,u N1 ,u N2 ] T m×1 Where m = 2N, and after conversion u∈[0,1] is the equivalent equalization current coefficient or the duty cycle of the MOSFET. In this embodiment, the action space of the agent is a = [u 11 +u 12 ,…,u 51 +u 52 ], converting the action space into a standard equalization control vector u = [u 11 ,u 12 ,…,u 51 ,u 52 ] T 10×1 .

[0071] Battery pack inconsistency status information includes the SOC difference between individual cells. diff Given the range of individual battery cells ΔSOC, the state space of the agent is s = [SOC]. diff [,a,ΔSOC];where SOC diff =[SOC diff1 SOC diff2 ,…,SOC diffj ,...,SOC diffN SOC diffj Let SOC be the difference between the j-th and (j+1)-th individual cells. diffN This represents the SOC difference between the first and last individual cells.

[0072] In this embodiment, the SOC difference between individual cells is... diff =[SOC diff1 SOC diff2 SOC diff3 SOC diff4 SOC diff5 SOC diff It can be obtained from the following formula:

[0073] SOC diff =T ref x(t)

[0074] Where x is the SOC state matrix of a single cell in the balanced battery pack, and T ref Represented as:

[0075]

[0076] In step 3, the deep learning network based on the deep reinforcement learning equalization strategy of the dual-delay deep deterministic policy gradient algorithm consists of one Actor network and two Critic networks. The Critic network consists of two input layers, four fully connected layers and one output layer, with the ReLU activation function. The Actor network consists of one input layer, three fully connected layers and one output layer, with the ReLU activation function. The Actor network parameters are updated with a delay through the dual Critic network.

[0077] After the state space and action space of the intelligent agent in the battery pack balancing system are determined, the parameters of the Critic network and Actor network are initialized.

[0078] Furthermore, a deep learning network architecture based on the Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithm is designed. The deep learning network consists of one Actor network and two Critic networks, and the pseudocode of the TD3 algorithm is shown in Table 2.

[0079] Table 2 TD3 Pseudocode

[0080]

[0081]

[0082] In step 4, the reward function of the battery balancing system is:

[0083] R = CJ

[0084] Where C is a constant, the function J is expressed as:

[0085]

[0086] The coefficients of function J are K = [k1 k2…k n ] 1×N The specific value of K is obtained by the following formula:

[0087] When (SOC) i -SOC i+1 )×(SOC i _ init -SOC i+1 _ init When k ≥ 0, i =a1;

[0088] When (SOC) n -SOC1)×(SOC n _ init -SOC1_ init When k ≥ 0, n =a1;

[0089] When (SOC) i -SOC i+1 )×(SOC i _ init -SOC i+1 _ init When k < 0, i =a1×a2;

[0090] When (SOC) n -SOC1)×(SOC n _ init -SOC1_init When k < 0, n = a1 × a2, where a1 ≥ 1, a2 > 1;

[0091] In the above relationship, i = 1, 2, ..., n-1; SOC i The current SOC value of the battery. i_init The initial value of the battery's SOC is used to determine whether the battery pack has experienced over-balancing during the balancing process. a1 and a2 are penalty coefficients. a1 is used to penalize inconsistencies to quickly improve consistency, and a2 is used to penalize balancing processes that have experienced over-balancing.

[0092] In this embodiment, the training parameters of the TD3 algorithm agent are shown in Table 3:

[0093] Table 3. Training parameters of the TD3 algorithm agent

[0094]

[0095] In this embodiment, five initial SOC states of the batteries were randomly generated: SOC1 = 0.554, SOC2 = 0.621, SOC3 = 0.570, SOC4 = 0.637, and SOC5 = 0.601. Figure 4 As shown, the rule-based active battery balancing method achieves an balancing time of 527 seconds, and over-balancing and repeated charging / discharging occur both during and after the balancing process. However, as... Figure 5 As shown, the battery pack active balancing system of the present invention reaches balancing time of 256s after training, and no balancing or repeated charging and discharging phenomenon occurs during or after the battery pack balancing process.

[0096] like Figure 5 As shown, the discharge rate of battery 2 is significantly faster than that of battery 5. Furthermore, batteries 2 and 5 are not adjacent to each other, so the intersection of their SOC curves does not affect the output result.

[0097] Both the rule-based equalization strategy and the reinforcement learning-based equalization strategy have a current limit of 1A. The rule-based equalization strategy formulated in this embodiment is as follows:

[0098] Rule 1: When ΔSOC>0.04, the equalization current is 1A, the purpose of which is to achieve fast equalization and shorten the equalization time;

[0099] Rule 2: When 0.03 < ΔSOC < 0.04, the balancing current is 0.7A, the purpose of which is to prevent the inconsistency inside the battery pack from increasing rapidly;

[0100] Rule 3: When 0.02 < ΔSOC < 0.03, the balancing current is 0.4A, the purpose of which is to quickly improve the inconsistency between battery packs;

[0101] Rule 4: When ΔSOC < 0.02, the balancing current is 0.2A, the purpose of which is to avoid over-discharge or over-charge.

[0102] This embodiment performs equalization on a 5-cell serial cell array with a capacity of 3000mAh and a SOC range of 7.4%. Figure 6 As shown, the battery pack capacity loss based on rules is 17.7mAh, while the battery pack capacity loss based on the equalization strategy of TD3 reinforcement learning in this invention is 14.7mAh. Compared with the rule-based approach, the equalization time is improved by 52% and the energy loss is reduced by 17%. The equalization strategy based on TD3 reinforcement learning in this invention can better shorten the time required for battery pack equalization, avoid output oscillation and over-equalization due to over-reliance on experience, and reduce the energy loss and waste of the battery pack during the equalization process.

[0103] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0104] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0105] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1The function specified in one or more boxes.

[0106] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0107] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.

Claims

1. A battery pack equalization method based on reinforcement learning, characterized in that, Includes the following steps: Step 1: Determine the balancing objective and constraints of the battery pack balancing process based on the rated capacity of the individual cells in the battery pack and the balancing topology parameters in the balancing system. Step 2: Establish the action space of the equalization system agent using the equalization current control quantity of the battery pack equalizer, and establish the state space of the equalization system agent using the inconsistent state information of the battery pack and the equalization current control quantity generated by the agent under this state information. Step 3: Build a deep learning network with an Actor-Critic architecture, and based on this, construct a deep reinforcement learning equilibrium strategy based on a dual-delay deep deterministic policy gradient algorithm; Step 4: Design the reward function of the battery equalization system, initialize the training parameters of the deep reinforcement learning equalization strategy, train the deep reinforcement learning equalization strategy, and randomly initialize the SOC state of a single battery in each training round. Step 5: Use the trained reinforcement learning balancing strategy to perform battery pack balancing control; In step 1, the rated capacity of a single battery cell is determined by... Indicated by C, which represents the rated capacity of the corresponding single cell, the bidirectional adjacent equalization topology parameters are expressed as follows: Where m = 2n, η d For battery charging coulomb efficiency, η c The battery discharge coulomb efficiency; The equilibrium objective of the equilibrium process is: Where ε is the maximum allowable error for battery pack consistency, ∆SOC = SOC max -SOC min Indicates the range of battery SOC; The constraints of the equilibrium process are: Among them, I ch,i For the charging current, I ch,max I is the maximum rechargeable current of the battery. dis,i For the discharge current, I dis,max I is the maximum discharge current of the battery. eq,i For ICE current, I eq,max This is the maximum equalizing current for the battery equalizer. In step 2, the equalization current control quantity is expressed as u=[u1,u2,…,u i ,…,u N ] represents, where u∈[-1,1], and the positive or negative sign indicates the charging or discharging state, u i This represents the current equalization current control value of the current equalization battery, u i Further represented as u i1 + u i2 In a bidirectional adjacent topology, u i1 u i2 =0, when u i When u ≤ 0 i1 =0, u i2 =|u i |, when u i When >0, u i1 =|u i |,u i2 =0; The action space of the agent is a = [u 11 +u 12 , …, u i1 +u i2 , …, u N1 +u N2 ], converting the action space into a standard equalization control vector u = [u 11 , u 12 , u i1 , u i2 , …, u N1 , u N2 ] T m×1 Where m=2N, and after conversion u∈[0,1] is the equivalent equalization current coefficient or the duty cycle of the MOSFET; Battery pack inconsistency status information includes the SOC difference between individual cells. diff Given the range ∆SOC of a single battery cell, the state space of the agent is s = [SOC]. diff [, a, ∆SOC], where SOC diff = [SOC diff1 SOC diff2 , …,SOC diffj , ..., SOC diffN SOC diffj Let SOC be the difference between the j-th and (j+1)-th individual cells. diffN This represents the SOC difference between the first and last individual cells.

2. The battery pack equalization method based on reinforcement learning according to claim 1, characterized in that, SOC difference between individual cells diff It can be obtained from the following formula: Where x is the SOC state matrix of a single cell in the balanced battery pack, and T ref Represented as: 。 3. The battery pack equalization method based on reinforcement learning according to claim 1, characterized in that, In step 3, the deep learning network based on the deep reinforcement learning equalization strategy of the dual-delay deep deterministic policy gradient algorithm consists of one Actor network and two Critic networks. The Critic network consists of two input layers, four fully connected layers and one output layer, with the ReLU activation function. The Actor network consists of one input layer, three fully connected layers and one output layer, with the ReLU activation function. The Actor network parameters are updated with a delay through the dual Critic network. After the state space and action space of the intelligent agent in the battery pack balancing system are determined, the parameters of the Critic network and Actor network are initialized.

4. The battery pack equalization method based on reinforcement learning according to claim 1, characterized in that, In step 4, the reward function of the battery balancing system is: Where C is a constant, the function J is expressed as: The coefficients of function J are K = [k1 k2… k n ] 1×N The specific value of K is obtained by the following formula: when At that time, k i = a1; when At that time, k n = a1; when At that time, k i = a1×a2; when At that time, k n = a1×a2, where a1≥1, a2>1; In the above relationship, i = 1, 2, …, n-1; SOC i The current state of charge (SOC) of the battery. i_init The initial value of the battery's SOC is used to determine whether the battery pack has experienced over-balancing during the balancing process. a1 and a2 are penalty coefficients. a1 is used to penalize inconsistencies to quickly improve consistency, and a2 is used to penalize balancing processes that have experienced over-balancing.

Citation Information

Patent Citations

  • Electric quantity self-adaptive optimization balance control method of storage battery

    CN110303945A

  • Two-layer MPC method for improved module-based CPC equalization system

    CN113507148A