Training method and system for reconfigurable battery pack balancing model based on reinforcement learning
By adopting a reinforcement learning-based intelligent control model in the reconfigurable battery pack, the switching actions in the battery pack string are optimized, and the service life reduction and voltage fluctuations caused by frequent switching actions are solved, and the battery pack is more efficient, stable and safe equalization control is achieved.
Patent Information
- Application Number
- CN202111171899.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-08
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2041-10-08
AI Technical Summary
When the existing reconfigurable battery packs achieve battery power balance, frequent switching operations lead to a decrease in the service life of the control switch and fluctuate the bus voltage, affecting the performance, stability and safety of the battery pack.
Using a reconstructible battery pack equalization model based on reinforcement learning, multiple agents control the switching actions of the switching devices in the battery pack string separately, and optimize the parameter training of the agent to reduce the number of switch actions according to the state of charge deviation of the battery cell and the number of switch actions.
On the basis of ensuring the balance of the state of charge of the battery cell, it effectively reduces the number of switching operations of the switching device, extends the service life of the switching device, reduces the impact on the bus voltage, and improves the performance, stability and safety of the battery pack.
Smart Images

Figure CN113964897B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of battery control technology, and specifically provides a training method and system for a reconfigurable battery pack balancing model based on reinforcement learning. Background Art
[0002] In order to meet the requirements of the working voltage, output power and usage capacity of the energy storage system, a large number of battery cells are usually used in series and parallel to form a battery pack for use. However, during use, due to the inconsistency of the manufacturing process of the battery cells, there will be slight differences in parameters between different battery cells. After long-term cyclic use, the battery cells connected in series and parallel will age to varying degrees, resulting in increased differences in parameters such as the internal resistance and state of charge of the battery cells, causing the battery pack to have problems with insufficient state of charge or overcharge / over-discharge of the battery cells, which will affect the state of charge and cycle life of the battery pack during use, and may also cause serious safety problems.
[0003] The reconfigurable battery pack adds control switches around the battery cells, which can modify the connection mode of the battery cells in real time according to the status of the battery. The reconfigurable battery pack is also equipped with a certain number of redundant batteries for hot standby, which can effectively reduce the impact of battery cell exhaustion or damage on the battery pack. However, such a control method requires frequent switching actions to achieve the battery cell power balance effect, and frequent switching actions will inevitably reduce the service life of the control switch, and also cause fluctuations in the bus voltage, affecting the bus voltage.
[0004] Accordingly, the art needs a new reconfigurable battery pack balancing control solution to solve the above problems. Summary of the invention
[0005] The present invention aims to solve the above technical problem, that is, how to avoid frequent switching actions while ensuring the power balancing effect of battery cells in a reconfigurable battery pack.
[0006] In a first aspect, the present invention provides a training method for a reconfigurable battery pack balancing model based on reinforcement learning, wherein the reconfigurable battery pack balancing model includes a plurality of intelligent agents, the reconfigurable battery pack includes a plurality of battery strings, the plurality of battery strings are connected in parallel, each battery string includes a plurality of battery cell branches connected in series, each of the battery cell branches includes a battery cell and two switching devices, the battery cell is connected in series with a switching device to form a series branch, and the series branch is connected in parallel with another switching device; each of the intelligent agents is respectively configured to control the switching action of the switching device in the corresponding battery string according to the charge state of the battery cells of different battery strings in the reconfigurable battery pack, so as to control the access state of the battery cell, realize the access of the battery cell to the battery string or remove the battery cell from the battery string;
[0007] The training method comprises:
[0008] In each iterative training, for each battery string, the charge state of all battery cells in the battery string is obtained, and the access state of the battery cells in the previous iterative training and the current iterative training is obtained;
[0009] Acquire a state of charge deviation of a battery cell of a corresponding battery string according to the state of charge;
[0010] According to the access status of the battery cell in the previous iteration training and the current iteration training, the number of switching actions of the switching device in the corresponding battery string is obtained;
[0011] According to the state of charge deviation and the number of switch actions, an actual reward value output by the intelligent agent corresponding to the battery string is obtained;
[0012] The loss value of the agent is calculated according to the actual reward value and the target reward value, the gradient of the agent is calculated according to the loss value, and the parameters of the agent are updated according to the gradient back propagation to realize the training of the agent.
[0013] In a technical solution of the above training method, the step of "obtaining the actual reward value of the corresponding battery string according to the battery cell capacity deviation and the number of switch actions" includes obtaining the actual reward value by the following formula:
[0014] R=α1F1+α2F2
[0015] Wherein, R is the actual reward value, α1 and α2 are preset reward coefficients, F1 is the charge state deviation of the battery cell, and F2 is the number of switching actions of the battery string.
[0016] In a technical solution of the above training method, the step of "obtaining the state of charge deviation of the battery cell of the corresponding battery string according to the state of charge" includes obtaining the state of charge deviation of the battery cell by the following formula:
[0017]
[0018] Among them, SOC max The state of charge of the battery cell with the largest state of charge in the battery string, SOC min SOC is the state of charge of the battery cell with the smallest state of charge in the battery string. mean The SOC is the average value of the state of charge of all battery cells in the battery string. end It is the state of charge of the battery cell when charging or discharging is completed.
[0019] In a technical solution of the above training method, the step of "obtaining the number of switching actions of the switching device in the corresponding battery string according to the access status of the battery cell in the previous iterative training and the current iterative training" includes obtaining the number of switching actions of the switching device in the battery string by the following formula:
[0020]
[0021] Wherein, N is the number of battery cells in the battery string, is the access state of battery cell i in the t-th iteration training. When battery cell i is connected to the battery string, The value of is 1. When the battery cell i is removed from the battery string, The value of is 0; is the access status of battery cell i in the t-1th iteration training.
[0022] In a technical solution of the above training method, the method further comprises, in each iterative training, for each battery string, determining the access status of the battery cells in the battery string in the next iterative training according to the following steps:
[0023] The control agent adopts a greedy algorithm probability and determines the access state of the battery cell in the next iterative training according to the actual reward value, so that the number of switching actions of the switching device can be determined according to the access state in the next iterative training to iteratively train the agent;
[0024] The switching action includes connecting a battery cell to a battery string or removing a battery cell from a battery string.
[0025] In a second aspect, the present invention provides a training system for a reconfigurable battery pack balancing model based on reinforcement learning, wherein the reconfigurable battery pack balancing model includes a plurality of intelligent agents, the reconfigurable battery pack includes a plurality of battery strings, the plurality of battery strings are connected in parallel, each battery string includes a plurality of battery cell branches connected in series, each of the battery cell branches includes a battery cell and two switching devices, the battery cell is connected in series with a switching device to form a series branch, and the series branch is connected in parallel with another switching device; each of the intelligent agents is respectively configured to control the switching action of the switching device in the corresponding battery string according to the charge state of the battery cells of different battery strings in the reconfigurable battery pack, so as to control the access state of the battery cell, realize the access of the battery cell to the battery string or remove the battery cell from the battery string;
[0026] The training system comprises:
[0027] A battery cell access status acquisition module, which is configured to acquire the charge status of all battery cells in each battery string in each iterative training, and acquire the access status of the battery cells in the previous iterative training and the current iterative training;
[0028] A state of charge deviation acquisition module, configured to acquire a state of charge deviation of a battery cell of a corresponding battery string according to the state of charge;
[0029] A switch action number acquisition module, configured to acquire the number of switch actions of the switch device in the corresponding battery string according to the access status of the battery cell in the previous iteration training and the current iteration training;
[0030] an actual reward value acquisition module, configured to acquire an actual reward value output by the intelligent agent corresponding to the battery string according to the state of charge deviation and the number of switch actions;
[0031] An agent training module is configured to calculate the loss value of the agent according to the actual reward value and the target reward value, calculate the gradient of the agent according to the loss value, and update the parameters of the agent by back-propagation according to the gradient to achieve training of the agent.
[0032] In a technical solution of the above training system, the actual reward value acquisition module is further configured to obtain the actual reward value through the following formula:
[0033] R=α1F1+α2F2
[0034] Wherein, R is the actual reward value, α1 and α2 are preset reward coefficients, F1 is the charge state deviation of the battery cell, and F2 is the number of switching actions of the battery string.
[0035] In a technical solution of the above training system, the state of charge deviation acquisition module is further configured to acquire the state of charge deviation of the battery cell by the following formula:
[0036]
[0037] Among them, SOC max The state of charge of the battery cell with the largest state of charge in the battery string, SOC min SOC is the state of charge of the battery cell with the smallest state of charge in the battery string. mean The SOC is the average value of the state of charge of all battery cells in the battery string. end It is the state of charge of the battery cell when charging or discharging is completed.
[0038] In a technical solution of the above training system, the switch action number acquisition module is further configured to acquire the number of switch actions of the switch device in the battery string by the following formula:
[0039]
[0040] Wherein, N is the number of battery cells in the battery string, is the access state of battery cell i in the t-th iteration training. When battery cell i is connected to the battery string, The value of is 1. When the battery cell i is removed from the battery string, The value of is 0; is the access status of battery cell i in the t-1th iteration training.
[0041] In a technical solution of the above training system, the training system further includes a greedy algorithm control module, which is configured to determine, in each iterative training, for each battery string, the access status of the battery cells in the battery string in the next iterative training according to the following steps:
[0042] The control agent adopts a greedy algorithm probability and determines the access state of the battery cell in the next iterative training according to the actual reward value, so that the number of switching actions of the switching device can be determined according to the access state in the next iterative training to iteratively train the agent;
[0043] The switching action includes connecting a battery cell to a battery string or removing a battery cell from a battery string.
[0044] When adopting the above technical solution, the present invention can obtain the actual reward value output by the intelligent agent that controls the access state of the battery cell according to the charge state deviation of the battery cell in each battery string and the number of switching actions of the switching device, calculate the loss value of the intelligent agent according to the actual reward value and the target reward value, further calculate the gradient of the intelligent agent according to the loss value, and update the parameters of the intelligent agent according to the gradient back propagation to realize the training of the intelligent agent. Through the above configuration, the present invention can effectively train the intelligent agent controlling the battery string according to the charge state deviation of the battery cells in each battery string and the number of switching actions of the switching device, so that the actual reward value output by the intelligent agent controlling each battery string continues to increase, so that when the intelligent agent controls the access status of the battery cells in the battery string, it can minimize the number of switching actions of the switching device while ensuring that the charge state between the battery cells in the reconfigurable battery pack meets the balance requirement. In other words, the intelligent agent after effective training can effectively reduce the number of switching actions of the switching device while ensuring the balance of the charge state in the reconfigurable battery pack, thereby improving the service life of the switching device, while also reducing the fluctuation of the bus voltage caused by the switching action of the switching device, reducing the impact on the bus voltage, and improving the performance, stability and safety of the reconfigurable battery pack. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] The disclosure of the present invention will become more easily understood with reference to the accompanying drawings. It is easy for those skilled in the art to understand that these drawings are only for illustrative purposes and are not intended to limit the scope of protection of the present invention. Among them:
[0046] Figure 1 is a flowchart of main steps of a training method for a reconfigurable battery pack balancing model based on reinforcement learning according to an embodiment of the present invention;
[0047] Figure 2 is a schematic diagram of a reconfigurable battery pack according to an embodiment of the present invention;
[0048] Figure 3 is a schematic diagram of an equivalent model of battery packs connected in series and parallel according to an implementation of an embodiment of the present invention;
[0049] Figure 4 is a schematic diagram of a framework for training an intelligent agent according to an implementation of an embodiment of the present invention;
[0050] Figure 5 It is a main structural block diagram of a training system for a reconfigurable battery pack balancing model based on reinforcement learning according to an embodiment of the present invention. DETAILED DESCRIPTION
[0051] Some embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the protection scope of the present invention.
[0052] In the description of the present invention, "module" and "processor" may include hardware, software or a combination of the two. A module may include hardware circuits, various suitable sensors, communication ports, and memory, and may also include software parts, such as program code, or a combination of software and hardware. The processor may be a central processing unit, a microprocessor, a digital signal processor, or any other suitable processor. The processor has data and / or signal processing functions. The processor may be implemented in software, hardware, or a combination of the two. Non-temporary computer-readable storage media include any suitable medium that can store program code, such as a disk, a hard disk, an optical disk, a flash memory, a read-only memory, a random access memory, and the like. The term "A and / or B" means all possible combinations of A and B, such as only A, only B, or A and B. The term "at least one A or B" or "at least one of A and B" has a similar meaning to "A and / or B", and may include only A, only B, or A and B. The singular terms "one" and "the" may also include plural forms.
[0053] The terms involved in the present invention are explained below.
[0054] Switching devices refer to electronic components that can disconnect a circuit, interrupt current, connect a circuit, or allow current to flow to other circuits. Switching devices include power electronic control switches and relays. Power electronic control switches refer to switches built with power electronic devices, which can be fully controlled power semiconductor devices such as metal-oxide-semiconductor field-effect transistors (MOSFET), insulated gate bipolar transistors (IGBT), or integrated gate commutated thyristors (IGCT). A relay is an automatic switching device with an isolation function.
[0055] First, refer to the attached Figure 2 , Figure 2 FIG. 1 is a schematic diagram of a reconfigurable battery pack according to an embodiment of the present invention. Figure 2As shown, the reconfigurable battery pack balancing model may include multiple intelligent agents, the reconfigurable battery pack may include multiple battery strings, the multiple battery strings are connected in parallel, each battery string may include multiple battery cell branches connected in series, and each battery cell branch may include a battery cell B ij (where i = 1, 2, ..., m; j = 1, 2, ..., n; n represents the number of battery cells included in the battery string, and m represents the number of battery strings included in the reconfigurable battery pack) and two switching devices, a battery cell and a switching device S ij The series connection forms a series branch, which is connected to another switch device S - ij By controlling the switching device S ij and S - ij Realize battery cell B ij Connect to the battery string or connect the battery cell B ij Remove the battery cell B from the battery string. ij When connecting to a battery string, close S ij And disconnect S - ij ; Battery cell B needs to be ij When disconnecting from the battery string, disconnect S ij And close S - ij The reconfigurable battery pack can realize redundant configuration of battery cells. Specifically, if the number of redundant battery cells in a battery string is k, then the data of the battery cells actually working in the battery string is nk. The battery string can use redundant battery cells to achieve charge state balance of the battery string. When the reconfigurable battery pack is charging, if it is detected that the state of charge of some battery cells is large, while the state of charge of other battery cells is small, the battery cells with large state of charge can be removed from the battery string, and the redundant battery cells in the battery string can be connected to ensure that the number of battery cells connected in series in the battery string remains nk; similarly, when the reconfigurable battery pack is discharging, if it is detected that the state of charge of some battery cells is small, while the state of charge of other battery cells is large, the battery cells with small state of charge can be removed from the battery string, and the redundant battery cells in the battery string can be connected to ensure that the number of battery cells connected in series in the battery string remains nk; if a large number of battery cells in a battery string have low state of charge or are damaged, all battery cells in the battery string can be removed, and other battery strings can be used to power the load.
[0056] See attached Figure 3 , Figure 3 Schematic diagram of an equivalent model of battery packs connected in series and parallel according to an embodiment of the present invention. Figure 3 As shown, the battery string composed of battery cells connected in series is simplified to a series structure with an open circuit voltage E and an equivalent internal resistance R. Then the voltage V across the battery string can be calculated by formula (1):
[0057] V=E+I×R (1)
[0058] Wherein, I is the branch current flowing through the battery string.
[0059] Since the voltages at both ends of the two battery strings connected in parallel are the same, the equivalent open circuit voltages, equivalent internal resistances and branch currents of the two battery strings are related by the formula (2):
[0060] E1+I1×R1=E2+I2×R2 (2)
[0061] Among them, E1 and E2 are the equivalent open-circuit voltages of the two battery strings, R1 and R2 are the equivalent internal resistances of the two battery strings, and I1 and I2 are the currents flowing through the two battery string branches.
[0062] According to formula (2), the DC currents I1 and I2 flowing through the two battery strings are calculated as shown in formulas (3) and (4) respectively.
[0063]
[0064]
[0065] Among them, I=I1+I2.
[0066] Further, according to formula (3) and formula (4), formula (5) can be obtained.
[0067]
[0068] It can be seen from formula (5) that if E1 and E2 are not equal, R1 and R2 are not equal, then the current flowing through the two battery string branches will be different, and the difference in current will affect the charge state of the battery cell. Taking the charging process as an example, the charge state of the battery cell is defined as shown in formula (6).
[0069]
[0070] Among them, SOC(t) is the state of charge of the battery cell at time t, SOC(t0) is the state of charge of the battery cell at time t0, i(τ) is the charging function of the battery cell that changes with time, It is the cumulative amount of electricity charged into the battery cells from t0 to t.
[0071] If the effect of temperature on the battery cells is not considered, the open circuit voltage E of the battery string can be equivalent to the state of charge function of the battery cells as shown in formula (7).
[0072]
[0073] Among them, a i and b are coefficients, n is the order of the polynomial, n is the number of battery cells in the battery string, j represents the jth battery cell in the battery string, a j is the coefficient of the jth battery cell.
[0074] It can be seen from the above formula that if the initial state of charge of the parallel battery strings is the same, then E1 = E2. If R2>R1, it will lead to I1>I2. When the initial state of charge of the parallel battery strings is the same, since I1>I2, the rate of increase of the state of charge of battery string 1 will be higher than that of battery string 2, that is, E1>E2 will be caused, so that the difference between I1 and I2 will gradually decrease. However, even if the difference between I1 and I2 can be reduced to 0 during the charging process, the difference in the cumulative state of charge in the early stage already exists. Only at the end of the charge and discharge period can I2>I1 be achieved. In the actual use of the reconfigurable battery pack, the battery cells in the battery string are generally not fully charged or fully exhausted. Therefore, the battery strings used in parallel will not be able to achieve the balance of the state of charge between the battery strings due to the difference in internal resistance between the battery strings, which makes the battery strings work for a long time in different states of charge, and thus leads to different degrees of aging of the battery cells in the battery string.
[0075] The multi-agent double deep Q-network (DDQN) reinforcement learning algorithm can achieve joint balancing of battery strings in reconfigurable battery packs. DDQN is a reinforcement learning algorithm based on Q learning, which is an improvement of DQN. It contains two neural networks, online network and target network. In the process of deep learning using the DDQN method, small batches of data can be uniformly extracted from the stored experience data to update the training parameters of DDQN for training. The agent is an entity that performs control operations on the control object through environmental state parameters. Multi-agent reinforcement learning uses multiple agents to control multiple targets. Each agent receives all or part of the environment state of the agent (the access state and charge state deviation of the battery cell in the current iterative training) to exert different effects on the environment in which the agent is located (the access state of the battery cell in the next iterative training). Through the synergy between multiple agents, more complex control tasks can be achieved.
[0076] In this embodiment, each battery string is respectively communicated with an intelligent agent in the reconfigurable battery pack balancing model, and the intelligent agent is configured to control the switching action of the switching device in the battery string according to the charge state of the battery cell in the reconfigurable battery pack, so as to control the access state of the battery cell, realize the access of the battery cell to the battery string or cut off the battery cell from the battery string.
[0077] Specifically, if the number of battery cells in a battery string in a reconfigurable battery pack is n and the number of redundant battery cells is k, then the number of actions for controlling the battery string (different combinations of battery cells that can be connected to the battery string can be determined by switching the switching devices, that is, the number of available battery cell combinations) is as shown in formula (8).
[0078]
[0079] Among them A number is the number of actions that the agent performs on a single battery string. For example, if a battery string includes three battery cells: battery cell a, battery cell b, and battery cell c, and the number of redundant battery cells k in the battery string is 1, then by switching the switch devices in the battery string, different combinations of battery cells that can be connected to the battery string can be determined, including combinations 1-3, where combination 1 includes battery cell a and battery cell b, combination 2 includes battery cell b and battery cell c, and combination 3 includes battery cell c and battery cell a.
[0080] As the total number of battery cells and the number of redundant battery cells increase, the task of achieving charge state balance between battery cells becomes more complex, and the number of actions to control the battery string also increases. For the case where the battery strings from the 1st to the i-th string are connected in parallel, the number of actions to control the battery cells is shown in formula (9).
[0081]
[0082] Among them, A number总 is the number of actions taken by the agent to control the reconfigurable battery pack, and i is the number of battery strings connected in parallel.
[0083] Since the number of actions increases exponentially, but the environmental state of the reconfigurable battery pack is relatively small, the input information (environmental state of the reconfigurable battery pack) in the DDQN network is much smaller than the output information (the number of actions of the agent), resulting in the DDQN network being unable to achieve effective learning and updating. Therefore, in this embodiment, an agent is set up for each battery string to control it, which can effectively reduce the number of actions of a single agent to control the reconfigurable battery pack, so that when the environmental state of the reconfigurable battery pack is the same, the number of actions of a single agent to control the reconfigurable battery pack is reduced, and the agent can effectively learn and update according to the environmental state of the reconfigurable battery.
[0084] See attached Figure 1 , Figure 1 FIG. 1 is a flow chart of the main steps of a training method for a reconfigurable battery pack balancing model based on reinforcement learning according to an embodiment of the present invention. Figure 1 As shown, the training method of the reconfigurable battery pack balancing model based on reinforcement learning in the embodiment of the present invention mainly includes the following steps S101 to S105.
[0085] Step S101: in each iterative training, for each battery string, the charge state of all battery cells in the battery string is obtained, and the access state of the battery cells in the previous iterative training and the current iterative training is obtained.
[0086] In this embodiment, in each iteration training of the agent, for each battery string, the charge state of multiple battery cells in the battery string can be obtained, and the access state of the battery cell in the previous iteration training and the current iteration training can be obtained. The access state refers to the state of the battery cell being connected to the battery string or the state of the battery cell being removed from the battery string.
[0087] Step S102: acquiring the state of charge deviation of the battery cells of the corresponding battery string according to the state of charge.
[0088] In this embodiment, the state of charge deviation of the battery cells in the battery string can be obtained according to the state of charge of multiple battery cells in the battery string. The state of charge deviation refers to the degree of difference between the states of charge of the battery cells in the battery string.
[0089] Step S103: according to the access status of the battery cell in the previous iteration training and the current iteration training, the number of switching actions of the switching device in the corresponding battery string is obtained.
[0090] In this embodiment, the number of switching actions of the switch device in the corresponding battery string can be obtained according to the access state of the battery cell in the previous iterative training and the current iterative training. That is, in the previous iterative training and the current iterative training, if the access state of the battery cell is the same, the number of switching actions of the switch device controlling the battery cell remains unchanged; if the access state of the battery cell is different in the previous iterative training and the current iterative training, the number of switching actions of the switch device controlling the battery cell is increased by 1.
[0091] Step S104: According to the state of charge deviation and the number of switch actions, the actual reward value output by the intelligent agent corresponding to the battery string is obtained.
[0092] In this embodiment, the intelligent agent controls the switch device to connect the battery cell to the battery string or remove the battery cell from the battery string. By connecting the battery cell to the battery string or removing the battery cell from the battery string, the charge state of the battery cell can be balanced. However, if the closing and opening of the switch device is frequently controlled, the service life of the switch device will be affected, and the stability of the bus voltage will also be affected. Therefore, in this embodiment, the actual reward value output by the intelligent agent corresponding to the battery string can be obtained according to the charge state deviation of the battery cell in the battery string and the number of switching actions of the switch device, that is, the influence of the charge state deviation and the number of switching actions on the actual reward value output by the intelligent agent is comprehensively considered.
[0093] Step S105: Calculate the loss value of the agent according to the actual reward value and the target reward value, calculate the gradient of the agent according to the loss value, and update the parameters of the agent according to the gradient back propagation to achieve training of the agent.
[0094] In this embodiment, the loss value of the agent can be calculated according to the actual reward value obtained in step S104 and the preset target reward value, and the gradient of the agent can be calculated according to the loss value, and the parameters of the agent can be updated according to the gradient back propagation to achieve the training of the agent. Specifically, according to the back propagation mechanism of the DDQN network, based on the gradient descent method, the parameters of the agent can be updated according to the loss value of the actual reward value and the target reward value, and the training of the agent can be achieved through cyclic iterative training.
[0095] Based on the above steps S101 to S105, the present invention can obtain the actual reward value output by the intelligent agent that controls the access state of the battery cell according to the charge state deviation of the battery cell in each battery string and the number of switching actions of the switching device, calculate the loss value of the intelligent agent according to the actual reward value and the target reward value, further calculate the gradient of the intelligent agent according to the loss value, and update the parameters of the intelligent agent according to the gradient back propagation to realize the training of the intelligent agent. Through the above configuration, the present invention can effectively train the intelligent agent controlling the battery string according to the charge state deviation of the battery cells in each battery string and the number of switching actions of the switching device, so that the actual reward value output by the intelligent agent controlling each battery string continues to increase, so that when the intelligent agent controls the access status of the battery cells in the battery string, it can minimize the number of switching actions of the switching device while ensuring that the charge state between the battery cells in the reconfigurable battery pack meets the balance requirement. In other words, the intelligent agent after effective training can effectively reduce the number of switching actions of the switching device while ensuring the balance of the charge state in the reconfigurable battery pack, thereby improving the service life of the switching device, while also reducing the fluctuation of the bus voltage caused by the switching action of the switching device, reducing the impact on the bus voltage, and improving the performance, stability and safety of the reconfigurable battery pack.
[0096] Step S102, step S103 and step S104 are further described below.
[0097] In one implementation of the embodiment of the present invention, step S102 may include obtaining the state of charge deviation of the battery cell by using the following formula (10).
[0098]
[0099] Among them, F1 is the charge state deviation of the battery cell, SOC max SOC is the state of charge of the battery cell with the largest state of charge in the battery string. min SOC is the state of charge of the battery cell with the smallest state of charge in the battery string. mean The SOC is the average value of the state of charge of all battery cells in the battery string. end It is the state of charge of a battery cell at the end of charging or discharging.
[0100] In this embodiment, in order to achieve the charge state balance of each battery cell, the battery cells in the reconfigurable battery pack cannot be charged or discharged at the same time. Ideally, the balance between the battery cells should show a trend of gradually increasing during the charging or discharging process of the reconfigurable battery pack. Therefore, a parameter for calculating the actual reward value output by the intelligent agent is set to the charge state deviation of the battery cell, and the charge state deviation of the battery cell is calculated by the above formula (10). Among them, during the charging stage, SOC end Can be set to 100, discharge stage, SOC end It can be set to 0. This will make the charge state deviation of the battery cell smaller in the initial stage of charging or discharging, and increase continuously during the charging or discharging process.
[0101] In one implementation of the embodiment of the present invention, step S103 may include obtaining the number of switching actions of the switching device in the battery string by the following formula (11):
[0102]
[0103] Wherein, F2 is the number of switching actions of the switching device in the battery string, N is the number of battery cells in the battery string, is the access status of battery cell i in the tth iteration training. When battery cell i is connected to the battery string, The value of is 1. When the battery cell i is removed from the battery string, The value of is 0; is the access status of battery cell i in the t-1th iteration training.
[0104] In this embodiment, the charge state balance between battery cells is achieved by controlling the switch device to control the battery cells to connect to the battery string or remove the battery cells from the battery string. However, if the switch device is frequently switched, it will affect the service life of the switch device and the stability of the bus voltage. Therefore, the number of switching actions of the switch device can be minimized when the degree of charge state imbalance of the battery cells is acceptable. The number of switching actions of the switch device can be calculated according to the above formula (11). In two adjacent iterative trainings, if the switch device has a switching action, then The value of is 1; if the switching device does not switch, then The value of is 0.
[0105] In one implementation of the present invention, step S104 may include obtaining the actual reward value output by the agent through the following formula (12):
[0106] R=α1F1+α2F2 (12)
[0107] Among them, R is the actual reward value, and α1 and α2 are the preset reward coefficients.
[0108] In this embodiment, a reward system can be set for the charge state deviation of the battery cell and the number of switching actions of the switching device respectively to obtain a formula for calculating the actual reward value output by the intelligent agent, as shown in formula (12).
[0109] In one embodiment, see the attached Figure 4 , Figure 4 FIG. 1 is a schematic diagram of a framework for training an agent according to an implementation of an embodiment of the present invention. Figure 4 As shown, when training the agent, the agent can be trained based on factors such as the environment (the access status of the battery cells in the previous iteration training of the reconfigurable battery pack), the state (the charge state of the battery cells in each battery string in the previous iteration training), the reward (the actual reward value obtained according to formula (12)), and the action (the access status of the battery cells in each battery string in the current iteration training). Multiple agents can be set, each of which controls a battery string. In the current iteration training, after the action is executed, the state is updated accordingly; the agent receives the new state and actual reward value, and calculates the loss value of the agent based on the actual reward value and the target reward value, calculates the gradient of the agent based on the loss value, and applies the back propagation mechanism of the DDQN network to update the parameters of the agent based on the gradient of the agent, that is, to update the two reward coefficients (α1 and α2) in formula (12) to achieve the purpose of training the agent. Figure 4 In the above example, S refers to the current iteration training, S -1 Refers to the previous training iteration.
[0110] In one implementation, a target actual reward value may be set, and when the actual reward value output by the agent reaches the target actual reward value, the training of the agent may be terminated. A threshold for the number of iterative trainings may also be set, and when the number of iterative trainings of the agent reaches the threshold, the training of the agent may be terminated. Those skilled in the art may select the values of the target actual reward value and the threshold for the number of iterative trainings according to the needs of the actual application.
[0111] In one implementation of the embodiment of the present invention, in addition to the above steps S101 to S105, the present invention may further include the following step S106, according to which in each iteration training, for each battery string, the access status of the battery cells in the battery string in the next iteration training is confirmed;
[0112] Step S106: The control agent adopts the greedy algorithm probability and determines the access status of the battery cell in the next iterative training according to the actual reward value, so that the switching action of the switching device can be determined according to the access status to iteratively train the agent in the next iterative training; wherein the switching action includes connecting the battery cell to the battery string or removing the battery cell from the battery string.
[0113] In this embodiment, in order to enable the agent to fully explore unknown situations and enable the training process of the agent to obtain better training results, an ε-greedy algorithm can be used to set the probability of the greedy algorithm to add a certain degree of randomness to the control process of the agent, that is, in the next iterative training, a random switch action is selected with probability ε, and a switch action with the maximum reward is selected with probability 1-ε. Among them, the ε-greedy algorithm refers to an algorithm that selects a random switch action with probability ε and selects the optimal switch action with probability 1-ε.
[0114] In one embodiment, the switching device may be a power electronically controlled switch.
[0115] In another embodiment, the switching device may be a relay.
[0116] In the embodiment of the present invention, after the reconfigurable battery pack balancing model based on reinforcement learning is trained through the above steps S101 to S106, the trained reconfigurable battery pack balancing model can be directly used to Figure 2 The reconfigurable battery pack shown in the figure performs charge state balancing control. Specifically, each intelligent agent in the reconfigurable battery pack balancing model is controlled to obtain the charge state deviation of the battery cell of a battery string and the number of switching actions of the switching device, and the switching action of the switching device in the battery string is determined according to the charge state deviation and the number of switching actions to change the access state of the battery cell in the battery string, so as to realize the charge state balancing control of the reconfigurable battery pack.
[0117] It should be pointed out that although the various steps in the above embodiments are described in a specific order, those skilled in the art can understand that in order to achieve the effects of the present invention, different steps do not have to be performed in such an order. They can be performed simultaneously (in parallel) or in other orders. These changes are within the scope of protection of the present invention.
[0118] Furthermore, the present invention also provides a training system for a reconfigurable battery pack balancing model based on reinforcement learning.
[0119] See attached Figure 5 , Figure 5 FIG. 1 is a main structural block diagram of a training system for a reconfigurable battery pack balancing model based on reinforcement learning according to an embodiment of the present invention. Figure 5As shown, the reconfigurable battery pack balancing model in the embodiment of the present invention may include multiple agents, the reconfigurable battery pack may include multiple battery strings, the multiple battery strings are connected in parallel, each battery string includes multiple battery cell branches connected in series, each battery cell branch includes a battery cell and two switch devices, the battery cell is connected in series with one switch device to form a series branch, and the series branch is connected in parallel with another switch device; each agent can be configured to control the switching action of the switch device in the corresponding battery string according to the state of charge of the battery cells of different battery strings in the reconfigurable battery pack, so as to control the access state of the battery cell, realize the access of the battery cell to the battery string or remove the battery cell from the battery string. The training system may include a battery cell access state acquisition module, a state of charge deviation acquisition module, a switch action number acquisition module, an actual reward value acquisition module and an agent training module. In this embodiment, the battery cell access state acquisition module can be configured to obtain the state of charge of all battery cells in the battery string for each battery string in each iterative training, and obtain the access state of the battery cell in the previous iterative training and the current iterative training. The state of charge deviation acquisition module can be configured to acquire the state of charge deviation of the battery cell of the corresponding battery string according to the state of charge. The switch action number acquisition module can be configured to acquire the number of switch actions of the switch device in the corresponding battery string according to the access status of the battery cell in the previous iteration training and the current iteration training. The actual reward value acquisition module can be configured to acquire the actual reward value output by the intelligent agent corresponding to the battery string according to the state of charge deviation and the number of switch actions. The intelligent agent training module can be configured to calculate the loss value of the intelligent agent according to the actual reward value and the target reward value, calculate the gradient of the intelligent agent according to the loss value, and update the parameters of the intelligent agent according to the gradient back propagation to realize the training of the intelligent agent.
[0120] In one embodiment, the actual reward value acquisition module may be further configured to acquire the actual reward value by the following formula:
[0121] R=α1F1+α2F2
[0122] Among them, R is the actual reward value, α1 and α2 are preset reward coefficients, F1 is the charge state deviation of the battery cell, and F2 is the number of switching actions of the battery string.
[0123] In one embodiment, the state of charge deviation acquisition module may be further configured to acquire the battery cell capacity deviation by the following formula:
[0124]
[0125] Among them, SOC max SOC is the state of charge of the battery cell with the largest state of charge in the battery string.min SOC is the state of charge of the battery cell with the smallest state of charge in the battery string. mean The SOC is the average value of the state of charge of all battery cells in the battery string. end It is the state of charge of a battery cell at the end of charging or discharging.
[0126] In one embodiment, the switch action times acquisition module may be further configured to acquire the number of switch actions of the switch device in the battery string by using the following formula:
[0127]
[0128] Where N is the number of battery cells in the battery string, is the access status of battery cell i in the tth iteration training. When battery cell i is connected to the battery string, The value of is 1. When the battery cell i is removed from the battery string, The value of is 0; is the access status of battery cell i in the t-1th iteration training.
[0129] In one embodiment, the training system may further include a greedy algorithm control module, which may be configured to determine, for each battery string in each iterative training, the access status of the battery cells in the battery string in the next iterative training according to the following steps:
[0130] The control agent adopts the greedy algorithm probability and determines the access state of the battery cell in the next iterative training according to the actual reward value, so that the number of switching actions of the switching device can be determined according to the access state in the next iterative training to iteratively train the agent; wherein the switching action includes connecting the battery cell to the battery string or removing the battery cell from the battery string.
[0131] The training system of the above-mentioned reconfigurable battery pack balancing model based on reinforcement learning is used to execute Figure 1 The embodiment of the training method for the reconfigurable battery pack balancing model based on reinforcement learning shown in the figure has similar technical principles, technical problems solved and technical effects produced. Technicians in this technical field can clearly understand that for the convenience and conciseness of description, the specific working process and related instructions of the training system for the reconfigurable battery pack balancing model based on reinforcement learning can refer to the contents described in the embodiment of the training method for the reconfigurable battery pack balancing model based on reinforcement learning, which will not be repeated here.
[0132] It is understood by those skilled in the art that the present invention implements all or part of the processes in the method of the above embodiment, and can also be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of each of the above method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device, medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory, random access memory, electric carrier signal, telecommunication signal and software distribution medium that can carry the computer program code. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.
[0133] Further, it should be understood that since the setting of each module is only for illustrating the functional units of the device of the present invention, the physical devices corresponding to these modules may be the processor itself, or a part of the software in the processor, a part of the hardware, or a part of the combination of software and hardware. Therefore, the number of each module in the figure is only schematic.
[0134] Those skilled in the art will appreciate that the modules in the device can be adaptively split or merged. Such splitting or merging of specific modules will not cause the technical solution to deviate from the principle of the present invention, and therefore, the technical solutions after splitting or merging will fall within the protection scope of the present invention.
[0135] So far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.
Claims
1. A training method for a reconfigurable battery pack balancing model based on reinforcement learning, characterized in that: The reconfigurable battery pack balancing model includes a plurality of intelligent agents, the reconfigurable battery pack includes a plurality of battery strings, the plurality of battery strings are connected in parallel, each battery string includes a plurality of battery cell branches connected in series, each of the battery cell branches includes a battery cell and two switching devices, the battery cell is connected in series with one switching device to form a series branch, and the series branch is connected in parallel with another switching device; each of the intelligent agents is respectively configured to control the switching action of the switching device in the corresponding battery string according to the charge state of the battery cells of different battery strings in the reconfigurable battery pack, so as to control the access state of the battery cell, realize the access of the battery cell to the battery string or remove the battery cell from the battery string; The training method comprises: In each iterative training, for each battery string, the charge state of all battery cells in the battery string is obtained, and the access state of the battery cells in the previous iterative training and the current iterative training is obtained; Acquire a state of charge deviation of a battery cell of a corresponding battery string according to the state of charge; According to the access status of the battery cell in the previous iteration training and the current iteration training, the number of switching actions of the switching device in the corresponding battery string is obtained; According to the state of charge deviation and the number of switch actions, an actual reward value output by the intelligent agent corresponding to the battery string is obtained; Calculating the loss value of the agent according to the actual reward value and the target reward value, calculating the gradient of the agent according to the loss value, and updating the parameters of the agent according to the gradient back propagation to achieve training of the agent; The step of "obtaining the actual reward value output by the intelligent agent corresponding to the battery string according to the state of charge deviation and the number of switch actions" includes obtaining the actual reward value by the following formula: R=α1F1+α2F2 Wherein, R is the actual reward value, α1 and α2 are preset reward coefficients, F1 is the charge state deviation of the battery cell, and F2 is the number of switching actions of the battery string.
2. The training method according to claim 1, characterized in that: The step of “obtaining the state of charge deviation of the battery cell of the corresponding battery string according to the state of charge” includes obtaining the state of charge deviation of the battery cell by the following formula: Among them, SOC max The state of charge of the battery cell with the largest state of charge in the battery string, SOC min SOC is the state of charge of the battery cell with the smallest state of charge in the battery string. mean The SOC is the average value of the state of charge of all battery cells in the battery string. end It is the state of charge of the battery cell when charging or discharging is completed.
3. The training method according to claim 2, characterized in that: The step of "obtaining the number of switching actions of the switching device in the corresponding battery string according to the access state of the battery cell in the previous iteration training and the current iteration training" includes obtaining the number of switching actions of the switching device in the battery string by the following formula: Wherein, N is the number of battery cells in the battery string, is the access state of battery cell i in the t-th iteration training. When battery cell i is connected to the battery string, The value of is 1. When the battery cell i is removed from the battery string, The value of is 0; is the access status of battery cell i in the t-1th iteration training.
4. The training method according to claim 1, characterized in that: The method further includes determining, for each battery string in each iterative training, the access status of the battery cells in the battery string in the next iterative training according to the following steps: The control agent adopts a greedy algorithm probability and determines the access state of the battery cell in the next iterative training according to the actual reward value, so that the number of switching actions of the switching device can be determined according to the access state in the next iterative training to iteratively train the agent; The switching action includes connecting a battery cell to a battery string or removing a battery cell from a battery string.
5. A training system for a reconfigurable battery pack balancing model based on reinforcement learning, characterized in that: The reconfigurable battery pack balancing model includes a plurality of intelligent agents, the reconfigurable battery pack includes a plurality of battery strings, the plurality of battery strings are connected in parallel, each battery string includes a plurality of battery cell branches connected in series, each of the battery cell branches includes a battery cell and two switching devices, the battery cell is connected in series with one switching device to form a series branch, and the series branch is connected in parallel with another switching device; each of the intelligent agents is respectively configured to control the switching action of the switching device in the corresponding battery string according to the charge state of the battery cells of different battery strings in the reconfigurable battery pack, so as to control the access state of the battery cell, realize the access of the battery cell to the battery string or remove the battery cell from the battery string; The training system comprises: A battery cell access status acquisition module, which is configured to acquire the charge status of all battery cells in each battery string in each iterative training, and acquire the access status of the battery cells in the previous iterative training and the current iterative training; A state of charge deviation acquisition module, configured to acquire a state of charge deviation of a battery cell of a corresponding battery string according to the state of charge; A switch action number acquisition module, configured to acquire the number of switch actions of the switch device in the corresponding battery string according to the access status of the battery cell in the previous iteration training and the current iteration training; an actual reward value acquisition module, configured to acquire an actual reward value output by the intelligent agent corresponding to the battery string according to the state of charge deviation and the number of switch actions; An agent training module is configured to calculate a loss value of the agent according to the actual reward value and the target reward value, calculate a gradient of the agent according to the loss value, and update the parameters of the agent by back-propagation according to the gradient to implement training of the agent; The actual reward value acquisition module is further configured to acquire the actual reward value by the following formula: R=α1F1+α2F2 Wherein, R is the actual reward value, α1 and α2 are preset reward coefficients, F1 is the charge state deviation of the battery cell, and F2 is the number of switching actions of the battery string.
6. The training system according to claim 5, characterized in that The state of charge deviation acquisition module is further configured to acquire the battery cell capacity deviation by the following formula: Among them, SOC max The state of charge of the battery cell with the largest state of charge in the battery string, SOC min SOC is the state of charge of the battery cell with the smallest state of charge in the battery string. mean The SOC is the average value of the state of charge of all battery cells in the battery string. end It is the state of charge of the battery cell when charging or discharging is completed.
7. The training system according to claim 6, characterized in that The switching action times acquisition module is further configured to acquire the number of switching actions of the switching device in the battery string by using the following formula: Wherein, N is the number of battery cells in the battery string, is the access state of battery cell i in the t-th iteration training. When battery cell i is connected to the battery string, The value of is 1. When the battery cell i is removed from the battery string, The value of is 0; is the access status of battery cell i in the t-1th iteration training.
8. The training system according to claim 5, characterized in that The training system further includes a greedy algorithm control module, which is configured to determine, in each iterative training, for each battery string, the access state of the battery cells in the battery string in the next iterative training according to the following steps: The control agent adopts a greedy algorithm probability and determines the access state of the battery cell in the next iterative training according to the actual reward value, so that the number of switching actions of the switching device can be determined according to the access state in the next iterative training to iteratively train the agent; The switching action includes connecting a battery cell to a battery string or removing a battery cell from a battery string.
Citation Information
Patent Citations
Training method and device of battery pack equalization control model and medium
CN113987921A