Training method of SOH balance control model of reconfigurable battery system and balance control method

By constructing Markov decision model and DQN algorithm training neural network, the problem of SOH differences in battery cells in dynamic reconfigurable battery systems is solved, and efficient balance control and dynamic decision-making of battery systems are realized.

CN120373403APending Publication Date: 2025-07-25WUHAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510373994.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The prior art is difficult to effectively solve the SOH difference between battery cells in dynamic reconfigurable battery systems, resulting in reduced system energy utilization efficiency and reliability problems.

Method used

A SOH equalization control method based on Markov decision model is constructed, sample parameters are obtained through simulated battery system, and neural network model is trained using DQN algorithm to avoid frequent measurement or estimation of SOH, so as to achieve equalization control of the battery system.

Benefits of technology

The SOH balance control is simplified, the energy utilization efficiency and reliability of the system are improved, and simple balance control of SOH and real-time dynamic decision-making are achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373403A_ABST
    Figure CN120373403A_ABST
Patent Text Reader

Abstract

The invention provides a training method of a reconfigurable battery system SOH balance control model and a balance control method, and relates to the technical field of battery management. The training method of the SOH balance control model of the reconfigurable battery system comprises the following steps: constructing a Markov decision model based on SOH balance control of the reconfigurable battery system; acquiring sample parameters based on a Markov decision model and a simulation battery system, and storing the sample parameters in an experience playback area until the data quantity stored in the experience playback area reaches a preset threshold value; and taking data in the experience playback area as a training sample, and training a neural network model based on a DQN algorithm to obtain a target balance control model. Through the training method, the neural network model can be trained according to historical data, a control decision can be made, and the SOH equalization problem of the dynamic reconfigurable battery system is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of battery management, and in particular, to a training method and an equalization control method for a state of health (SOH) equalization control model of a reconfigurable battery system. Background Art

[0002] With the transformation of the energy structure and the increasing urgency of sustainable development, battery energy storage technology has become an important topic in the global energy field due to its characteristics such as fast response speed, flexible storage and scheduling capabilities, and environmental friendliness. However, due to manufacturing tolerances, environmental differences, operating condition differences, etc., during the operation of the system, the SOH differences between battery cells gradually emerge, resulting in a significant reduction in the energy utilization efficiency of the system and even posing a huge challenge to the reliability and safety of the system.

[0003] The dynamic reconfigurable battery system can flexibly adjust the battery configuration according to the actual operating conditions of the battery pack by real-time sensing the battery state and combining advanced control algorithms, so as to achieve optimal energy distribution and then achieve system state equalization. However, currently, the SOH equalization control schemes for reconfigurable battery systems are mainly rule-based and optimization methods, which require frequent measurement or estimation of SOH, and it is difficult to measure or estimate SOH, resulting in difficulty in effectively equalizing the control of the battery system. Summary of the Invention

[0004] In view of this, the present application proposes a training method and an equalization control method for a state of health (SOH) equalization control model of a reconfigurable battery system.

[0005] In a first aspect, the present application provides a training method for a state of health (SOH) equalization control model of a reconfigurable battery system. The reconfigurable battery system includes a plurality of battery cells and a plurality of switch units, and the operating state of the battery cell is related to the switch state of the corresponding switch unit. The method includes:

[0006] Constructing a Markov decision model based on the SOH equalization control of the reconfigurable battery system. The Markov decision model includes a state space, an action space, and a reward function. The state space includes a set of state parameters of the reconfigurable battery system. The state parameters of the reconfigurable battery system include a state of charge parameter, a relative state of charge parameter, and a switch state parameter. The relative state of charge parameter is respectively related to the state of charge parameter and the maximum discharge depth of the battery cell. The action space includes a set of switch state parameters. The reward function reflects the difference between the relative state of charge parameter variable and the average relative state of charge parameter during the discharge process of the battery.

[0007] Obtain sample parameters based on the Markov decision model and the simulated battery system, and save the sample parameters in the experience replay area until the number of data stored in the experience replay area reaches a preset threshold. The sample parameters include state parameters, action parameters, reward parameters, and predicted state parameters of the reconfigurable battery system;

[0008] Use the data in the experience replay area as training samples to train a neural network model based on the DQN algorithm to obtain a target balancing control model.

[0009] In one embodiment, before obtaining the sample parameters based on the Markov decision model and the simulated battery system, the method further includes:

[0010] Obtain the initial SOH values of each battery cell in the reconfigurable battery system and the reference discharge depth of the target battery cell, where the target battery cell is the battery cell with the highest initial SOH value;

[0011] Determine the discharge depths of the other battery cells according to the reference discharge depth and the initial SOH values of each battery cell;

[0012] Set the initial state-of-charge parameters of each battery cell to the same predetermined parameter, use the reference cycle number of the target battery cell as the expected cycle number of the reconfigurable battery system, and construct a simulated battery system based on the discharge depths of each battery cell.

[0013] In one embodiment, obtaining the sample parameters based on the Markov decision model and the simulated battery system and saving the sample parameters in the experience replay area includes:

[0014] Use the simulated battery system as the environment of the Markov decision model, obtain the state parameters and action parameters of the simulated battery system at each time step of the cycle, and solve the Markov decision model to obtain the reward parameters and predicted state parameters at each time step of the cycle;

[0015] Save the state parameters, action parameters, reward parameters, and predicted state parameters at the same time step in the experience replay area in the form of a quadruple storage structure.

[0016] In one embodiment, using the data in the experience replay area as training samples to train a neural network model based on the DQN algorithm to obtain a target balancing control model includes:

[0017] Obtain the sample data in the experience replay area;

[0018] Transfer the state parameters and action parameters in the sample data to the valuation neural network of the neural network model, predict the state parameters and transfer them to the target neural network of the neural network model, and combine the reward parameters and the target neural network to generate a training label value;

[0019] Update the weight parameters of the neural network model according to the training label value, the preset loss function, and the optimization function;

[0020] Repeat the above training steps until the number of cycle periods of the reconfigurable battery system reaches the expected number of cycle periods and the relative state of charge parameter reaches 0%.

[0021] In one embodiment, the calculation formula for the reference cycle number of the target battery cell is:

[0022]

[0023] Among them, C acu,min is the reference cycle number at the time of achieving balance, SOHt,retire is the SOH value at the retirement level, a and b are battery parameters related to the battery type, SOHmax,init is the maximum initial SOH of the battery cell, and DODlimit is the maximum discharge depth of the battery cell.

[0024] In one embodiment, determining the discharge depth of other battery cells according to the reference discharge depth and the initial SOH values of each battery cell includes:

[0025] Based on the reference discharge depth and the initial SOH values of each battery cell, determine the discharge depth of other battery cells based on a preset depth calculation formula;

[0026] Among them, the preset depth calculation formula is:

[0027]

[0028] Among them, DOD max,i is the maximum discharge depth of the target battery cell, DOD max,j is the maximum discharge depth of other battery cells to be determined, SOH i,init is the initial SOH value of the target battery cell, SOH j,init is the initial SOH value of other battery cells to be determined, C acu,i is the reference cycle number of the target battery cell, a and b are battery parameters related to the battery type, SOHi,t = SOHj,t at time t, and Cacu,i = Cacu,j.

[0029] In one embodiment, the correlation formula between the relative state of charge parameter and the state of charge parameter and the maximum discharge depth of the battery is as follows:

[0030]

[0031] Among them, RSOC is the relative state of charge parameter of the battery cell, SOC is the state of charge parameter of the battery cell, and DOD max is the maximum depth of discharge of the battery cell.

[0032] In a second aspect, the present application also provides a method for SOH balancing control of a reconfigurable battery system, and the method for SOH balancing control of the reconfigurable battery system includes:

[0033] Obtain the initial parameters of the reconfigurable battery system in the initial state;

[0034] Input the initial parameters into the target balancing control model to obtain the control strategy of the reconfigurable battery system in a single cycle period, where the target balancing control model is trained by the method described in the first aspect;

[0035] Control the working state of the reconfigurable battery system according to the control strategy.

[0036] In a third aspect, the present application also provides an electronic device, including a processor and a memory; the memory stores a computer program, and when the computer program is executed by the processor, it implements the training method of the SOH balancing control model of the reconfigurable battery system described in the first aspect or the SOH balancing control method of the reconfigurable battery system described in the second aspect.

[0037] In a fourth aspect, the present application also provides a computer storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the training method of the SOH balancing control model of the reconfigurable battery system described in the first aspect or the SOH balancing control method of the reconfigurable battery system described in the second aspect.

[0038] The training method of the SOH balancing control model of the reconfigurable battery system of the present application has the following beneficial effects compared with the related art:

[0039] 1. By constructing a Markov decision model for the SOH equalization control of a reconfigurable battery system, the SOH equalization control problem of the reconfigurable battery system is abstracted into a Markov decision process, providing a systematic framework for the optimization of control decisions and a theoretical basis for the reinforcement learning process. Additionally, by correspondingly determining appropriate state spaces, action spaces, and reward functions, the relative state of charge parameter is used to characterize the battery SOH, and the relative state of charge parameter is respectively related to the state of charge parameter and the maximum discharge depth of the battery cell. Therefore, frequent measurement or estimation of SOH can be avoided, which is conducive to the simple implementation of SOH equalization control and the training of the equalization control model.

[0040] 2. By constructing a simulated battery system, based on the simulated battery system and the Markov decision model, sufficient sample parameters can be obtained and stored in the experience replay area. Then, the data in the experience replay area can be used as training samples to train a neural network model based on the DQN algorithm to obtain a target equalization control model. Thus, the neural network model can be trained according to the historical data situation and control decisions can be made to solve the SOH equalization problem of the dynamic reconfigurable battery system. This data-driven algorithm does not require the establishment of complex rules or optimization models for the reconfigurable battery system, but learns from historical data and updates its own strategy, with advantages such as simple implementation and high real-time performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0042] Figure 1 It is a schematic flowchart of the training method for the SOH equalization control model of the reconfigurable battery system in an embodiment of the present application;

[0043] Figure 2 It is a schematic structural diagram of the reconfigurable battery system in an embodiment of the present application;

[0044] Figure 3 It is a schematic training framework diagram of the neural network model based on the DQN algorithm in an embodiment of the present application;

[0045] Figure 4 It is a schematic construction process diagram of the simulated battery system in an embodiment of the present application;

[0046] Figure 5 It is a schematic interaction environment diagram of the simulated battery system and the Markov decision model in an embodiment of the present application;

[0047] Figure 6 It is a schematic flow chart of step S103 in an embodiment of the present application;

[0048] Figure 7 It is a schematic flow chart of the SOH balancing control method for a reconfigurable battery system in an embodiment of the present application;

[0049] Figure 8 It is a schematic structural diagram of a training device for an SOH balancing control model of a reconfigurable battery system in an embodiment of the present application;

[0050] Figure 9 It is a schematic structural diagram of an SOH balancing control device for a reconfigurable battery system in an embodiment of the present application;

[0051] Figure 10 It is a schematic structural diagram of an electronic device in an embodiment of the present application. Detailed implementation manners

[0052] Next, in combination with the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0053] In some embodiments, as Figure 1 shown, Figure 1 It is a schematic flow chart of a training method for an SOH balancing control model of a reconfigurable battery system in an embodiment of the present application; The present application provides a training method for an SOH (state of health) balancing control model of a reconfigurable battery system. The reconfigurable battery system includes a plurality of battery units and a plurality of switch units, and the operating state of the battery unit is related to the switch state of the corresponding switch unit. Exemplarily, the reconfigurable battery system can be a structure as Figure 2 shown. The reconfigurable battery system includes at least one series battery pack, and each series battery pack is connected in parallel. The series battery pack includes a plurality of battery units (such as B1 to Bn in Figure 2 ), a plurality of first switch units (such as S1 to Sn in Figure 2 ) and a plurality of second switch units (such as S1' to Sn' in Figure 2 ). The first switch unit is connected in series with the corresponding battery unit, and the second switch unit is connected in parallel with the corresponding battery unit. By controlling the conduction states of the first switch unit and the second switch unit, it is possible to control whether the corresponding battery unit is connected to the reconfigurable battery system. This training method includes the following steps S101 to S103.

[0054] S101: Construct a Markov decision model for SOH balancing control of a reconfigurable battery system. The Markov decision model includes a state space, an action space, and a reward function. The state space includes a set of state parameters of the reconfigurable battery system. The state parameters of the reconfigurable battery system include a state of charge parameter, a relative state of charge parameter, and a switch state parameter. The relative state of charge parameter is related to the state of charge parameter and the maximum discharge depth of the battery cell respectively. The action space includes a set of switch state parameters. The reward function reflects the difference between the relative state of charge parameter variable and the average relative state of charge parameter during the discharge process of the battery.

[0055] Among them, due to the difficulty of SOH measurement and the complexity of establishing an estimation model, this application proposes an intermediate parameter, the relative state of charge parameter (denoted by RSOC). RSOC is an intermediate variable formed by combining the SOC of the battery cell and the maximum discharge depth under the SOH balancing condition, as shown in the following formula:

[0056]

[0057] Among them, DODmax is the maximum discharge depth of the battery cell, which is calculated based on the initial SOH and the expected number of cycle periods of the battery cell. RSOC can uniformly represent the discharge depth of all batteries at the end of each cycle period, which is beneficial to the implementation of the control algorithm. For a battery cell in a fully charged initial state, the initial value of SOC is 100%, and the initial value of RSOC is 100%. When the predetermined maximum discharge depth of the battery cell is reached within a single cycle period, the end value of SOC discharge is 100% - DODmax, and the end value of RSOC discharge is 0%. This application maps the discharge process of the battery from 100% state of charge to (100% - DODmax) state of charge to the discharge process of RSOC from 100% to 0%. When all batteries reach the state where RSOC is equal to 0% simultaneously within the expected cycle period, the SOH balance of the entire reconfigurable battery system can be achieved.

[0058] It should be noted that the Markov decision process is a mathematical model used to describe a system with randomness and decision-making processes. In this application, the SOH balancing control problem of the reconfigurable battery system is abstracted as a Markov decision process, providing a systematic framework for optimizing control decisions and a theoretical basis for the reinforcement learning process. During the operation of the reconfigurable battery system, there are multi-dimensional state values and various behavior patterns. Storing the state value of the reconfigurable battery system at a certain moment in a set forms the state space of the reconfigurable battery system. Similarly, storing the switching behavior patterns of the reconfigurable battery system in a set forms the action space of the reconfigurable battery system. The state values of the reconfigurable battery system can include the state of charge parameter (denoted as S0C), relative state of charge parameter (denoted as RSOC), depth of discharge (denoted as DOD), the closed or off state of the switching unit of the reconfigurable battery system (denoted as SS), etc. For example, for a branch composed of 10 series-connected battery cells, SOC1, SOC2, SOC3,..., SOC10 can represent the SOC of each of the 10 batteries, RSOC1, RSOC2, RSOC3,..., RSOC10 can represent the RSOC values of the 10 batteries respectively, and SS1, SS2, SS3,..., SS10 can represent the switching states of the 10 batteries respectively. Among them, SS is a Boolean variable with values of 0 or 1, representing the disconnection and connection of the battery cell in the reconfigurable battery system respectively.

[0059] The state space is the set of state parameters of the reconfigurable battery system at a certain moment, denoted by S, and can be specifically expressed as the following formula:

[0060] S = {RSOC1, RSOC2,..., RSOC n , SOC1, SOC2..., SOC n , SS1, SS2,..., SS n} (2)

[0061] Among them, n represents the number of series-connected battery monomers in the reconfigurable battery system.

[0062] The action of the reconfigurable battery system refers to the switching operation of whether the battery cell is connected to the circuit. For a system with n series-connected battery cells, in order to ensure that the system has the ability of fault isolation and flexible control, n - k spare batteries can be designed, where k represents the number of batteries required to meet the external load demand of the system. The action of the reconfigurable battery system is determined according to the state value of the battery. Therefore, the action space can be expressed as a set related to the state, denoted as A(s), and is expressed by the following formula.

[0063] A(s) = {SS1, SS2, SS3,..., SSn} (3)

[0064] According to the binomial theorem, the number of behavioral patterns contained in A(s) can be determined as shown in the following formula:

[0065]

[0066] where Na refers to the number of all actions of the reconfigurable battery system. Taking the reconfigurable battery system with 10 battery cells as an example, if 3 batteries are used as redundant batteries of the system, there are 120 possible action spaces in total. Specifically, as shown in the following table:

[0067]

[0068] where 1 and 0 represent that the battery cell is connected to the reconfigurable battery system or disconnected from the reconfigurable battery system.

[0069] The design of the reward function needs to reflect the balanced control objective of the reconfigurable battery system. This application designs a reward function reward based on the RSOC variable t , specifically as follows:

[0070]

[0071] This reward function has two characteristics: First, it reflects the difference between the RSOC variable and the average RSOC value during the discharge process of the battery, that is, the balance degree of RSOC. The greater the balance degree of RSOC of the behavior agent, the greater the reward it can obtain. Second, as the discharge process progresses, RSOC gradually decreases to 0. At this time, under the constraint of the reward function, the DOD of each battery cell will tend to the preset maximum value, that is, the necessary condition for SOH balance.

[0072] After establishing the state space, action space, and reward function of the balanced control task of the reconfigurable battery system, the Markov decision process can be realized according to this information.

[0073] At any time step t, starting from state s, the expected cumulative reward following policy π is calculated, and the state value function is as follows:

[0074]

[0075] where γ represents the discount factor, a t represents the action a at time t. The state selection principle will determine the long-term benefits of different current states according to the magnitude of Vπ(s). The larger Vπ(s) is, the higher the long-term benefit of state s is.

[0076] After state s executes action a, the expected cumulative reward following policy π is calculated, and the action value function is as follows:

[0077]

[0078] S102: Obtain sample parameters based on the Markov decision model and the simulated battery system, and save the sample parameters in the experience replay area until the number of data stored in the experience replay area reaches a preset threshold. The sample parameters include the state parameters, action parameters, reward parameters, and predicted state parameters of the reconfigurable battery system.

[0079] It can be understood that based on the Markov decision model and the simulated battery system, simulation parameters can be obtained. Taking the simulation parameters as sample parameters and saving them in the experience replay area can save historical interaction data until the number of data stored in the experience replay area reaches the corresponding data volume requirement, thereby providing data support for subsequent model training.

[0080] S103: Use the data in the experience replay area as training samples to train a neural network model based on the DQN algorithm to obtain a target equilibrium control model.

[0081] Exemplarily, as Figure 3 shown, Figure 3 is a schematic diagram of the training framework of the neural network model based on the DQN algorithm in an embodiment of the present application. The training process of the target equilibrium control model can adopt the training framework as Figure 3 shown.

[0082] It should be noted that before the start of each training cycle, the data in the experience replay area can be randomly shuffled at the batch level to reduce the negative impact of temporal correlation on neural network training by destroying the temporal continuity of the original data. This mechanism not only improves the reuse frequency of valuable samples but also effectively prevents the model from falling into a local optimal solution.

[0083] It can be understood that batch samples can be drawn from the experience replay area through random sampling techniques, and these samples can be used to iteratively train the valuation neural network of DQN. At the same time, the target neural network parameters are updated regularly to stabilize the training process. This method effectively alleviates the common training instability problem in reinforcement learning by breaking the temporal correlation between consecutive samples, and finally enables the neural network model to learn the optimal approximation of the state-action value function, forming a target model with equilibrium control ability, which is applicable to energy management scenarios that require dynamic decision-making, realizing continuous optimization and stable equilibrium of the system state. The whole process integrates the sample reuse characteristics of experience replay and the deep representation ability of DQN, ensuring the generalization and robustness of the model in complex environments.

[0084] The training method of the above-mentioned SOH balancing control model for a reconfigurable battery system constructs a Markov decision model based on the SOH balancing control of the reconfigurable battery system, thereby abstracting the SOH balancing control problem of the reconfigurable battery system into a Markov decision process, providing a systematic framework for the optimization of control decisions, and providing a theoretical basis for the reinforcement learning process. At the same time, by correspondingly determining the appropriate state space, action space, and reward function, the relative state of charge parameter is used to characterize the battery SOH, and the relative state of charge parameter is respectively related to the state of charge parameter and the maximum discharge depth of the battery cell. Therefore, frequent measurement or estimation of SOH can be avoided, which is beneficial to the simple implementation of the SOH balancing control and the training of the balancing control model. In addition, by constructing a simulation battery system, sufficient sample parameters can be obtained based on the simulation battery system and the Markov decision model and stored in the experience replay area. Then, the data in the experience replay area can be used as training samples to train the neural network model based on the DQN algorithm to obtain the target balancing control model, so that the neural network model can be trained according to the historical data situation and control decisions can be made to solve the SOH balancing problem of the dynamic reconfigurable battery system.

[0085] In some embodiments, before step S102, as Figure 4 shown, the training method of the SOH balancing control model for the reconfigurable battery system may further include the following steps S401 to S403.

[0086] S401: Obtain the initial SOH values of each battery cell in the reconfigurable battery system and the reference discharge depth of the target battery cell. The target battery cell is the battery cell with the highest initial SOH value.

[0087] In applications, the aging of the battery is affected by factors such as ambient temperature, discharge rate, and discharge depth. SOH can usually be defined from aspects such as capacity, internal resistance, and cycle life. In this application, SOH can be defined from the perspective of cycle life as follows:

[0088]

[0089] Among them, Cleft refers to the remaining discharge cycle times of the battery from the current state to the retirement level, and Ctotal refers to the discharge cycle times experienced by the battery from a new battery to the retirement level.

[0090] For the SOH at any initial time t, it can be derived from the SOH in the initial state as shown in the following formula:

[0091]

[0092] Among them, SOHt represents the SOH value of the battery cell at time t, SOHinit refers to an initial state that is not limited to the new battery state, and Cacu refers to the number of discharge cycles experienced by the battery cell from SOHinit to SOHt. Therefore, the initial SOH value of each battery cell in the reconfigurable battery system can be determined according to formula (8) and formula (9). In addition, the above Cleft, Ctotal, and Cacu are all measured at the same discharge depth.

[0093] In order to maximize the overall release capacity of the reconfigurable battery system, the reference discharge depth of the target battery cell may be the maximum discharge depth of the target battery cell, and the maximum discharge depth may be preset. For example, in order to avoid damage to the electrode caused by excessive discharge depth, the maximum discharge depth of the target battery cell may be set to 80%.

[0094] S402: Determine the discharge depth of other battery cells according to the reference discharge depth and the initial SOH value of each battery cell.

[0095] In the application, taking the ambient temperature of 25 degrees Celsius as an example, the relationship between the total number of cycles of the battery cell and the depth of discharge (DOD) is as follows:

[0096] C total =a·DOD -b (10)

[0097] Wherein a and b are parameters of the battery, which are related to the type of battery. In one example, the corresponding parameters a of a lithium battery may be 694, and b may be 0.795.

[0098] Combining the above three equations (8), (9) and (10), we get the calculation formula of SOH at any initial state at time t:

[0099]

[0100] For two different batteries i and j, their SOH at time t can be known as follows:

[0101]

[0102] According to the equilibrium condition, that is, SOHi,t=SOHj,t at time t, and Cacu,i=Cacu,j, it can be obtained that when the discharge depth and cycle number of battery i are known, the theoretical value of the discharge depth of battery j under the same cycle number can be calculated as follows:

[0103]

[0104] Taking the i-th battery as the target battery cell and the j-th battery as the other battery cell to be determined, the discharge depth of the other battery cell, i.e., DOD, can be determined based on Formula (13) according to the reference discharge depth and the initial SOH values of each battery cell. max,i is the maximum discharge depth of the target battery cell, DOD max,j is the maximum discharge depth of the other battery cell to be determined, SOH i,init is the initial SOH value of the target battery cell, SOH j,init is the initial SOH value of the other battery cell to be determined, C acu,i is the reference cycle number of the target battery cell.

[0105] It can be understood that the inconsistency of the initial SOH values determines the difference in the discharge depth of battery cells in each cycle. After determining the discharge depth of the other battery cells according to the reference discharge depth and the initial SOH values of each battery cell, the balanced control of the SOH of the reconfigurable battery system can be achieved by controlling the discharge depths of different batteries at a specific cycle number. For example, during the discharge process, calculate the RSOC value at each time step and design a control scheme so that all battery cells can reach the state where RSOC is 0 in each cycle until the discharge cycle number reaches the predetermined cycle number.

[0106] S403: Set the initial state of charge parameters of each battery cell to the same predetermined parameter, take the reference cycle number of the target battery cell as the expected cycle number of the reconfigurable battery system, and construct a simulation battery system based on the discharge depths of each battery cell.

[0107] Exemplarily, as Figure 5 shown, Figure 5 is a schematic diagram of the interaction environment between the simulation battery system and the Markov decision model in an embodiment of the present application. The interaction process between the simulation battery system and the Markov decision model can adopt the Figure 5 shown process.

[0108] Among them, the initial state of charge parameters of each battery cell can be set to 100%. The cycle number of the reconfigurable battery system is the aforementioned specific cycle number. Setting the initial state of charge parameters of each battery cell to the same predetermined parameter, taking the reference cycle number of the target battery cell as the expected cycle number of the reconfigurable battery system, and constructing a simulation battery system based on the discharge depths of each battery cell can achieve the balanced control of the SOH of the simulation battery system by controlling the discharge depths of different batteries at a specific cycle number, so that a simulation battery system based on the principle of SOH balanced control of the reconfigurable battery system can be constructed.

[0109] When the reference cycle number of the target battery cell is the minimum expected cycle number for achieving balance, and the reference depth of discharge of the target battery cell is the maximum depth of discharge of the target battery cell, based on the above SOH balance theory, this simulated battery system can further maximize the total released capacity of the reconfigurable battery system. The specific principle is as follows:

[0110] Define the discharge capacity value of the battery cell in a single cycle as follows:

[0111] Q i = DOD i ·Q rated (14)

[0112] Where Qi is the total released capacity of the ith battery in a single cycle, and Qrated is the rated capacity of the battery cell. Then the calculation method of the total released capacity of the battery cell in the predetermined cycle can be expressed by the following formula:

[0113]

[0114] Where Qi,total refers to the total released capacity of battery cell i in the entire predetermined cycle, and Cacu is the predetermined number of cycles.

[0115] According to the derivation process above and formulas (11), (14), and (15), Qi,total can be expressed as the following formula:

[0116]

[0117] The total released capacity of the reconfigurable battery system can be expressed as:

[0118]

[0119] Among them, for the balance of n batteries, it is ultimately necessary to ensure that they are balanced to the same level, and the battery with the highest SOH must have the largest discharge capacity. Then, in formula (17), it is equivalent to This part on the right side of the dot multiplication symbol will not change. In formula (17), A decreasing function, that is, the smaller the number of cycle periods, the larger the discharge capacity. According to Equation (11), if both the initial value and the final value are determined, the number of cycle periods and the depth of discharge show an opposite trend. The larger the DOD, the smaller the number of cycle periods. Since the maximum depth of discharge of the target battery cell has been predetermined, for example, preset to 80%, the minimum number of cycle periods of the target battery voltage can be calculated, and the other battery cells can be calculated according to Equation (13). On this basis, for the overall reconfigurable battery system, when the SOH of all batteries reaches the retirement level during system balancing, and when the predetermined number of cycle periods Cacu reaches the minimum value under this limit condition, the released energy of the overall system reaches the maximum value.

[0120] According to the initial SOH of the target battery cell, the SOH value at the retirement level, and the depth of discharge DOD, the minimum number of cycle periods can be calculated by the following formula:

[0121]

[0122] where C acu,min is the minimum number of cycle periods for achieving balance (used as the reference number of cycle periods here), SOHt,retire is the SOH value at the retirement level, a and b are battery parameters related to the battery type, SOHmax,init is the maximum initial SOH of the target battery cell, and DODlimit is the maximum depth of discharge of the target battery cell.

[0123] It can be understood that after constructing a simulation system based on the RSOC-based balance control and capacity maximization method, sample parameters are obtained based on the Markov decision model and the simulation battery system, and the sample parameters are saved in the experience replay area. Then, the data in the experience replay area is used as training samples to train the neural network model based on the DQN algorithm, and a target balance control model that takes into account the maximization of the released energy of the reconfigurable battery system can be obtained.

[0124] It should also be noted that the constructed simulation battery system has functions such as monitoring the SOC, RSOC, and switch status of battery cells, can simulate the action process of switches and the cycle aging process of batteries, and can update the state values of the reconfigurable battery system in combination with switch actions and established rules. Before and after the switch action, the change in the SOC of the battery cell follows the following formula:

[0125]

[0126] where SOCi,t and SOCi,t+1 respectively refer to the SOC values of battery cell i at time t and time t + 1, I is the current value from time t to time t + 1, SSi,t represents the switch status of single battery i from time t to time t + 1, and Qrated is the rated capacity of the battery.

[0127] The RSOC change of the battery cell follows the law of the following formula:

[0128]

[0129] Among them, RSOCi,t+1 represents the RSOC value of the battery cell at time t+1. Based on formula (19) and formula (20), the state parameters at different time steps of the cycle period can be determined, so that the sample parameters can be obtained based on the Markov decision model and the simulated battery system, providing data support for subsequent model training.

[0130] In some embodiments, sample parameters are obtained based on the Markov decision model and the simulated battery system, and the sample parameters are stored in the experience replay area, including: using the simulated battery system as the environment of the Markov decision model, obtaining the state parameters and action parameters of the simulated battery system at each time step of the cycle period, and solving the Markov decision model to obtain the reward parameters and predicted state parameters at each time step of the cycle period; the step of storing the state parameters, action parameters, reward parameters and predicted state parameters at the same time step in the experience replay area in the form of a quadruple storage structure.

[0131] It can be understood that in the training of the battery equalization control model based on deep reinforcement learning, first, the simulated battery system is used as the dynamic environment of the Markov decision model. By simulating the charge and discharge process of the battery pack during the cycle period, the state parameters (such as SOC and DOD, etc.) and action parameters (such as the switch state of the reconfigurable battery system) at each time step are collected in real time; at the same time, the Markov decision model is used to calculate the reward parameters corresponding to the time step, and the state parameters at the next moment (i.e., the predicted state parameters) are obtained through model prediction. Subsequently, the state parameter (s), action parameter (a), reward parameter (r) and predicted state parameter (s') at the same time step are encapsulated into a quadruple (s, a, r, s'), and stored in the circular buffer of the experience replay area in the order of time series. This process continuously collects, calculates and saves multiple groups of quadruple data by traversing all time steps of the cycle period, forming an experience dataset containing the dynamic behavior of the reconfigurable battery system and the response of the control strategy. This structured storage method not only retains the time correlation of the battery state evolution, but also provides a standardized input for the batch sampling training of the subsequent DQN algorithm, supports the repeated use of historical interaction data in offline training, effectively improves the data efficiency and alleviates the sample correlation problem in reinforcement learning, and finally provides high-quality training samples for the optimization of the target equalization control model. The entire process realizes the closed-loop integration of the simulation environment, mathematical model and experience storage, ensuring the dynamic adaptability and control accuracy of the battery equalization control strategy in practical applications.

[0132] In some embodiments, in step S103, such as Figure 6As shown, the data in the experience replay area is used as training samples to train the neural network model of the DQN algorithm to obtain the target equilibrium control model, including the following steps S601 to S604.

[0133] S601: Obtain the sample data in the experience replay area.

[0134] Among them, the experience replay area stores sample data composed of quadruples of state parameters (s), action parameters (a), reward parameters (r), and predicted state parameters (s'). Therefore, the sample data can be directly obtained from the experience replay area. For example, during the iterative training process, a batch of (s, a, r, s') data can be randomly selected from the experience replay area as training data, which can break the temporal correlation between the data, avoid the training process falling into a local optimal solution, and enable the model to learn the optimal strategy under different states more comprehensively.

[0135] S602: Pass the state parameters and action parameters in the sample data to the evaluation neural network of the neural network model, and pass the predicted state parameters to the target neural network of the neural network model. The reward parameter and the target neural network are combined to generate a training label value.

[0136] Among them, the DQN network is represented by two neural networks with the same structure, which are called the evaluation neural network and the target neural network respectively. The evaluation neural network updates the network weights in each step of the action, while the target neural network copies the corresponding weights from the evaluation neural network according to a manually set interval. The output of the target neural network and the reward value together constitute the label of the training evaluation network. The existence of the target neural network helps to reduce the target fluctuation, improve the stability of the algorithm, and make the training process more reliable.

[0137] It should be noted that when the state s and the action a are input into the evaluation neural network, the evaluation neural network can estimate the Q value obtained by taking the action a in the state s, that is, evaluate the value of the current state-action pair. The Q value reflects the expected cumulative reward obtained from the state until the end after taking a certain action in a certain state. The structure of the target neural network is basically the same as that of the evaluation neural network, but the parameter update frequency is lower. Therefore, passing the state s' to the target neural network can keep the target value relatively stable within a certain period of time, which helps the convergence of the training process.

[0138] The reward r and the target neural network are combined to generate a training label value. Specifically, according to the Bellman equation, the maximum Q value in the state s' is calculated using the target neural network, and then the current obtained reward r is added to obtain a target Q value, which is the training label value.

[0139] The training label value y of the networki It can be expressed by the following formula:

[0140]

[0141] S603: Update the weight parameters of the neural network model according to the training label value, the preset loss function, and the optimization function.

[0142] It can be understood that after determining the training label value, by comparing the difference between the Q value output by the evaluation neural network and the training label value, a loss function such as the mean square error can be used to measure this difference. Then, using the backpropagation algorithm, the parameters of the evaluation neural network are updated according to the gradient of the loss function, so that the Q value output by the evaluation neural network gradually approaches the training label value. As the iterative training progresses, the evaluation neural network will more and more accurately estimate the Q value of each state-action pair, enabling the intelligent agent to learn the optimal strategy and make the most favorable decisions in different environmental states.

[0143] Among them, the loss function during the training process can be:

[0144] L(θ) = (y i DQN -Q(s,a;θ)) 2 (22)

[0145] The gradient descent method is adopted during the training process, which can be expressed as:

[0146]

[0147] The update formula for the weight parameters of the neural network is as follows:

[0148]

[0149] where η is the learning rate.

[0150] It should also be noted that during the learning process of the DQN network, the agent selects multiple possible actions based on the current state. However, each action will result in different Q-values output by the neural network. To achieve optimal control decisions, the action with the highest Q-value should be selected during the training process and used as the control signal to be applied to the environment, so that the environment can make corresponding actions. As the agent continues to train, the Q-network will gradually converge, and the agent can make decisions based on the knowledge in the current network and select the best action in a given state. To encourage the agent to explore other possible action paths and avoid getting stuck in local optimal solutions, this application can adopt the ε-greedy strategy. This strategy can balance the trade-off between exploring new states and exploiting the existing optimal strategy by adjusting the value of ε. When ε is large, the agent tends to choose random actions to explore new possibilities; when ε is small, the agent will tend to choose the optimal action estimated by the current Q-network, thus exploiting the known experience. The ε-greedy strategy is specifically implemented according to the following formula:

[0151]

[0152] where ε ∈ [0, 1], and Na is the number of actions in state s. π(a|s) represents the probability of taking action a in state s.

[0153] S604: Repeat the above training steps until the number of cycle periods of the reconfigurable battery system reaches the expected number of cycle periods and the relative state of charge parameter reaches 0%. It should be noted that considering error factors, when the relative state of charge parameter reaches near 0%, it can be considered that the relative state of charge parameter reaches 0%. Specifically, a preset interval with 0% as the lower limit value can be set, and when the relative state of charge parameter is within this preset interval, it is determined that the relative state of charge parameter reaches 0%.

[0154] It can be understood that based on the above description, this application can perform balancing control based on RSOC, and can use the minimum number of cycle periods of the target battery cell as the expected number of cycle periods of the reconfigurable battery system to maximize the capacity of the reconfigurable battery system. Then, taking the number of cycle periods of the battery reaching the expected number of times and the RSOC reaching 0% as the convergence condition can enable the control decision made by the target balancing control model to take into account both balancing control and the maximization of the capacity of the reconfigurable battery system.

[0155] In some embodiments, this application also provides a method for balancing the SOH of a reconfigurable battery system, as Figure 7 shown. This method for balancing the SOH of a reconfigurable battery system may include the following steps S701 to S703.

[0156] S701: Obtain the initial parameters of the reconfigurable battery system in the initial state.

[0157] Among them, the initial parameters may include state parameters and action parameters. The state parameters may include state of charge parameters (denoted as S0C), relative state of charge parameters (denoted as RSOC), depth of discharge (denoted as DOD), etc. The action parameters may include the closed or off state of the reconfigurable switch (denoted as SS).

[0158] S702: Input the initial parameters into the target equalization control model to obtain the control strategy of the reconfigurable battery system in a single cycle. Among them, the target equalization control model is trained by using the training method of the SOH equalization control model of the reconfigurable battery system in any of the above embodiments.

[0159] It should be noted that the target equalization control model can be trained in an online manner.

[0160] It can be understood that the target equalization control model has been trained by using the training method of the SOH equalization control model of the reconfigurable battery system in any of the above embodiments. Therefore, by inputting the initial parameters into the target equalization control model, the target equalization control model can output the optimal control strategy of the reconfigurable battery system in a single cycle, thereby realizing the equalization control of the reconfigurable battery system.

[0161] S703: Control the working state of the reconfigurable battery system according to the control strategy.

[0162] In this embodiment, the working state of the reconfigurable battery system may be the working state of the switch unit in the reconfigurable battery system. Controlling the working state of the reconfigurable battery system according to the control strategy can realize the equalization control of the reconfigurable battery system. When the training method of the SOH equalization control model of the reconfigurable battery system takes into account the maximization of the released energy of the reconfigurable battery system, using this control strategy to control the working state of the reconfigurable battery system can also take into account the realization of the maximization of the released energy of the reconfigurable battery system.

[0163] In some embodiments, please refer to Figure 8 , the embodiment of the present application also provides a training device 80 for the SOH equalization control model of the reconfigurable battery system. The training device 80 includes: a modeling module 81, a first acquisition module 82, and a training module 83.

[0164] The modeling module 81 is used to construct a Markov decision model for SOH balancing control of a reconfigurable battery system. Among them, the Markov decision model includes a state space, an action space, and a reward function. The state space includes a set of state parameters of the reconfigurable battery system. The state parameters of the reconfigurable battery system include a state of charge parameter, a relative state of charge parameter, and a switch state parameter. The relative state of charge parameter is respectively related to the state of charge parameter and the maximum discharge depth of the battery cell. The action space includes a set of switch state parameters. The reward function reflects the difference between the relative state of charge parameter variable and the average relative state of charge parameter of the battery during the discharge process.

[0165] The first acquisition module 82 is used to obtain sample parameters based on the Markov decision model and the simulated battery system, and store the sample parameters in the experience replay area until the number of data stored in the experience replay area reaches a preset threshold. Among them, the sample parameters include the state parameters, action parameters, reward parameters, and predicted state parameters of the reconfigurable battery system.

[0166] The training module 83 is used to use the data in the experience replay area as training samples to train a neural network model based on the DQN algorithm to obtain a target balancing control model.

[0167] In some embodiments, the modeling module 81 is further used to obtain the initial SOH value of each battery cell in the reconfigurable battery system and the reference discharge depth of the target battery cell. The target battery cell is the battery cell with the highest initial SOH value; determine the discharge depth of other battery cells according to the reference discharge depth and the initial SOH value of each battery cell; set the initial state of charge parameters of each battery cell to the same predetermined parameter, use the reference cycle number of the target battery cell as the expected cycle number of the reconfigurable battery system, and construct a simulated battery system based on the discharge depth of each battery cell.

[0168] In some embodiments, the first acquisition module 82 is further used to use the simulated battery system as the environment of the Markov decision model, obtain the state parameters and action parameters of the simulated battery system at each time step of the cycle, and solve the Markov decision model to obtain the reward parameters and predicted state parameters at each time step of the cycle; store the state parameters, action parameters, reward parameters, and predicted state parameters at the same time step in the experience replay area in the form of a quadruple storage structure.

[0169] In some embodiments, the training module 83 is further configured to obtain sample data from the experience replay area; transfer the state parameters and action parameters in the sample data to the value neural network of the neural network model, predict the state parameters and transfer them to the target neural network of the neural network model, and combine the reward parameters and the target neural network to generate a training label value; update the weight parameters of the neural network model according to the training label value, a preset loss function, and an optimization function; repeat the above training steps until the number of cycle periods of the reconfigurable battery system reaches the expected number of cycle periods and the RSOC reaches 0%.

[0170] It should be noted that the training device 80 for the SOH equalization control model of the reconfigurable battery system provided in the embodiments of the present application and the training method for the SOH equalization control model of the reconfigurable battery system provided in the embodiments of the present application are based on the same inventive concept. Therefore, the specific implementation of this embodiment can refer to the implementation of the aforementioned training method for the SOH equalization control model of the reconfigurable battery system, and the repeated parts will not be elaborated here.

[0171] In some embodiments, please refer to Figure 9 , the embodiments of the present application further provide a SOH equalization control device 90 for a reconfigurable battery system. The equalization control device 90 includes: a second acquisition module 91, an output module 92, and a control module 93.

[0172] The second acquisition module 91 is configured to obtain the initial parameters of the reconfigurable battery system in the initial state.

[0173] The output module 92 is configured to input the initial parameters into the target equalization control model to obtain the control strategy of the reconfigurable battery system in a single cycle period, where the target equalization control model is trained by the method as described in the first aspect.

[0174] The control module 93 is configured to control the working state of the reconfigurable battery system according to the control strategy.

[0175] It should be noted that the SOH equalization control device 90 for the reconfigurable battery system provided in the embodiments of the present application and the SOH equalization control method for the reconfigurable battery system provided in the embodiments of the present application are based on the same inventive concept. Therefore, the beneficial effects of this embodiment compared with the related technologies can refer to the description of the SOH equalization control method for the reconfigurable battery system, and will not be elaborated here.

[0176] In some embodiments, please refer to Figure 10 , Figure 10It is a schematic structural diagram of an electronic device in an embodiment of this application. An electronic device 100 provided in an embodiment of this application includes a processor 101 and a memory 102; the memory 102 stores a computer program, wherein the computer program, when executed by the processor, implements the training method of the SOH equalization control model of the reconfigurable battery system according to any of the above solutions or the SOH equalization control method of the reconfigurable battery system according to any of the above solutions.

[0177] Specifically, the processor 101 may include, for example, a general microprocessor, an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), and so on. The processor 101 may also include on-board memory for caching purposes. The processor 101 may be a single processing unit or multiple processing units for performing different actions of the method flow according to the embodiment of this application.

[0178] The memory 102 may be, for example, any medium capable of containing, storing, transmitting, propagating, or transporting instructions. For example, the memory 102 may include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, components, or propagation media. Specific examples of the memory 102 include: magnetic storage devices, such as magnetic tapes or hard disk drives (HDDs); optical storage devices, such as compact discs (CD-ROMs); may also be, such as random access memory (RAM) or flash memory; and / or wired / wireless communication links.

[0179] This application also provides a computer-readable medium, on which a computer program is stored, and when the program is executed by the processor, it is the training method of the SOH equalization control model of the reconfigurable battery system according to any of the above solutions or the SOH equalization control method of the reconfigurable battery system according to any of the above solutions. The computer-readable medium may be included in the device / device / system described in the above embodiment; or it may exist separately without being assembled into the device / device / system. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed, the method according to the embodiment of this application is implemented.

[0180] According to an embodiment of the present application, a computer-readable medium may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wired, optical fiber cable, radio frequency signal, etc., or any suitable combination of the above.

[0181] Those skilled in the art can understand that the features recited in the various embodiments and / or claims of the present application can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly recited in the present application. In particular, without departing from the spirit and teachings of the present application, the features recited in the various embodiments and / or claims of the present application can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present application. Therefore, the scope of the present application should not be limited to the above embodiments, but should be determined not only by the appended claims but also by the equivalents of the appended claims. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. A training method for a SOH equalization control model of a reconfigurable battery system, characterized in that The reconfigurable battery system includes a plurality of battery cells and a plurality of switch units, and the operating state of the battery cells is related to the switch states of the corresponding switch units; the method includes: Construct a Markov decision model based on the SOH equalization control of the reconfigurable battery system. The Markov decision model includes a state space, an action space, and a reward function. The state space includes a set of state parameters of the reconfigurable battery system. The state parameters of the reconfigurable battery system include state-of-charge parameters, relative state-of-charge parameters, and switch state parameters. The relative state-of-charge parameters are respectively related to the state-of-charge parameters and the maximum discharge depth of the battery cells. The action space includes a set of switch state parameters. The reward function reflects the difference between the relative state-of-charge parameter variable and the average relative state-of-charge parameter during the discharge process of the battery; Obtain sample parameters based on the Markov decision model and the simulated battery system, and save the sample parameters in the experience replay area until the number of data stored in the experience replay area reaches a preset threshold. The sample parameters include state parameters, action parameters, reward parameters, and predicted state parameters of the reconfigurable battery system; Use the data in the experience replay area as training samples to train a neural network model based on the DQN algorithm to obtain a target equalization control model.

2. The training method of the SOH equalization control model for the reconfigurable battery system according to claim 1, characterized in that, Before obtaining the sample parameters based on the Markov decision model and the simulated battery system, the method further includes: Obtain the initial SOH values of each battery cell in the reconfigurable battery system and the reference discharge depth of the target battery cell. The target battery cell is the battery cell with the highest initial SOH value; Determine the discharge depths of other battery cells according to the reference discharge depth and the initial SOH values of each battery cell; Set the initial state-of-charge parameters of each battery cell to the same predetermined parameter, use the reference cycle number of the target battery cell as the expected cycle number of the reconfigurable battery system, and construct a simulated battery system based on the discharge depths of each battery cell.

3. The training method of the SOH equalization control model for the reconfigurable battery system according to claim 1, wherein, The obtaining the sample parameters based on the Markov decision model and the simulated battery system and saving the sample parameters in the experience replay area includes: Use the simulated battery system as the environment of the Markov decision model, obtain the state parameters and action parameters of the simulated battery system at each time step of the cycle, and solve the Markov decision model to obtain the reward parameters and predicted state parameters at each time step of the cycle; Save the state parameters, action parameters, reward parameters, and predicted state parameters at the same time step in the experience replay area in the form of a quadruple storage structure.

4. The training method of the SOH equalization control model of the reconfigurable battery system according to claim 1, characterized in that The using the data in the experience replay area as training samples to train a neural network model based on the DQN algorithm to obtain a target equalization control model includes: Obtain the sample data in the experience replay area; Transfer the state parameters and action parameters in the sample data to the evaluation neural network of the neural network model, transfer the predicted state parameters to the target neural network of the neural network model, and combine the reward parameters and the target neural network to generate a training label value; Update the weight parameters of the neural network model according to the training label values, a preset loss function, and an optimization function; Repeat the above training steps until the number of cycle periods of the reconfigurable battery system reaches the expected number of cycle periods and the relative state of charge parameter reaches 0%.

5. The training method of the SOH equalization control model of the reconfigurable battery system according to claim 2, characterized in that, The calculation formula for the reference number of cycle periods of the target battery cell is: Among them, C acu,min is the reference cycle number at equilibrium, SOHt,retire is the SOH value at the retirement level, a and b are battery parameters related to the battery type, SOHmax,init is the maximum initial SOH of the battery cell, and DODlimit is the maximum depth of discharge of the battery cell.

6. The training method of the SOH equalization control model of the reconfigurable battery system according to claim 2, characterized in that, Determine the discharge depths of other battery cells according to the reference discharge depth and the initial SOH values of the battery cells, including: Based on the reference discharge depth and the initial SOH values of the battery cells, determine the discharge depths of other battery cells based on a preset depth calculation formula; wherein, the preset depth calculation formula is: Among them, DOD max,i is the maximum depth of discharge of the target battery cell, DOD max,j is the maximum depth of discharge of other battery cells to be determined, SOH i,init is the initial SOH value of the target battery cell, SOH j,init is the initial SOH value of other battery cells to be determined, C acu,i is the reference cycle number of the target battery cell, a and b are battery parameters related to the battery type, at time t, SOHi,t = SOHj,t, and Cacu,i = Cacu,j.

7. The training method of the SOH equalization control model of the reconfigurable battery system according to claim 1, wherein, The association formula between the relative state of charge parameter and the state of charge parameter and the maximum discharge depth of the battery is as follows: Among them, RSOC is the relative state of charge parameter of the battery cell, SOC is the state of charge parameter of the battery cell, and DOD max is the maximum depth of discharge of the battery cell.

8. A method for SOH balancing control of a reconfigurable battery system, characterized in that, The SOH equalization control method for the reconfigurable battery system includes: Obtain the initial parameters of the reconfigurable battery system in the initialization state; Input the initial parameters into the target equalization control model to obtain the control strategy of the reconfigurable battery system in a single cycle period, wherein the target equalization control model is trained by using the method according to any one of claims 1 to 7; Control the working state of the reconfigurable battery system according to the control strategy.

9. An electronic device, comprising a processor and a memory; the memory stores a computer program, wherein, The computer program, when executed by the processor, implements the training method of the SOH equalization control model of the reconfigurable battery system according to any one of claims 1 to 7 or the SOH equalization control method of the reconfigurable battery system according to claim 8.

10. A computer storage medium, characterized in that, A computer program is stored thereon, wherein the computer program, when executed by the processor, implements the training method of the SOH equalization control model of the reconfigurable battery system according to any one of claims 1 to 7 or the SOH equalization control method of the reconfigurable battery system according to claim 8.