Balanced charging method, device and equipment for power battery pack
Through the multi-head self-attention network and the Actor-Critic deep reinforcement learning framework, combined with the multi-stage equalization circuit structure, dynamically generates an equalization strategy, solving the overcharge or over-discharge problems caused by the performance differences of single batteries in the power battery pack, and achieving efficient and safe balanced charging effect.
Patent Information
- Application Number
- CN202510258643.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-03-06
AI Technical Summary
The performance parameters of each single battery in the power battery pack lead to overcharging or overdischarge, which affects service life and safety. The traditional balanced charging method is inefficient and cannot adapt to different temperatures and working conditions.
A multi-head self-attention network is used to extract features, generate a state-associated feature matrix, and dynamically generate a combination strategy of equalization current, mode and time through the Actor-Critic deep reinforcement learning framework, combining the coordinated work of the main equalization loop and the secondary equalization loop to achieve adaptive adjustment of equalization control.
It improves the equalization efficiency, reduces energy loss, realizes the temperature adaptive control of the equalization circuit, ensures the safety and reliability of the equalization charging process, and extends the service life of the battery pack.
Smart Images

Figure CN119765584B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of battery management, and in particular to a balanced charging method, device and equipment for a power battery pack. Background Art
[0002] With the rapid development of the new energy vehicle industry, the performance and life of power battery packs as core components have attracted much attention. In actual applications, due to the influence of factors such as manufacturing process, ambient temperature, and operating conditions, the performance parameters of each single cell in the power battery pack will gradually differ. This inconsistency will cause some cells to be overcharged or over-discharged, seriously affecting the service life and safety of the battery pack.
[0003] Traditional equalization charging methods mainly use passive equalization or simple active equalization strategies. Although these methods can improve the consistency of cells to a certain extent, the equalization efficiency is low and cannot adapt to the dynamic change characteristics of the battery pack under different temperatures, charge and discharge rates and cycle times. Especially for large-scale parallel power battery packs, due to the large number of cells, traditional equalization methods are difficult to accurately identify the cells that need to be balanced, which easily leads to waste of equalization resources. Summary of the invention
[0004] The present invention provides a balanced charging method, device and equipment for a power battery pack, which reduces energy loss while ensuring the balanced efficiency of the power battery pack and realizes temperature adaptive control of a balanced circuit.
[0005] In a first aspect, the present invention provides a balanced charging method for a power battery pack, the balanced charging method for a power battery pack comprising:
[0006] Collect parameters of the power battery pack, measure the terminal voltage, temperature, and charge and discharge current of each single battery, and calculate the state of charge and internal resistance change rate of each single battery;
[0007] The terminal voltage value, the temperature value, the charge and discharge current value, the state of charge value and the internal resistance change rate are constructed into a state feature sequence, and input into a multi-head self-attention network for feature extraction to generate a state correlation feature matrix;
[0008] Based on the state association feature matrix, a three-layer fully connected Actor network and a dual-Q structure Critic network are constructed, and a balanced current value, a constant current or constant voltage balanced mode and a balanced time are output through the Actor network;
[0009] According to the balancing current value, the balancing mode and the balancing time, the DC-DC converter in the main balancing loop is controlled to generate a constant current output, or the Buck-Boost converter in the secondary balancing loop is controlled to generate a constant voltage output.
[0010] In a second aspect, the present invention provides a balanced charging device for a power battery pack, and the balanced charging device for the power battery pack includes:
[0011] A parameter acquisition module, configured to acquire parameters of the power battery pack, measure the terminal voltage values, temperature values, and charge and discharge current values of each single battery, and calculate the state of charge values and internal resistance change rates of each single battery;
[0012] A feature extraction module, configured to construct the terminal voltage values, the temperature values, the charge and discharge current values, the state of charge values, and the internal resistance change rates into a state feature sequence, and input the state feature sequence into a multi-head self-attention network for feature extraction to generate a state correlation feature matrix;
[0013] A construction module, configured to construct a three-layer fully connected Actor network and a Critic network with a double Q structure based on the state correlation feature matrix, and output an equalization current value, a constant current or constant voltage equalization mode, and an equalization time through the Actor network;
[0014] A control module, configured to control a DC-DC converter in a main equalization circuit to generate a constant current output, or control a Buck-Boost converter in a secondary equalization circuit to generate a constant voltage output according to the equalization current value, the equalization mode, and the equalization time.
[0015] In a third aspect of the present invention, there is provided a balanced charging device for a power battery pack, including: a memory and at least one processor, wherein instructions are stored in the memory; the at least one processor calls the instructions in the memory to enable the balanced charging device for the power battery pack to execute the above-mentioned balanced charging method for the power battery pack.
[0016] In the technical solution provided by the present invention, by introducing a multi-head self-attention network for feature extraction, the efficient processing of the state information of the power battery pack is realized, the mutual correlation between different monomers is accurately captured, and the expression ability of the state features is improved. The Actor-Critic deep reinforcement learning framework is adopted to dynamically generate a combined strategy of balancing current, balancing mode and balancing time, realizing the adaptive adjustment of the balancing control, so that the balancing strategy can be optimized according to the real-time state of the battery pack. A multi-level balancing circuit structure in which the main balancing circuit and the secondary balancing circuit work together is designed, and a balancing mode switching mechanism based on voltage difference is established, reducing the energy loss while ensuring the balancing efficiency. A complete balancing charging protection system is constructed. By real-time monitoring of the monomer voltage, temperature and balancing current and setting multi-level protection thresholds, the safety and reliability of the balancing charging process are ensured. The proximal policy optimization algorithm is used to update the network parameters online. By introducing the advantage function and policy entropy, the exploration efficiency of the reinforcement learning is improved and the convergence performance of the algorithm is enhanced. The temperature adaptive control of the balancing circuit is realized. By real-time monitoring of the junction temperature of the switching tube and dynamic adjustment of the duty cycle, overheating of the device is effectively prevented and the service life of the system is extended. Through the comprehensive evaluation of the voltage standard deviation, temperature distribution range and state of charge consistency index, the quantitative evaluation of the balancing effect and strategy optimization are realized. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0018] Figure 1 It is a schematic flow chart of the balancing charging method for the power battery pack provided by the embodiment of the present application;
[0019] Figure 2 It is a schematic block diagram of the structure of the balancing charging device for the power battery pack provided by the embodiment of the present application;
[0020] Figure 3 It is a schematic block diagram of the structure of the balancing charging equipment for the power battery pack provided by the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the protection scope of the present invention.
[0022] The flowcharts shown in the accompanying drawings are only illustrative examples and do not necessarily include all contents and operations / steps, nor are they necessarily executed in the described order. For example, some operations / steps can be decomposed, combined, or partially merged, so the actual execution order may change based on the actual situation.
[0023] It should also be understood that the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. As used in the specification of this application and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.
[0024] It should be further understood that the term "and / or" used in the specification of this application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0025] The following will, with reference to the accompanying drawings, elaborate on some embodiments of this application. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0026] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of the equalization charging method for the power battery pack provided by the embodiment of this application. As Figure 1 shown, the equalization charging method for the power battery pack provided by the embodiment of this application includes steps S100 to S600.
[0027] Step S100: Collect parameters of the power battery pack, measure the terminal voltage values, temperature values, and charge and discharge current values of each single battery, and calculate the state of charge values and internal resistance change rates of each single battery;
[0028] It can be understood that the execution subject of the present invention can be an equalization charging device for the power battery pack, or a terminal or a server. Specifically, it is not limited here. The embodiment of the present invention takes the server as the execution subject as an example for illustration.
[0029] Specifically, the sampling time sequence of multiple sampling points in the power battery pack is set to obtain a parameter acquisition time sequence matrix, which is used to coordinate the data acquisition of multiple sensors, enabling the comparison and analysis of data at different time points on the same time axis. Since the power battery pack is composed of multiple battery cells connected in series and parallel, and the states of each battery cell are different, it is necessary to ensure that the construction of the sampling time sequence matrix meets the requirements of balanced charging to avoid the failure of the control algorithm due to data asynchronization. During the sampling process, voltage data is collected from the sampling points according to the parameter acquisition time sequence matrix. A high-precision voltage sensor is used to measure the terminal voltage value of each single battery cell, and it is digitally processed through an ADC (analog-to-digital converter) for subsequent feature extraction and calculation. At the same time, temperature data is collected from the sampling points according to the parameter acquisition time sequence matrix. A thermistor or a temperature sensor is used to collect the temperature value of the battery under working conditions, and temperature correction is performed in combination with the battery thermal management system to ensure the accuracy of the temperature data. At the same time, current data is collected from the sampling points according to the same parameter acquisition time sequence matrix. A Hall current sensor or a shunt resistor is used to measure the charge and discharge current value, and the collected original signal is filtered to eliminate the influence of high-frequency noise on the measurement result. The charge and discharge current value is subjected to time integration operation to calculate the state of charge (SOC) of the battery. The calculation of the state of charge adopts the Coulomb integration method, that is, by integrating the charge and discharge current over time, the accumulated charge quantity is obtained, and it is superimposed with the initial state of charge value to calculate the remaining charge quantity of the battery. Combining with the open-circuit voltage characteristic of the battery for correction to improve the accuracy of SOC calculation and obtain the target value. The target value is normalized by normalizing it with the nominal capacity of the battery to obtain a relative state of charge value between 0 and 1. During the calculation of the internal resistance change rate, the difference operation of adjacent sampling points is performed on the terminal voltage value and the charge and discharge current value to obtain the voltage change rate and current change rate of the battery in a short period of time. This difference operation can reflect the dynamic characteristics of the battery, especially during pulse discharge or high-current charge and discharge, and can effectively extract the instantaneous impedance information of the battery. The difference result is subjected to Kalman filtering to reduce the measurement error and enhance the signal stability. Kalman filtering dynamically estimates the measurement value through recursive update, thereby improving the calculation accuracy of the battery internal resistance change rate. After the filtering process, the ratio of the differential signals is calculated, and the internal resistance change rate of the battery is estimated by calculating the ratio of the voltage change rate to the current change rate, reflecting the dynamic evolution of the internal impedance of the battery.
[0030] Step S200: Construct the terminal voltage value, temperature value, charge and discharge current value, state of charge value, and internal resistance change rate into a state feature sequence and input it into the multi-head self-attention network for feature extraction to generate a state correlation feature matrix;
[0031] Specifically, the terminal voltage value, temperature value, charge-discharge current value, state of charge value, and internal resistance change rate are subjected to feature splicing and dimensional transformation to obtain a state feature sequence. Before splicing, each feature is normalized to ensure that the scales of different physical quantities match and to avoid a particular feature having an excessive impact on model learning. After normalization is completed, all features are spliced in time series to form a state feature sequence, which reflects the overall state evolution of the power battery pack. A linear transformation is performed on the state feature sequence to generate a query matrix (Query), a key matrix (Key), and a value matrix (Value). A learnable weight matrix is used to project the original features into different feature spaces, enhancing the model's ability to understand the relationships between different features. By performing a linear transformation on the state feature sequence, the query matrix, key matrix, and value matrix are obtained respectively, and based on this, a matrix multiplication operation is carried out. The query matrix is dot-product calculated with the key matrix to obtain an attention score matrix. A scaling operation is performed on the attention score matrix, that is, dividing by the square root of the matrix column dimension, to prevent the occurrence of gradient vanishing or gradient explosion. After the matrix multiplication is completed, a Softmax normalization operation is performed on the attention score matrix to ensure that the sum of weights between different sampling points is 1, obtaining an attention weight matrix. The attention scores are converted into a probability distribution, enabling the model to learn the importance of each feature at different time points and dynamically adjust the weight allocation. The attention weight matrix is multiplied by the value matrix in a matrix multiplication operation to calculate the output features of single-head attention and extract the temporal dependence relationships between different state parameters of the power battery pack. The state feature sequence is input into multiple parallel attention calculation units in the multi-head self-attention network respectively. Each attention calculation unit independently performs attention feature extraction to obtain multiple different attention output features. Different attention calculation units use independent query, key, and value projection matrices, and each head focuses on different feature relationships, enabling the model to learn the state evolution patterns of the power battery pack from multiple perspectives. For each attention calculation unit, the calculation process is similar to the single-head attention mechanism, including linear transformation, matrix multiplication, scaling, Softmax normalization, and finally weighted summation. Since N attention calculation units independently calculate N different attention output features, at the final stage, a splicing operation is performed on the N attention output features to construct a more rich feature vector. After the splicing operation is completed, a normalization process of the LayerNorm layer is performed on the target feature vector to improve training stability and accelerate model convergence. The role of LayerNorm is to normalize the features within each sample, enabling the model to maintain a good learning effect under different batches of data input. The normalized target feature vector finally forms a state correlation feature matrix, which contains state information such as the terminal voltage, temperature, current, state of charge, and internal resistance change rate of each single cell in the power battery pack, and extracts the correlation relationships between these features through the multi-head self-attention mechanism.
[0032] Step S300: Construct a three-layer fully connected Actor network and a Critic network with a double Q structure based on the state correlation feature matrix, and output the balanced current value, the balancing mode of constant current or constant voltage, and the balancing time through the Actor network;
[0033] Specifically, the state-associated feature matrix is input into the first fully connected layer of the Actor network. The first fully connected network contains 512 neurons and uses the ReLU activation function for non-linear transformation to obtain the first layer of action features. The role of the ReLU function is to introduce non-linearity, enabling the model to learn complex state mapping relationships. Moreover, due to the good gradient propagation characteristics of ReLU on the positive semi-axis, it can avoid the problem of gradient vanishing and ensure the effective training of deep networks. The first layer of action features is input into the second fully connected layer of the Actor network, which contains 256 neurons and also uses the ReLU activation function for non-linear transformation. This layer further extracts more abstract action features, enabling the network to more precisely fit the complex mapping relationship of the equilibrium strategy and obtain the second layer of action features. The second layer of action features is input into the third fully connected layer of the Actor network, which contains 32 neurons and uses the Tanh activation function for non-linear transformation to generate action output features. The Tanh activation function limits the output value within the range of [−1,1], ensuring that the output of the Actor network has a reasonable numerical range. Also, because the gradient of the Tanh function is relatively large near zero, it helps to improve the learning efficiency. The action output features are input into three parallel action mapping layers to calculate the equilibrium current value, equilibrium mode, and equilibrium time respectively. Among them, the output of the first action mapping layer corresponds to the equilibrium current value, which is used to control the current regulation during the equalization charging of the power battery pack to ensure that the voltages of different battery cells tend to be consistent. The output of the second action mapping layer is used to determine the equilibrium mode, that is, to select between the constant current mode and the constant voltage mode. The constant current mode is suitable for equalizing batteries with large voltage differences, while the constant voltage mode is more suitable for fine-grained equalization control to reduce battery loss and improve charging efficiency. The output of the third action mapping layer determines the equilibrium time, that is, to control the duration of the action of the equilibrium current or voltage to achieve precise equalization and avoid over-compensation or under-compensation. At the same time, after the Actor network generates action output features, the value of this action is evaluated. The state-associated feature matrix and the action output features are combined and input into two Q networks with the same structure, namely the double Q structure of the Critic network. Each Q network contains three fully connected layers. The first fully connected layer contains 512 neurons, the second fully connected layer contains 256 neurons, and the third fully connected layer contains only one neuron and finally outputs a Q value. All layers use the ReLU activation function for non-linear transformation to ensure that the Critic network effectively fits the value function of the state-action pair. During the calculation process, each Q network independently evaluates the input state and action, calculates the first Q value and the second Q value respectively. The introduction of the double Q structure is to reduce the overestimation bias problem during the policy update process, enabling the Actor network to learn a more stable and reliable equilibrium strategy during training.Perform a minimum selection operation on the first Q value and the second Q value, that is, select the smaller value from the two Q values as the target Q value, thereby reducing the over-optimistic problem of Q value estimation and improving the stability of the equilibrium strategy. The target Q value serves as the value evaluation index for the actions output by the Actor network, guiding the training of the Actor network so that it can continuously optimize the equilibrium strategy to ensure that the selection of the equilibrium current, equilibrium mode, and equilibrium time can maximize the overall performance of the power battery pack.
[0034] Step S400: Control the DC-DC converter in the main equalization circuit to generate a constant current output according to the equalization current value, equalization mode, and equalization time, or control the Buck-Boost converter in the secondary equalization circuit to generate a constant voltage output.
[0035] Specifically, the balanced current value is compared and analyzed with the voltage difference between adjacent single cells to determine the selection of the balancing circuit. The terminal voltage value of each single cell is obtained, the voltage difference between adjacent cells is calculated, and then it is compared with a preset voltage value. If the voltage difference is greater than the preset voltage value, it indicates that there is a large voltage imbalance. At this time, the main balancing circuit is selected for balanced charging to quickly reduce the voltage deviation between battery cells. If the voltage difference is less than or equal to the preset voltage value, it means that the voltage between battery cells is close to balance. At this time, it is suitable to use the secondary balancing circuit for fine adjustment to ensure that the balancing process is more stable and efficient. Through this decision-making process, the type of the target balancing circuit is determined. A target duty cycle numerical sequence is generated according to the balancing mode. Since the operating modes of both the DC-DC converter and the Buck-Boost converter rely on PWM (pulse width modulation) signals, an appropriate duty cycle is calculated to ensure that the converter operates in the required balancing mode. For the constant current mode of the DC-DC converter, the calculation of the duty cycle is adjusted according to the input voltage and the target balanced current to maintain a constant output current. For the constant voltage mode of the Buck-Boost converter, the calculation of the duty cycle is dynamically adjusted based on the input voltage and the target output voltage to ensure the stability of the output voltage. After the target duty cycle numerical sequence is calculated, it is modulated with a PWM signal to generate the drive signal for the switching tube. When the target balancing circuit is determined to be the main balancing circuit, the generated drive signal for the switching tube is input to the main switching tube of the DC-DC converter, and the constant current output is achieved by precisely controlling the on-off timing of the main switching tube. The constant current control of the DC-DC converter adopts a current closed-loop feedback mechanism to ensure that the output current can be stably maintained at the set value. By real-time monitoring the output current and dynamically adjusting the PWM duty cycle, the influence of load changes on the balanced current is effectively compensated to ensure the stability and reliability of the balancing process. When the target balancing circuit is determined to be the secondary balancing circuit, the generated drive signal for the switching tube is input to the four switching tubes of the Buck-Boost converter, and the constant voltage output is achieved by precisely controlling the on-off timing of the four switching tubes. The constant voltage mode of the Buck-Boost converter adopts a voltage feedback control method, that is, by real-time detecting the output voltage and comparing it with the target voltage, the PWM duty cycle is adjusted so that the output voltage can be stably maintained at the set value. Since the Buck-Boost converter has the ability to step up and step down voltages, a stable output voltage can be achieved under different input voltage conditions, making it more flexible to adapt to the states of different battery cells during the balanced charging process. During the balancing process, the junction temperature of the switching tube in the target balancing circuit is detected in real time to prevent the switching tube from being damaged due to overheating. A thermistor or an infrared temperature sensor is used to monitor the junction temperature of the switching tube, and the temperature data is input into the control system for real-time analysis.When it is detected that the junction temperature of the switching transistor exceeds the target temperature, derating calculation is performed on the target duty cycle numerical sequence according to the temperature overrun amplitude to reduce the working load of the switching transistor and reduce heat loss, and the first duty cycle numerical sequence is obtained. The derating calculation method adopts linear attenuation or exponential attenuation method, that is, as the junction temperature rises, the PWM duty cycle is gradually reduced to reduce power loss and lower the temperature of the switching transistor, so as to ensure the safe operation of the equalization circuit. The first duty cycle numerical sequence is input into the steady-state duty cycle compensator to perform feedback compensation on the derated duty cycle value. The steady-state duty cycle compensator fine-tunes the duty cycle according to historical data and real-time working conditions to avoid the decline of equalization performance caused by derating calculation. The compensator adopts PID (Proportional-Integral-Derivative) control algorithm or adaptive control algorithm based on neural network to ensure the accuracy and stability of duty cycle adjustment. After the compensation is completed, the second duty cycle numerical sequence is finally obtained and used to update the switching transistor drive signal, so as to optimize the equalization charging process, enabling it to maximize the equalization efficiency and battery life while ensuring safety.
[0036] Perform standard deviation operation on the terminal voltage value to calculate the voltage standard deviation, which is used to measure the voltage consistency of each single battery in the power battery pack. A larger standard deviation indicates a large voltage deviation within the battery pack, so a more aggressive balancing strategy is needed to reduce this unbalanced state. At the same time, perform maximum and minimum operations on the temperature value to calculate the temperature distribution range, which is used to measure the uniformity of the internal temperature distribution of the power battery pack. If the temperature distribution range is large, there is a problem of local overheating, which will affect the battery performance and even pose a safety hazard. During the balanced charging process, the charging strategy needs to be appropriately adjusted to reduce the temperature difference inside the battery pack. Calculate the root mean square deviation of the state of charge value to obtain the consistency index, which is used to evaluate the overall uniformity of the battery pack's SOC (state of charge). The imbalance of SOC will cause some batteries to be overcharged or over-discharged, thus affecting the service life of the battery. After completing the above index calculations, weight-combine the voltage standard deviation, temperature distribution range, and consistency index to generate a reward function value, which is used to guide the Actor-Critic network to optimize the balanced charging strategy. In the reinforcement learning framework, the reward function is the core of the optimization goal, and its design needs to ensure that the charging strategy develops in the direction of reducing the imbalance of the battery pack, improving energy utilization efficiency, and extending the battery life. After calculating the reward function value, perform a difference operation between it and the target reward value to obtain the advantage function value, which is used to measure the improvement space of the current charging strategy compared to the optimal strategy. The larger the advantage function value, the closer the current strategy is to the optimal solution. Otherwise, the balanced charging strategy needs to be further optimized. Calculate the policy gradient based on the advantage function value and perform truncation processing on the policy gradient to prevent gradient explosion or gradient disappearance problems. The purpose of policy gradient truncation is to limit the amplitude of gradient change, making the training process more stable and avoiding the problem that the update amplitude of the Actor network is too large due to extreme states, which affects the convergence. After obtaining the truncated policy gradient, calculate the policy entropy value according to the logarithmic probability density of the output action of the Actor network in the action space. The policy entropy value is used to measure the exploratory nature of the policy. A larger policy entropy means that the balanced charging strategy still maintains a certain degree of randomness in different states, thus exploring a more optimal balanced scheme, while a smaller policy entropy indicates that the balanced charging strategy tends to converge and is more inclined to use the learned balanced mode. During the optimization process, comprehensively consider the policy gradient and policy entropy to achieve a balance between exploration and exploitation. Perform weighted superposition on the truncated policy gradient and policy entropy value, and update the weight parameters of the Actor network according to the set learning rate to obtain the updated weight parameters of the Actor network. The update of the Actor network can adjust the output of the balancing current, balancing mode, and balancing time, making the balancing strategy more accurately adapt to different battery states, thus improving the effect of balanced charging.While optimizing the Actor network, calculate the temporal difference error according to the advantage function value, and calculate the mean squared error of the temporal difference error to obtain the value loss function. The calculation of the value loss function is the key to the update of the Critic network, and its goal is to minimize the error between the predicted value and the actual reward, thereby improving the value evaluation ability of the Critic network for the balanced charging strategy. Based on the value loss function, update the weight parameters of the Critic network according to the learning rate to obtain the updated weight parameters of the Critic network. According to the updated weight parameters of the Actor network and the updated weight parameters of the Critic network, correct the balanced current value, the balancing mode, and the balancing time to obtain an optimized balanced charging strategy. Through continuous iterative optimization, the balanced charging strategy is dynamically adjusted under different working conditions to minimize the imbalance state inside the power battery pack to the greatest extent, improve the energy utilization efficiency, and extend the service life of the battery pack.
[0037] Data collection is carried out on the single - cell voltage, temperature, and equalization current in the main equalization circuit and the secondary equalization circuit, and a single - cell monitoring data matrix is constructed to reflect the operating state of the power battery pack during the equalization charging process. During the data collection process, the terminal voltage of each single - cell battery is measured in real - time through a high - precision voltage sensor, and its variation over time is recorded. At the same time, a temperature sensor is used to obtain the temperature data of each battery cell to monitor temperature abnormalities during charging. The equalization current in the equalization circuit is measured, and high - precision current data is obtained through a Hall sensor or a shunt resistor to ensure the stability and safety of the equalization process. After the construction of the single - cell monitoring data matrix is completed, a comparison operation is performed between the single - cell voltage and the preset voltage range to determine whether there is a voltage abnormality in the battery cell. When the voltage of a certain battery cell exceeds the preset range, a voltage over - limit flag is generated to take appropriate protection measures for this battery cell in the subsequent control process. The temperature difference between adjacent single - cells is calculated and compared with the preset temperature difference threshold to identify local overheating phenomena. If the temperature difference between a certain battery cell and its adjacent single - cell exceeds the set threshold, a temperature - difference over - limit flag is generated to ensure that no safety hazards are caused by overheating during the charging process. At the same time, the equalization current is compared with the current threshold of the rated current to determine whether the equalization current exceeds the safe range, and a current over - limit flag is generated accordingly. An over - limit level matrix is constructed based on the voltage over - limit flag, the temperature - difference over - limit flag, and the current over - limit flag, and priority sorting is performed on this matrix to generate corresponding protection action instructions. During the generation of the protection action instructions, weighted calculations are performed on different types of over - limit situations to determine the severity of various abnormalities and set corresponding processing strategies accordingly. For example, if the voltage of a certain battery cell deviates from the normal range but does not exceed the safety limit, priority is given to reducing the equalization current to alleviate the voltage difference. If the battery temperature exceeds the set safety limit, the equalization charging is immediately interrupted to prevent overheating from damaging the battery. When the protection action instruction is to reduce the equalization current, a decrement calculation is performed on the current equalization current value to obtain the derated equalization current value, and it is updated to the main equalization circuit or the secondary equalization circuit to ensure that the equalization charging process can be carried out within a safe range. The derating calculation of the equalization current adopts a linear decrement or exponential decrement strategy, that is, the current is gradually reduced while maintaining the equalization effect to reduce the excessive voltage difference between battery cells and prevent over - current from damaging the battery. During the adjustment of the equalization current, dynamic adjustment is carried out in combination with the overall state of the battery pack to ensure that the equalization process can maintain the equalization effect without causing additional negative impacts on the battery pack. When the protection action instruction is to interrupt the equalization charging, an equalization interruption signal sequence is immediately generated, and this signal sequence is respectively input to the control terminals of the switching tubes in the main equalization circuit and the secondary equalization circuit to make the corresponding switching tubes enter the cut - off state, thereby stopping the output of the equalization current.The purpose of balance interruption is to prevent abnormal conditions from deteriorating continuously. For example, when the battery temperature reaches the critical point or the voltage of a single battery deviates significantly from the normal value, interrupting the balance charging can effectively avoid further imbalance and provide time for subsequent fault recovery. During the balance interruption process, ensure the rapid response of the control signal to cut off the balance circuit in a timely manner and prevent the abnormal state from causing irreversible damage to the battery pack. Conduct feedback detection on the derated balance current value or the execution status of the balance interruption signal sequence to ensure that the protection measures are correctly implemented. The content of the feedback detection includes whether the balance current decreases as expected, whether the switching tubes in the balance circuit successfully enter the cut-off state, and whether the voltage and temperature of the battery cells return to the safe range. Through these feedback data, optimize the balance charging strategy and adjust the protection measures when necessary to ensure the stability and safety of the charging process. After completing the protection action, record the protection execution result in the balance charging log for subsequent analysis and maintenance. The recorded content of the balance charging log includes the voltage and temperature change trends of the battery cells, the adjustment history of the balance current, the occurrence time of the abnormal state, and the handling measures, etc. By analyzing the balance charging log, optimize the balance strategy, improve the service life of the power battery pack, and reduce the safety risks during the charging process.
[0038] In the embodiment of the present invention, by introducing a multi-head self-attention network for feature extraction, the efficient processing of the state information of the power battery pack is realized, the mutual correlation between different monomers is accurately captured, and the expression ability of the state features is improved. The Actor-Critic deep reinforcement learning framework is adopted to dynamically generate the combined strategy of the balance current, balance mode, and balance time, realizing the adaptive adjustment of the balance control, so that the balance strategy can be optimized according to the real-time state of the battery pack. A multi-level balance circuit structure in which the main balance circuit and the secondary balance circuit work together is designed, and a balance mode switching mechanism based on the voltage difference is established, reducing the energy loss while ensuring the balance efficiency. A complete balance charging protection system is constructed. By real-time monitoring the monomer voltage, temperature, and balance current and setting multi-level protection thresholds, the safety and reliability of the balance charging process are ensured. The proximal policy optimization algorithm is used to update the network parameters online. By introducing the advantage function and policy entropy, the exploration efficiency of the reinforcement learning is improved, and the convergence performance of the algorithm is enhanced. The temperature adaptive control of the balance circuit is realized. By real-time monitoring the junction temperature of the switching tube and dynamically adjusting the duty cycle, the device overheating is effectively prevented, and the service life of the system is extended. Through the comprehensive evaluation of the voltage standard deviation, temperature distribution range, and state of charge consistency index, the quantitative evaluation of the balance effect and the strategy optimization are realized.
[0039] In a specific embodiment, the process of executing step S100 may specifically include the following steps:
[0040] Set the sampling time sequence for multiple sampling points in the power battery pack to obtain a parameter acquisition time sequence matrix;
[0041] Collect voltage data for the sampling points according to the parameter acquisition time sequence matrix to obtain the terminal voltage value, collect temperature data for the sampling points according to the parameter acquisition time sequence matrix to obtain the temperature value, and collect charge and discharge current data for the sampling points according to the parameter acquisition time sequence matrix to obtain the charge and discharge current value;
[0042] Perform a time integral operation on the charge and discharge current value and superimpose it with the initial state of charge value to obtain a target value, and perform a normalization process on the target value and the nominal capacity of the battery to obtain the state of charge value;
[0043] Perform a differential operation on the terminal voltage value and the charge and discharge current value for adjacent sampling points to obtain a differential result, and perform a Kalman filter process and a ratio calculation on the differential result to obtain the internal resistance change rate.
[0044] Specifically, define a reasonable sampling time sequence to ensure that the measurement of the voltage, temperature, and charge and discharge current of each single battery is carried out under a unified time reference. Assume that the power battery pack consists of single batteries, and set the sampling period to . Then, within each sampling period, measure the voltage, temperature, and current of all single batteries in sequence, and arrange the collected data in chronological order to form a parameter acquisition time sequence matrix. Let the voltage, temperature, and current of the th battery in the th sampling period be , and respectively, then the parameter acquisition time sequence matrix is expressed as:
[0045]
[0046] Among them, represents the measurement data matrix of all single batteries in the th sampling period. Perform a time integral operation on the charge and discharge current value and superimpose it with the initial state of charge value to calculate the state of charge of the power battery. The calculation of SOC is based on the Coulomb integration method, and the mathematical expression is
[0047]
[0048] Among them, represents the state of charge in the th sampling period; is the initial state of charge, provided by the manufacturer or estimated according to the open-circuit voltage; represents the nominal capacity of the battery; is the The battery current within one sampling period; represents the sampling period. After calculating the state of charge, differential operations are performed on the terminal voltage value and the charge and discharge current value at adjacent sampling points to calculate the internal resistance change rate. Assume that the th single battery at adjacent sampling periods and The voltages and currents are respectively , , and , then the differential calculation is as follows:
[0049]
[0050]
[0051] After completing the differential calculation, the internal resistance is calculated by Ohm's law:
[0052]
[0053] Kalman filtering is performed on the voltage difference and the current difference to improve the measurement accuracy. The update equation of the Kalman filter is as follows:
[0054]
[0055] Among them, represents the estimated value of the internal resistance after filtering; is the Kalman gain, which is used to adjust the sensitivity of the estimated value to the new measurement value; is the internal resistance value calculated currently. After filtering, a more stable internal resistance change rate is obtained for evaluating the battery health state and optimizing the equalization strategy.
[0056] In a specific embodiment, the process of executing step S200 may specifically include the following steps:
[0057] Perform feature splicing and dimensional transformation on the terminal voltage value, temperature value, charge and discharge current value, state of charge value, and internal resistance change rate to obtain a state feature sequence;
[0058] Perform a linear transformation on the state feature sequence to generate a query matrix, a key matrix, and a value matrix, and perform a matrix multiplication operation on the query matrix and the key matrix to obtain an attention score matrix;
[0059] Perform a Softmax normalization operation on the attention score matrix to obtain an attention weight matrix, and perform a matrix multiplication operation on the attention weight matrix and the value matrix to obtain a single-head attention output feature;
[0060] The state feature sequence is respectively input into N parallel attention calculation units in the multi-head self-attention network, and each attention calculation unit independently performs attention feature extraction to obtain N attention output features;
[0061] Perform a concatenation operation on the N attention output features to obtain a target feature vector, and perform normalization processing on the target feature vector through a LayerNorm layer to obtain a state correlation feature matrix.
[0062] Specifically, perform feature concatenation and dimensionality transformation on the terminal voltage value, temperature value, charge and discharge current value, state of charge value, and internal resistance change rate. Assume that the power battery pack contains single cells, and at the th sampling moment, the terminal voltage, temperature, charge and discharge current, state of charge, and internal resistance change rate of each cell are respectively represented as 、 、 、 and , and the state feature matrix of the entire battery pack is represented as:
[0063]
[0064] Among them, The rows represent different single cells, and the columns represent different features. Perform feature concatenation and dimensionality transformation on the state feature matrix . Assume that the feature vector represents the state of the th cell:
[0065]
[0066] The state feature sequence of the entire battery pack is represented as:
[0067]
[0068] Among them The dimension of is , among which represents the number of features. Perform a linear transformation on the state feature sequence to generate a query matrix (Query), a key matrix (Key), and a value matrix (Value). Let 、 and be trainable weight matrices, which are respectively used to project the feature sequence to calculate the query, key, and value:
[0069]
[0070] Among them, is the query matrix, representing the importance evaluation of the input features; is the key matrix, representing the similarity between input features; is the value matrix, representing the information of the input features themselves. Calculate the attention score matrix , the calculation method of which is to perform a dot product operation on the query matrix and the key matrix and perform scaling:
[0071]
[0072] where is the feature dimension, and the scaling factor is used to stabilize the gradient and improve the stability of numerical calculations. Perform the Softmax normalization operation on the attention score matrix to generate the attention weight matrix :
[0073]
[0074] where the role of Softmax normalization is to convert the attention scores into a probability distribution, so that the weights of each input feature sum to 1, thus ensuring that the attention mechanism can effectively capture the importance of the input features. Multiply the attention weight matrix with the value matrix to calculate the output features of single-head attention:
[0075]
[0076] where represents the output of the state features after attention weighting. Input the state feature sequence into parallel attention calculation units in the multi-head self-attention network respectively, and each unit independently performs attention feature extraction. Let the query, key, and value projection matrices corresponding to the th attention head be , and respectively, and their calculation methods are:
[0077]
[0078] Calculate the weights of each attention head:
[0079]
[0080] and calculate the attention output of this head:
[0081]
[0082] All After the calculations of all attention heads are completed, their output features are concatenated:
[0083]
[0084] Among them is the output of the multi-head attention mechanism, and its dimension is . For the concatenated target feature vector perform LayerNorm normalization to improve training stability:
[0085]
[0086] Among them, the function of LayerNorm normalization is to normalize the features within each sample, so that the model can maintain a good learning effect under the input of data in different batches. Obtain the state correlation feature matrix , which contains state information such as the voltage, temperature, current, state of charge, and internal resistance change rate of each single battery in the power battery pack, and extracts the correlation relationships between these features through the multi-head self-attention mechanism.
[0087] In a specific embodiment, the process of executing step S300 may specifically include the following steps:
[0088] Input the state correlation feature matrix into the first fully connected layer of the Actor network. The first fully connected layer contains 512 neurons and performs a non-linear transformation through the ReLU activation function to obtain the first layer of action features;
[0089] Input the first layer of action features into the second fully connected layer of the Actor network. The second fully connected layer contains 256 neurons and performs a non-linear transformation through the ReLU activation function to obtain the second layer of action features;
[0090] Input the second layer of action features into the third fully connected layer of the Actor network. The third fully connected layer contains 32 neurons and performs a non-linear transformation through the Tanh activation function to obtain the action output features;
[0091] Input the action output features into three parallel action mapping layers respectively. The first action mapping layer outputs the balanced current value, the second action mapping layer outputs the balanced mode of constant current mode or constant voltage mode, and the third action mapping layer outputs the balanced time;
[0092] Input the state correlation feature matrix and the action output features into two Q networks with the same structure. Each Q network contains three fully connected layers, and the number of neurons in the three fully connected layers is 512, 256, and 1 respectively, to obtain the first Q value and the second Q value;
[0093] Perform a minimum selection operation on the first Q value and the second Q value to obtain the target Q value, and use the target Q value as the value evaluation index for the actions output by the Actor network.
[0094] Specifically, define a state-related feature matrix, which contains the key state parameters of the power battery pack, such as terminal voltage, temperature, charge and discharge current, state of charge, and internal resistance change rate. Assume that the power battery pack contains single cells, and at the th sampling moment, the state characteristics of each cell are represented as , where:
[0095]
[0096] Among them, is the terminal voltage of the th single cell, is the temperature, is the charge and discharge current, is the state of charge, is the internal resistance change rate. The state characteristic matrix of the entire power battery pack is represented as:
[0097]
[0098] This matrix serves as the input to the Actor network and enters the first fully connected layer, which contains 512 neurons and undergoes a non-linear transformation through the ReLU activation function. Let the input matrix be transformed by the first layer weight matrix and the bias , and calculate the first layer action features:
[0099]
[0100] Among them, is the output feature matrix of the first layer, is the first layer weight matrix, is the bias term. The role of the ReLU activation function is to introduce non-linearity and enable the model to more effectively learn the complex state-action mapping relationship. Input the first layer action features into the second fully connected layer, which contains 256 neurons and also uses the ReLU activation function for non-linear transformation. The calculation method is as follows:
[0101]
[0102] Among them, and They are the weight matrix and bias term of the second layer respectively. The main function of this layer is to further extract more abstract action features, enabling the Actor network to more precisely fit the complex mapping relationship of the equilibrium strategy. The action features of the second layer are input into the third fully connected layer, which contains 32 neurons and uses the Tanh activation function for non-linear transformation to generate the final action output features:
[0103]
[0104] Among them, the choice of the Tanh activation function is because its output range is between [-1, 1], making the output of the Actor network have better numerical stability, and Tanh has a large gradient near zero, which helps to improve the learning efficiency. After generating the action output features, they are input into three parallel action mapping layers to calculate the equilibrium current value, equilibrium mode, and equilibrium time respectively. Among them, the output of the first action mapping layer corresponds to the equilibrium current value:
[0105]
[0106] Among them, represents the equilibrium current value, and are the weights and biases of the mapping layer respectively. The output of the second action mapping layer is used to determine the equilibrium mode, that is, to select between the constant current mode and the constant voltage mode:
[0107]
[0108] Among them, represents the probability value of the equilibrium mode, represents the Sigmoid function, which is used to limit the output value between (0, 1) so as to judge whether it is the constant current mode or the constant voltage mode according to the threshold. The output of the third action mapping layer is used to determine the equilibrium time:
[0109]
[0110] Among them, represents the equilibrium time. The value of the action generated by the Actor network is evaluated, and the state-related feature matrix and the action output feature are combined and then input into two Q networks in the Critic network. Let the combined input matrix be:
[0111]
[0112] The matrix is input into two Q - networks with the same structure respectively. Each Q - network contains three fully - connected layers, and the number of neurons is 512, 256, and 1 respectively. For the first Q - network, its calculation process is as follows:
[0113]
[0114]
[0115]
[0116] Among them, is the output value of the first Q - network. Similarly, the calculation process of the second Q - network is as follows:
[0117]
[0118]
[0119]
[0120] Among them, is the output value of the second Q - network. After obtaining the first Q - value and the second Q - value, a minimum - selection operation is performed to reduce the problem of over - optimistic policies:
[0121]
[0122] Among them, is used as the final target Q - value to evaluate the value of the actions output by the Actor network. The Actor network continuously optimizes the equilibrium strategy according to the value evaluation of the Critic network, enabling the equilibrium charging process to dynamically adapt to different battery states and improving the efficiency and reliability of equilibrium charging.
[0123] In a specific embodiment, the process of executing step S400 may specifically include the following steps:
[0124] Compare and analyze the equilibrium current value with the voltage difference of adjacent single - cell batteries. When the voltage difference is greater than the preset voltage value, select the main equilibrium circuit for equilibrium charging; when the voltage difference is less than or equal to the preset voltage value, select the secondary equilibrium circuit for equilibrium charging to obtain the target equilibrium circuit;
[0125] Generate a target duty - cycle numerical sequence according to the type of the target equilibrium circuit and the equilibrium mode, and perform pulse - width modulation on the target duty - cycle numerical sequence to obtain the switching - tube drive signal;
[0126] When the target equilibrium circuit is the main equilibrium circuit, input the switching - tube drive signal into the main switching tube of the DC - DC converter to control the on - off timing of the main switching tube to obtain a constant - current output;
[0127] When the target equalization circuit is the secondary equalization circuit, input the switching tube drive signal to the four switching tubes of the Buck-Boost converter, control the on-off timing of the four switching tubes, and obtain a constant voltage output;
[0128] Real-time detect the junction temperature of the switching tubes in the target equalization circuit. When the junction temperature of the switching tubes exceeds the target temperature, perform derating calculation on the target duty ratio numerical sequence according to the temperature overrun amplitude to obtain the first duty ratio numerical sequence;
[0129] Input the first duty ratio numerical sequence into the steady-state duty ratio compensator, perform feedback compensation on the first duty ratio numerical sequence to obtain the second duty ratio numerical sequence, and update the switching tube drive signal.
[0130] Specifically, calculate the voltage difference between adjacent single cells and compare it with a preset voltage threshold. Assume that the power battery pack contains single cells, and at the sampling moment, the terminal voltage of each single cell is , then the voltage difference between adjacent single cells is calculated as follows:
[0131]
[0132] Among them, represents the voltage difference between the th and the th battery. Assume that the preset voltage threshold is , if , it indicates that the voltage imbalance between the batteries is relatively large, and select the main equalization circuit for equalization charging; if , it indicates that the voltage deviation is relatively small, and at this time select the secondary equalization circuit for fine-tuning equalization. The selection rule of the target equalization circuit is expressed as:
[0133]
[0134] Among them, represents the type of equalization circuit selected at the sampling moment. When the type of equalization circuit is determined, generate the target duty ratio numerical sequence according to the equalization mode (constant current mode or constant voltage mode), and generate a pulse width modulation (PWM) signal based on this to control the output of the equalization current or equalization voltage. For the main equalization circuit, adopt the DC-DC converter to output the constant current mode, and the calculation method of the duty ratio is:
[0135]
[0136] Among them, is the input voltage of the converter, is the target balanced voltage. For the secondary balancing circuit, the Buck-Boost converter is adopted to output in the constant voltage mode, and the duty cycle is calculated as follows:
[0137]
[0138] The calculated target duty cycle is converted into the switching tube drive signal by adopting the PWM modulation mode :
[0139]
[0140] wherein, is the PWM period, is the target duty cycle within the current sampling period. When the target balancing circuit is the main balancing circuit, the switching tube drive signal is input into the main switching tube of the DC-DC converter to control its on-off timing and realize the constant current output. The on-off behavior of the main switching tube is determined by the PWM signal, so that the balancing current is maintained at the set value:
[0141]
[0142] wherein, is the equivalent resistance of the balancing circuit. When the target balancing circuit is the secondary balancing circuit, the four switching tubes of the Buck-Boost converter are simultaneously controlled to switch synchronously according to the PWM signal to keep the balanced output voltage stable. The switching logic of the four switching tubes is based on the duty cycle sequence , so that the balanced voltage output by the Buck-Boost converter meets the target requirements:
[0143]
[0144] During the balancing process, the junction temperature of the switching tube in the target balancing circuit is detected in real time to prevent device damage caused by over-temperature. Suppose the current junction temperature of a certain switching tube is , and the set temperature threshold is , then when , the derating calculation is carried out on the target duty cycle to reduce the power loss. The derating calculation method adopts linear attenuation:
[0145]
[0146] wherein, is the maximum allowable temperature, is the derated duty cycle value. The derated duty cycle An input steady-state duty cycle compensator is used to dynamically adjust it to ensure that the system can still operate stably under different temperature conditions. The compensator performs feedback regulation based on the PID control algorithm:
[0147] ;
[0148] Among them, are the proportional, integral, and differential coefficients of the PID control respectively, is the duty cycle of the previous cycle. The compensated duty cycle is used to update the switching tube drive signal to ensure that the equalization circuit works properly within the temperature-controlled range and maintains the best equalization charging effect.
[0149] In this embodiment, controlling the DC-DC converter in the main equalization circuit to generate a constant current output or controlling the Buck-Boost converter in the secondary equalization circuit to generate a constant voltage output further includes: calculating the difference between the actual equalization current and the target equalization current of the main equalization circuit and the secondary equalization circuit to obtain an equalization current deviation sequence, constructing a Markov random jump model based on the equalization current deviation sequence to obtain a fault feature matrix; constructing a generalized martingale measure based on the fault feature matrix, performing a random walk analysis on the equalization current deviation sequence to obtain the probability of fault occurrence, and classifying the equalization circuit according to the probability of fault occurrence to obtain a fault level index; inputting the fault level index into a fault-tolerant equivalent controller to generate a fault compensation control quantity, and the fault-tolerant equivalent controller adopts a three-layer neural network structure, including an input layer, a hidden layer, and an output layer, and the hidden layer uses a radial basis function as an activation function; performing a combined operation on the fault compensation control quantity and the equalization current value to obtain a corrected equalization current value, and regenerating a PWM control signal according to the corrected equalization current value to obtain a compensated switching tube control signal; performing an online evaluation on the control effect of the compensated switching tube control signal, calculating the probability distribution difference between the actual equalization current and the target equalization current to obtain a control performance index, and adaptively updating the network parameters of the fault-tolerant equivalent controller according to the control performance index; constructing a probability limit constraint condition, calculating a fault influence coefficient according to the equalization current deviation sequence, comparing the fault influence coefficient with a preset threshold, and triggering a standby equalization circuit when it exceeds the preset threshold; dynamically adjusting the compensation strategy of the fault-tolerant equivalent controller based on the working state of the standby equalization circuit to generate a new fault compensation control quantity, and feeding back the new fault compensation control quantity to the equalization control system; recording the compensated switching tube control signal, the fault feature matrix, and the control performance index in a fault diagnosis database for subsequent fault mode analysis and control strategy optimization.
[0150] In a specific embodiment, the process of executing step S500 may specifically include the following steps:
[0151] Perform standard deviation operation on the terminal voltage value to obtain the voltage standard deviation, perform maximum and minimum operations on the temperature value to obtain the temperature distribution range, and perform root mean square deviation calculation on the state of charge value to obtain the consistency index;
[0152] Perform weighted combination on the voltage standard deviation, temperature distribution range and consistency index to generate the reward function value, and perform difference operation on the reward function value in the current state and the target reward value to obtain the advantage function value;
[0153] Calculate the policy gradient based on the advantage function value, perform truncation processing on the policy gradient to obtain the truncated policy gradient, and calculate the policy entropy value according to the logarithmic probability density of the output action of the Actor network in the action space;
[0154] Perform weighted superposition on the truncated policy gradient and the policy entropy value, and update the weight parameters of the Actor network according to the learning rate to obtain the updated weight parameters of the Actor network;
[0155] Calculate the temporal difference error based on the advantage function value, perform mean square error calculation on the temporal difference error to obtain the value loss function, and update the weight parameters of the Critic network according to the learning rate to obtain the updated weight parameters of the Critic network;
[0156] According to the updated weight parameters of the Actor network and the updated weight parameters of the Critic network, correct the equalization current value, equalization mode and equalization time to obtain the optimized equalization charging strategy.
[0157] Specifically, calculate the standard deviation of the terminal voltage value to measure the voltage consistency within the battery pack. Suppose the power battery pack contains individual battery cells, and within the sampling period, the terminal voltage of each individual battery cell is , then the voltage standard deviation is calculated as follows:
[0158]
[0159] Among them, is the average voltage of the battery pack:
[0160]
[0161] The larger the voltage standard deviation, the more serious the imbalance state of the battery pack, and a stronger equalization strategy is required to adjust the voltage difference of the individual battery cells. Calculate the temperature distribution range to measure the temperature non-uniformity inside the battery pack. Suppose the temperature of the th individual battery cell at time is , then the temperature distribution range The calculation is as follows:
[0162]
[0163] The greater the range of the temperature distribution, the more serious the local overheating problem exists inside the battery pack. At this time, the equalization charging strategy needs to be adjusted to reduce the heat accumulation in the high-temperature area. To evaluate the degree of equalization of the state of charge (SOC), the consistency index of SOC is calculated, and the root mean square deviation is used to measure the consistency degree of SOC. Let the SOC of each single battery be , then the SOC consistency index is calculated as follows:
[0164]
[0165] where is the average SOC:
[0166]
[0167] The smaller the SOC consistency index, the more balanced the SOC state in the battery pack, and the more reasonable the energy distribution during the charge and discharge process. The standard deviation of voltage, the range of temperature distribution, and the SOC consistency index are weighted and combined to construct a reward function for equalization charging. Let the reward function be the weighted sum of these three indicators:
[0168]
[0169] where are the weight coefficients, which respectively measure the relative importance of voltage equalization, temperature equalization, and SOC equalization. The goal of the reward function is to minimize the degree of imbalance, so a negative sign is taken, so that the better the equalization state, the larger the reward value . Calculate the advantage function value , which is used to measure the improvement degree of the current strategy relative to the benchmark strategy. Let the target reward value be , then the advantage function is calculated as follows:
[0170]
[0171] If is positive, it means that the current strategy is better than the benchmark strategy and needs to be strengthened; if is negative, it means that the current strategy is not good and needs to be adjusted. Based on the advantage function, calculate the policy gradient and perform truncation processing on the policy gradient to prevent the problem of gradient explosion. Let the policy gradient of the Actor network be :
[0172]
[0173] Among them, is the action probability density output by the policy network, and are the parameters of the Actor network. To prevent the gradient from being too large and causing unstable training, the policy gradient is truncated:
[0174]
[0175] Among them, is the gradient truncation threshold. Calculate the policy entropy of the Actor network to encourage the exploration of the policy. The policy entropy is calculated as follows:
[0176]
[0177] The larger the policy entropy, the more exploratory the Actor network has in the balanced charging strategy, which helps to improve the robustness of the charging strategy. For the truncated policy gradient and the policy entropy perform weighted superposition and update the parameters of the Actor network according to the learning rate :
[0178]
[0179] Among them, is the policy entropy weight, which is used to control the influence of exploration on policy optimization. At the same time, the Critic network needs to be updated to optimize the state value estimation. The goal of the Critic network is to minimize the temporal difference error (TD error), and the calculation method is as follows:
[0180]
[0181] Among them, is the state value estimated by the Critic network, is the discount factor. The loss function of the Critic network is calculated using the mean square error:
[0182]
[0183] Update the parameters of the Critic network according to the learning rate :
[0184]
[0185] After updating the parameters of the Actor and Critic networks, the equalization current, equalization mode, and equalization time are corrected according to the output of the optimized Actor network to optimize the equalization charging strategy. The corrected equalization charging strategy dynamically adjusts the equalization charging parameters according to the current battery state, thereby improving the equalization of the charging process, reducing the imbalance degree inside the battery pack, increasing the battery life, and optimizing the energy utilization efficiency.
[0186] In a specific embodiment, the process of executing step S600 may specifically include the following steps:
[0187] Collect data on the cell voltage, temperature, and equalization current in the main equalization circuit and the secondary equalization circuit to obtain a cell monitoring data matrix;
[0188] Perform a comparison operation on the cell voltage in the cell monitoring data matrix with a preset voltage range to obtain a voltage overlimit flag, perform a comparison operation on the temperature difference between adjacent cells with a preset temperature difference threshold to obtain a temperature difference overlimit flag, and perform a comparison operation on the equalization current with a current threshold of the rated current to obtain a current overlimit flag;
[0189] Generate an overlimit level matrix according to the voltage overlimit flag, temperature difference overlimit flag, and current overlimit flag, and perform a priority sorting on the overlimit level matrix to obtain a protection action instruction;
[0190] When the protection action instruction is to reduce the equalization current, perform a decrement calculation on the equalization current value to obtain a derated equalization current value, and update the derated equalization current value to the main equalization circuit or the secondary equalization circuit;
[0191] When the protection action instruction is to interrupt the equalization charging, generate an equalization interruption signal sequence, and input the equalization interruption signal sequence into the switch control terminals of the main equalization circuit and the secondary equalization circuit respectively to make the switch tubes enter the cut-off state;
[0192] Perform a feedback detection on the execution status of the derated equalization current value or the equalization interruption signal sequence to obtain a protection execution result, and record the protection execution result in the equalization charging log.
[0193] Specifically, high-precision sensors are arranged inside the power battery pack to monitor the working status of each single cell in real time. Assume that the power battery pack contains single cells, and at the sampling moment , the terminal voltage, temperature, and equalization current of each single cell are respectively represented as , and , where . Through the data acquisition system, a cell monitoring data matrix is constructed:
[0194]
[0195] Among them, each row corresponds to the monitoring data of a single battery cell, and each column represents different measurement parameters. For the cell voltage in the single-cell monitoring data matrix and the preset voltage range a comparison operation is performed to determine whether there is a situation where the voltage exceeds the limit. For each single battery cell, if exceeds the range:
[0196]
[0197] then mark the voltage over-limit status of this battery:
[0198]
[0199] Otherwise:
[0200]
[0201] Among them, represents the voltage over-limit identifier of the th battery. Calculate the temperature difference between adjacent single battery cells and compare it with the preset temperature difference threshold as follows:
[0202]
[0203] If:
[0204]
[0205] then mark the temperature difference over-limit status of this battery:
[0206]
[0207] Otherwise:
[0208]
[0209] Among them, represents the temperature difference over-limit identifier of the th battery. Similarly, compare the balancing current with the rated current to determine whether the balancing current exceeds the limit:
[0210]
[0211] If the balancing current exceeds the threshold, then:
[0212]
[0213] Otherwise:
[0214]
[0215] After calculating the voltage over-limit flag, temperature difference over-limit flag, and current over-limit flag, combine them into an over-limit level matrix:
[0216]
[0217] Perform a priority sorting on the over-limit level matrix to determine the trigger order of the protection action instructions. According to the sorting result, obtain the corresponding protection action instructions. If the protection action instruction requires reducing the equalizing current, then perform a decreasing calculation on the current equalizing current as follows:
[0218]
[0219] where is the derating ratio coefficient, and its value is within . Update the derated equalizing current to the main equalizing circuit or the secondary equalizing circuit. If the protection action instruction requires interrupting the equalizing charge, generate an equalizing interruption signal sequence and input it to the control terminals of the switching transistors in the main equalizing circuit and the secondary equalizing circuit to make the switching transistors enter the cut-off state:
[0220]
[0221] where represents the switching transistor state of the th battery, taking 0 to indicate off and 1 to indicate on. After performing the protection action, perform a feedback detection on the derated equalizing current value or the execution status of the equalizing interruption signal sequence to confirm whether the protection measure is effective. If the feedback detection shows that the equalizing current has been reduced to the set range or the equalizing circuit has been interrupted, record the successful execution:
[0222]
[0223] Otherwise, record the failed execution:
[0224]
[0225] Record the protection execution result in the equalizing charge log for subsequent analysis and optimization of the equalizing strategy. The equalizing charge log records are as follows:
[0226]
[0227] where records the monitoring data, over-limit status, adjusted equalizing current, switching transistor state, and execution result at the current moment.
[0228] Please refer to Figure 2 ,Figure 2 Schematic block diagram of the balanced charging device 200 for the power battery pack provided by the embodiment of the present application, as Figure 2 shown, the balanced charging device 200 for the power battery pack includes:
[0229] A parameter acquisition module 210, configured to acquire parameters of the power battery pack, measure the terminal voltage value, temperature value and charge-discharge current value of each single battery, and calculate the state of charge value and internal resistance change rate of each single battery;
[0230] A feature extraction module 220, configured to construct a state feature sequence from the terminal voltage value, temperature value, charge-discharge current value, state of charge value and internal resistance change rate, and input it into a multi-head self-attention network for feature extraction to generate a state correlation feature matrix;
[0231] A construction module 230, configured to construct a three-layer fully connected Actor network and a Critic network with a double Q structure based on the state correlation feature matrix, and output an equalization current value, a constant current or constant voltage equalization mode and an equalization time through the Actor network;
[0232] A control module 240, configured to control the DC-DC converter in the main equalization circuit to generate a constant current output according to the equalization current value, equalization mode and equalization time, or control the Buck-Boost converter in the secondary equalization circuit to generate a constant voltage output.
[0233] Through the collaborative cooperation of the above-mentioned various components, by introducing a multi-head self-attention network for feature extraction, the efficient processing of the state information of the power battery pack is realized, the mutual correlation between different monomers is accurately captured, and the expression ability of the state features is improved. Using the Actor-Critic deep reinforcement learning framework, a combined strategy of equalization current, equalization mode and equalization time is dynamically generated, and the adaptive adjustment of equalization control is realized, so that the equalization strategy can be optimized according to the real-time state of the battery pack. A multi-level equalization circuit structure in which the main equalization circuit and the secondary equalization circuit work together is designed, and an equalization mode switching mechanism based on the voltage difference is established, which reduces the energy loss while ensuring the equalization efficiency. A complete equalization charging protection system is constructed. By real-time monitoring the single cell voltage, temperature and equalization current, and setting multi-level protection thresholds, the safety and reliability of the equalization charging process are ensured. The proximal policy optimization algorithm is used to update the network parameters online. By introducing the advantage function and policy entropy, the exploration efficiency of reinforcement learning is improved, and the convergence performance of the algorithm is enhanced. The temperature adaptive control of the equalization circuit is realized. By real-time monitoring the junction temperature of the switching tube and dynamically adjusting the duty cycle, the device overheating is effectively prevented, and the service life of the system is extended. Through the comprehensive evaluation of the voltage standard deviation, temperature distribution range and state of charge consistency index, the quantitative evaluation of the equalization effect and strategy optimization are realized.
[0234] Please refer to Figure 3 , Figure 3 which is a schematic block diagram of the equalizing charging device 300 for a power battery pack provided by an embodiment of the present application. The equalizing charging device 300 for a power battery pack includes a processor 301 and a memory 302. The processor 301 and the memory 302 are connected through a device bus 303. Among them, the memory 302 may include a non-volatile storage medium and an internal memory.
[0235] The non-volatile storage medium can store a computer program. The computer program includes program instructions. When the program instructions are executed by the processor 301, the processor 301 can be made to execute any of the above-mentioned equalizing charging methods for a power battery pack.
[0236] The processor 301 is used to provide computing and control capabilities to support the operation of the entire equalizing charging device 300 for a power battery pack.
[0237] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor 301, the processor 301 can be made to execute any of the above-mentioned equalizing charging methods for a power battery pack.
[0238] Those skilled in the art can understand that Figure 3 the structure shown in
[0239] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the equalizing charging device 300 for a power battery pack involved in the solution of the present application. The specific equalizing charging device 300 for a power battery pack may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0240] It should be noted that those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process of the equalizing charging device 300 of the power battery pack described above can refer to the corresponding process of the equalizing charging method of the power battery pack described above, and will not be elaborated here.
[0241] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, systems and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated here.
[0242] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0243] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present application.
Claims
1. A balanced charging method for a power battery pack, characterized in that: include: Collect parameters of the power battery pack, measure the terminal voltage, temperature, and charge and discharge current of each single battery, and calculate the state of charge and internal resistance change rate of each single battery; The terminal voltage value, the temperature value, the charge and discharge current value, the state of charge value and the internal resistance change rate are constructed into a state feature sequence, and are input into a multi-head self-attention network for feature extraction to generate a state-related feature matrix; specifically comprising: performing feature splicing and dimension transformation on the terminal voltage value, the temperature value, the charge and discharge current value, the state of charge value and the internal resistance change rate to obtain a state feature sequence; performing linear transformation on the state feature sequence to generate a query matrix, a key matrix and a value matrix, and performing matrix multiplication operation on the query matrix and the key matrix to obtain an attention score matrix; The attention score matrix performs a Softmax normalization operation to obtain an attention weight matrix, and the attention weight matrix is subjected to a matrix multiplication operation with the value matrix to obtain a single-head attention output feature; the state feature sequence is respectively input into N parallel attention calculation units in the multi-head self-attention network, and each attention calculation unit independently performs attention feature extraction to obtain N attention output features; the N attention output features are concatenated to obtain a target feature vector, and the target feature vector is normalized by the LayerNorm layer to obtain a state-related feature matrix; Based on the state association feature matrix, a three-layer fully connected Actor network and a dual-Q structure Critic network are constructed, and a balanced current value, a constant current or constant voltage balanced mode and a balanced time are output through the Actor network; According to the balancing current value, the balancing mode and the balancing time, the DC-DC converter in the main balancing loop is controlled to generate a constant current output, or the Buck-Boost converter in the secondary balancing loop is controlled to generate a constant voltage output; specifically, the method comprises: comparing and analyzing the balancing current value and the voltage difference between adjacent single cells, selecting the main balancing loop for balancing charging when the voltage difference is greater than a preset voltage value, and selecting the secondary balancing loop for balancing charging when the voltage difference is less than or equal to the preset voltage value, to obtain a target balancing loop; generating a target duty cycle value sequence according to the type of the target balancing loop and the balancing mode, and performing pulse width modulation on the target duty cycle value sequence to obtain a switch tube driving signal; when the target balancing loop is the main balancing loop, the switch tube driving signal is switched to the secondary balancing loop; The tube driving signal is input into the main switch tube of the DC-DC converter to control the on-off timing of the main switch tube to obtain a constant current output; when the target balancing loop is the secondary balancing loop, the switch tube driving signal is input into the four switch tubes of the Buck-Boost converter to control the on-off timing of the four switch tubes to obtain a constant voltage output; the junction temperature of the switch tube in the target balancing loop is detected in real time, and when the junction temperature of the switch tube exceeds the target temperature, the target duty cycle value sequence is derated according to the temperature over-limit value to obtain a first duty cycle value sequence; the first duty cycle value sequence is input into a steady-state duty cycle compensator, the first duty cycle value sequence is feedback compensated to obtain a second duty cycle value sequence, and the switch tube driving signal is updated.
2. The balanced charging method for a power battery pack according to claim 1, characterized in that: The parameter collection of the power battery pack, measuring the terminal voltage value, temperature value and charge and discharge current value of each single battery, and calculating the state of charge value and internal resistance change rate of each single battery, includes: Setting a sampling sequence for a plurality of sampling points in the power battery pack to obtain a parameter acquisition sequence matrix; According to the parameter acquisition timing matrix, voltage data is collected at the sampling points to obtain terminal voltage values, according to the parameter acquisition timing matrix, temperature data is collected at the sampling points to obtain temperature values, and according to the parameter acquisition timing matrix, current data is collected at the sampling points to obtain charge and discharge current values; Performing a time integration operation on the charge and discharge current value and superimposing it with the initial state of charge value to obtain a target value, and normalizing the target value with the nominal capacity of the battery to obtain a state of charge value; A differential operation is performed on the terminal voltage value and the charge and discharge current value at adjacent sampling points to obtain a differential result, and the differential result is subjected to Kalman filtering and ratio calculation to obtain an internal resistance change rate.
3. The balanced charging method for a power battery pack according to claim 1, characterized in that: The method of constructing a three-layer fully connected Actor network and a dual-Q structured Critic network based on the state association feature matrix, and outputting a balanced current value, a constant current or constant voltage balanced mode, and a balanced time through the Actor network, includes: The state association feature matrix is input into the first fully connected layer of the Actor network, where the first fully connected layer includes 512 neurons, and a nonlinear transformation is performed through a ReLU activation function to obtain the first layer of action features; The first layer of action features is input into the second fully connected layer of the Actor network. The second fully connected layer contains 256 neurons and is nonlinearly transformed through the ReLU activation function to obtain the second layer of action features. The second layer of action features is input into the third fully connected layer of the Actor network, where the third fully connected layer includes 32 neurons, and a nonlinear transformation is performed through a Tanh activation function to obtain action output features; The action output characteristics are input into three parallel action mapping layers respectively, the first action mapping layer outputs a balanced current value, the second action mapping layer outputs a balanced mode of a constant current mode or a constant voltage mode, and the third action mapping layer outputs a balanced time; Inputting the state association feature matrix and the action output feature combination into two Q networks with the same structure, each Q network comprises three fully connected layers, and the number of neurons in the three fully connected layers is 512, 256, and 1 respectively, to obtain a first Q value and a second Q value; A minimum value selection operation is performed on the first Q value and the second Q value to obtain a target Q value, and the target Q value is used as a value evaluation indicator of the Actor network output action.
4. The balanced charging method for a power battery pack according to claim 1, characterized in that: The balanced charging method of the power battery pack also includes: Performing standard deviation calculation on the terminal voltage value to obtain the voltage standard deviation, performing maximum and minimum value calculation on the temperature value to obtain the temperature distribution range, and performing root mean square deviation calculation on the state of charge value to obtain the consistency index; The voltage standard deviation, the temperature distribution extreme difference and the consistency index are weightedly combined to generate a reward function value, and a difference operation is performed between the reward function value in the current state and the target reward value to obtain an advantage function value; Calculating a policy gradient based on the advantage function value, truncating the policy gradient to obtain a truncated policy gradient, and calculating a policy entropy value according to a logarithmic probability density of an output action of the Actor network in an action space; Performing weighted superposition on the truncated policy gradient and the policy entropy value, and updating the weight parameters of the Actor network according to the learning rate to obtain updated weight parameters of the Actor network; Calculating the time series difference error according to the advantage function value, and performing mean square error calculation on the time series difference error to obtain the value loss function, and updating the weight parameters of the Critic network according to the learning rate to obtain the updated weight parameters of the Critic network; According to the updated Actor network weight parameter and the updated Critic network weight parameter, the balancing current value, the balancing mode and the balancing time are corrected to obtain an optimized balancing charging strategy.
5. The balanced charging method for a power battery pack according to claim 1, characterized in that: The balanced charging method of the power battery pack also includes: Collecting data on the cell voltage, temperature and balancing current in the primary balancing loop and the secondary balancing loop to obtain a cell monitoring data matrix; Comparing the cell voltage in the cell monitoring data matrix with a preset voltage range to obtain a voltage over-limit mark, comparing the temperature difference between adjacent cells with a preset temperature difference threshold to obtain a temperature difference over-limit mark, and comparing the balanced current with a current threshold of the rated current to obtain a current over-limit mark; Generate an over-limit level matrix according to the voltage over-limit mark, the temperature difference over-limit mark and the current over-limit mark, and prioritize the over-limit level matrix to obtain a protection action instruction; When the protection action instruction is to reduce the balancing current, the balancing current value is decreased and calculated to obtain a reduced balancing current value, and the reduced balancing current value is updated to the main balancing circuit or the secondary balancing circuit; When the protection action instruction is to interrupt the balanced charging, a balanced interrupt signal sequence is generated, and the balanced interrupt signal sequence is input into the switch tube control terminals of the main balanced circuit and the secondary balanced circuit respectively, so that the switch tube enters the cut-off state; Feedback detection is performed on the derating equalizing current value or the execution state of the equalizing interrupt signal sequence to obtain a protection execution result, and the protection execution result is recorded in the equalizing charging log.
6. A balanced charging device for a power battery pack, characterized in that: A method for performing balanced charging of a power battery pack according to any one of claims 1 to 5, comprising: The parameter acquisition module is used to collect parameters of the power battery pack, measure the terminal voltage value, temperature value and charge and discharge current value of each single battery, and calculate the state of charge value and internal resistance change rate of each single battery; A feature extraction module, used to construct the terminal voltage value, the temperature value, the charge and discharge current value, the state of charge value and the internal resistance change rate into a state feature sequence, and input it into a multi-head self-attention network for feature extraction to generate a state correlation feature matrix; A construction module, used to construct a three-layer fully connected Actor network and a Critic network with a dual-Q structure based on the state association feature matrix, and output a balanced current value, a constant current or constant voltage balanced mode and a balanced time through the Actor network; The control module is used to control the DC-DC converter in the main balancing loop to generate a constant current output, or control the Buck-Boost converter in the secondary balancing loop to generate a constant voltage output according to the balancing current value, the balancing mode and the balancing time.
7. A balanced charging device for a power battery pack, characterized in that: The equalization charging device of the power battery pack includes: a memory and at least one processor, wherein the memory stores instructions; The at least one processor calls the instruction in the memory to enable the balanced charging device of the power battery pack to execute the balanced charging method for the power battery pack according to any one of claims 1 to 5.
Citation Information
Patent Citations
Battery pack equalization method based on reinforcement learning
CN116674431A
Equalization method of energy storage battery pack management system based on neural network and medium
CN117613421A
Cited By
Quick charging equipment for power battery
CN121133464A
A quick charging device for power battery
CN121133464B