Lithium ion battery low-temperature preheating charging segmented optimization method and system based on reinforcement learning

Through the deep reinforcement learning model based on reinforcement learning, the low-temperature preheating and charging process of lithium-ion batteries is optimized in segments, solving the problem of low battery charging efficiency in low temperature environments, and achieving rapid lossless charging and extended battery life.

CN120087183AActive Publication Date: 2025-06-03UNIV OF JINAN

Patent Information

Application Number
CN202510057380.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-06-03
Estimated Expiration
2045-01-14

AI Technical Summary

Technical Problem

In low temperature environments, the available capacity of lithium-ion batteries is low and the charging and discharging performance is poor. It is difficult for the prior art to achieve rapid lossless charging, and it is difficult to maintain the optimal charging temperature during the charging stage.

Method used

Using reinforcement learning-based methods, a deep reinforcement learning model is established and the low-temperature preheating and charging process is optimized in segments. Through high-frequency pulse current preheating optimization and heating and charging optimization of hybrid action space, heat generation and charging efficiency are achieved.

Benefits of technology

It improves the charging efficiency of lithium-ion batteries under low temperature conditions, extends battery life, and achieves fast lossless charging, solving the problem of low charging efficiency in low temperature environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120087183A_ABST
    Figure CN120087183A_ABST
Patent Text Reader

Abstract

The invention relates to a lithium ion battery low-temperature preheating charging segmented optimization method and system based on reinforcement learning, and the method comprises the steps: building a thermoelectric coupling model of a battery, and building a state transition expression of voltage, temperature and SOC; when the battery temperature is lower than the preset switching temperature, starting a high-frequency pulse preheating optimization process based on reinforcement learning; when the temperature of the battery reaches a preset switching temperature, a heating and charging process based on reinforcement learning is entered, and three modes of heating, charging and simultaneous heating and charging are adaptively switched, so that safe and rapid charging at a low temperature is realized. Through the deep reinforcement learning technology, the preheating and charging stages are optimized in a segmented mode, parameters in the preheating and charging processes are accurately controlled, rapid preheating and efficient charging are achieved, meanwhile, battery damage is reduced, and the service life of the battery is remarkably prolonged.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method and system for segment optimization of low-temperature preheating and charging of lithium-ion batteries based on reinforcement learning, belonging to the technical field of low-temperature charging of batteries. Background Art

[0002] Under low-temperature conditions, the available capacity of the battery is low and the charge and discharge performance is poor, and this attenuation will be particularly significant when the temperature reaches below zero. Therefore, non-destructive and rapid heating of the battery under low-temperature environment is the key way to solve this problem. Since the attenuation of the battery at low temperature is a reversible process, its performance can be restored when the battery temperature is raised to room temperature. Therefore, heating the lithium-ion battery is the mainstream solution to solve its poor low-temperature performance.

[0003] The alternating current heating method can heat the battery from the inside to the outside by using a charging current with a specific frequency and amplitude, and has the advantages of fast heating speed, high efficiency and uniform distribution. Especially, lithium insertion and deinsertion alternate within an alternating current cycle, which can effectively avoid lithium plating and will not cause irreversible damage to the battery capacity, etc., and has great development prospects. Therefore, the existing low-temperature charging methods usually adopt the method of constant current and constant voltage charging after heating or alternating heating and charging. Although this strategy can reduce the damage of low-temperature charging to the battery life to a certain extent, it takes a long time and it is difficult to maintain the optimal charging temperature during the charging stage. Therefore, how to realize the segmented optimization of the current during low-temperature preheating and charging is crucial for rapid non-destructive charging. In addition, when coupling the heating process and the charging process, how to achieve the fastest charging speed and the minimum loss is a challenging problem. Summary of the Invention

[0004] Aiming at the deficiencies of the prior art, the present invention provides a method for segment optimization of low-temperature preheating and charging of lithium-ion batteries based on reinforcement learning;

[0005] The present invention respectively establishes preheating optimization and heating and charging optimization based on deep reinforcement learning. Among them, in the preheating stage, the pulse and amplitude of the high-frequency pulse current are used as optimization variables to maximize heat generation; in the heating and charging stage, an improved algorithm for solving the hybrid action space is used to select among three modes of "heating", "heating and charging" and "charging", adjust the charging rate and heating loss weight according to user requirements, and optimize the parameters of the selected mode to achieve non-destructive rapid charging. This method is expected to solve the problems existing in the prior art during charging in a low-temperature environment, improve the charging efficiency and extend the battery life.

[0006] The present invention provides a system for segment optimization of low-temperature preheating and charging of lithium-ion batteries based on reinforcement learning.

[0007] Term Explanation:

[0008] 1. TD3, short for Twin Delayed Deep Deterministic policy gradient algorithm, is called Twin Delayed Deep Deterministic Policy Gradient in Chinese. It is a reinforcement learning algorithm, an improvement based on DDPG (Deep Deterministic Policy Gradient), aiming to solve the problem of overestimating Q values that may occur during the training of DDPG. TD3 improves the stability and performance of the algorithm by introducing techniques such as two evaluation networks, delayed update of the policy network, and target policy smoothing regularization.

[0009] 2. PDQN, namely Parametrized Deep Q-Networks Learning, effectively solves the reinforcement learning problem in the mixed action space by combining deep neural networks with parameterization techniques. It optimizes the learning process and improves the decision-making efficiency and stability in complex environments.

[0010] The technical solution of the present invention is as follows:

[0011] A method for segmentally optimizing the low-temperature preheating charging of lithium-ion batteries based on reinforcement learning, including:

[0012] Establish a thermoelectric coupling model of the battery and establish state transition expressions for voltage, temperature, and SOC;

[0013] When the battery temperature is lower than the preset switching temperature, start the high-frequency pulse preheating optimization process based on reinforcement learning; when the battery temperature reaches the preset switching temperature, enter the heating and charging process based on reinforcement learning.

[0014] Preferably according to the present invention, the thermoelectric coupling model of the battery is as follows:

[0015]

[0016] Among them, m represents the mass of the battery; C p represents the specific heat capacity of the battery; t is time; Q represents the heat generation rate of the battery; h represents the heat transfer coefficient of the battery; S represents the surface area of the battery; T represents the temperature of the battery; T 0 represents the ambient temperature.

[0017] Preferably according to the present invention, the state transition expressions for voltage, temperature, and SOC are as follows:

[0018] The voltage of the lithium-ion battery is calculated by the following method:

[0019] U O = U OCV + I ac · Z total (T, f, SOC)

[0020] Among them, U O is the output voltage of the lithium-ion battery; U OCV is the OCV of the lithium-ion battery; I ac is the peak-to-peak value of the AC component; Z total is the total AC impedance at the current moment, obtained by interpolation;

[0021] The calculation of the heat generation rate Q during the preheating stage is expressed as:

[0022]

[0023] Among them, Z Re is the real part of the AC impedance at the current moment, obtained by interpolation; SOC is defined as the ratio of the remaining capacity to the capacity in the fully charged state; f represents the frequency of the AC component;

[0024] Obtain the temperature T of the lithium-ion battery at the k + 1 moment k+1 :

[0025]

[0026] Among them, T s represents the time interval of the temperature;

[0027] The battery SOC is obtained by Coulomb counting; as follows:

[0028]

[0029] Among them, SOC int is the initial SOC of the battery, Q bat is the actual available capacity of the battery, and I is the current flowing through the battery, positive for charging and negative for discharging.

[0030] According to the preference of the present invention, during the optimization of high-frequency pulse preheating based on reinforcement learning, the network parameters of the reinforcement learning model are updated using the TD3 model; the TD3 model includes two evaluation networks and one action network; the evaluation network is used to evaluate the value of the action, and the action network is used to select the action; the action network outputs according to the current voltage, temperature, and SOC states, and according to the reward function, outputs the optimal pulse frequency as the action; the impedance information is obtained by interpolation, combined with the calculated heat generation rate Q, and the state transition is performed through the established state transition expressions of voltage, temperature, and SOC; by continuously interacting with the environmental state, the TD3 model agent learns the optimal preheating strategy, and the preheating strategy guides the selection of the pulse frequency to achieve the predetermined preheating effect; the entire optimization process is a cyclic operation until the preset number of iterations is reached.

[0031] Further preferably, during the optimization process of high-frequency pulse preheating based on reinforcement learning, the state space s(t) is defined as:

[0032] s(t) = {T(t)}

[0033] where T(t) represents the temperature of the battery and t represents the time step;

[0034] The action space a(t) is defined as:

[0035] a(t) = {A(t), f(t)}

[0036] where A(t) represents the amplitude of the high-frequency pulse current waveform and f(t) represents the frequency of the frequency pulse current waveform;

[0037] The reward function r(t) is set as:

[0038] r(t) = ω 1 r temp (t) + ω 2 r targ (t) + ω 3 r volt (t)

[0039] where ω 1 ~ω 3 are the weights of each part of the reward or punishment respectively; r temp (t) and r targ (t) are the target temperature difference penalty and the heating end reward respectively, and r volt (t) is the voltage overlimit penalty;

[0040] r temp (t) and r targ (t) expressions are respectively:

[0041] r temp (t) = T(t) - T target

[0042]

[0043] where T target is the preset target temperature and R is the heating end reward value;

[0044] r volt (t) expression is:

[0045]

[0046] where U limit is the charging cut-off voltage and U(t) is the current battery terminal voltage;

[0047] Preferably according to the present invention, during the heating and charging process based on reinforcement learning,

[0048] An action space a(t)' is established as follows:

[0049] a(t)' = {(k, x k )|k ∈ [1, 2, 3], x k ∈ [I ac (t), f(t), I dc (t)]}

[0050] Wherein, k represents different modes, and x k represents the action space. The corresponding relationship between discrete actions and continuous actions is:

[0051]

[0052] Among them, k = 1 indicates selecting the heating mode, and the associated parameters are the pulse amplitude I ac (t) and the frequency f(t); k = 2 indicates selecting the heating and charging mode, and the associated parameters are the pulse frequency f(t), the amplitude of the AC component I ac (t) and the amplitude of the DC component I dc (t); k = 3 indicates selecting the charging mode, and the associated parameter is only the amplitude of the DC current I dc (t);

[0053] The state space s(t)' is defined as:

[0054] s(t)' = {T(t), SOC(t)}

[0055] Wherein, T(t) represents the temperature of the battery, and SOC(t) represents the state of charge of the battery;

[0056] State transition is performed through the established state transition expressions of voltage, temperature, and SOC;

[0057] The calculation of the heat generation rate Q is related to the mode selected at the current moment:

[0058]

[0059] Wherein, Z Re and Z tot are the real part of the AC impedance and the total impedance at the current moment, obtained through interpolation; R i is the ohmic impedance at the current moment, obtained through interpolation; T represents the temperature of the battery, f represents the frequency of the AC current, and SOC refers to the state of charge of the battery.

[0060] The reward function is set as:

[0061]

[0062] Among them, α 1 to α 4 are the weights of the rewards or punishments of each part respectively; r SOC (t) and r Qac (t) are the target SOC difference penalty and the AC heat generation penalty respectively, and r time (t) is the charging duration penalty, and r volt (t) and r dclim (t) are the voltage overlimit penalty and the DC component overlimit penalty respectively;

[0063] r SOC (t) and r Q_ac (t) expressions are respectively:

[0064] r SOC (t) = SOC(t) - SOC target

[0065]

[0066] Among them, SOC target is the set SOC target value. The AC heat generation penalty is the penalty related to the AC component heat generation given when the heating or heating charging mode is selected. K is a scaling coefficient to ensure that the values of the SOC difference penalty and the AC heat generation penalty are in the same order of magnitude;

[0067] r time (t) expression is:

[0068] r time (t) = -(1 - e -tT )

[0069] Among them, T is the set maximum charging duration. As the charging progresses, this part of the reward gradually increases to complete the charging goal faster;

[0070] r volt (t) and r dclim (t) are the voltage overlimit penalty and the DC component overlimit penalty respectively, and the expressions are:

[0071]

[0072]

[0073] I dc_lim (T t , SOC t ) = [U limit - U ocv (T t , SOC t)] / R i (T t , SOC t )

[0074] wherein, U limit is the charging cut-off voltage; if the current terminal voltage value is greater than the cut-off voltage, a penalty with a value equal to the difference between the cut-off voltage and the current terminal voltage is given; I dc_lim is the DC boundary. When the heating charging or charging mode is selected and the DC component is greater than the DC boundary, a penalty is given.

[0075] Preferably according to the present invention, the heating charging process based on reinforcement learning includes:

[0076] Collect the initial state information of the current lithium battery and input these initial state information into the target policy network model; the target policy network model is a PDQN model;

[0077] At each time step, the PDQN model receives the current state as input and outputs the Q values of each possible action; the agent selects an action to execute based on the Q values and gives feedback according to the current state;

[0078] The PDQN model uses the gradient descent optimization algorithm to update the network parameters of the PDQN model according to the reward value calculated by the reward function and the new state information; this process continues, the agent continuously interacts with the environmental state, accumulates experience, and gradually improves the heating charging strategy through the update of the PDQN model; when the preset number of iterations is satisfied, the PDQN model has learned a better Q value function.

[0079] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of the method for segmental optimization of low-temperature preheating charging of lithium-ion batteries based on reinforcement learning.

[0080] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps of the method for segmental optimization of low-temperature preheating charging of lithium-ion batteries based on reinforcement learning.

[0081] The system for segmental optimization of low-temperature preheating charging of lithium-ion batteries based on reinforcement learning includes:

[0082] A thermoelectric coupling model establishment module configured to: establish a thermoelectric coupling model of the battery and establish state transition expressions of voltage, temperature, and SOC;

[0083] The preheating charging segmented optimization module is configured to: when the battery temperature is lower than the preset switching temperature, start the high-frequency pulse preheating optimization process based on reinforcement learning; when the battery temperature reaches the preset switching temperature, enter the heating charging process based on reinforcement learning.

[0084] The beneficial effects of the present invention are as follows:

[0085] In the low-temperature charging technology of lithium-ion batteries, the present invention proposes a solution with significant advantages and innovation, and specifically has the following advantages:

[0086] 1. Improve charging efficiency and extend battery life: Through the deep reinforcement learning technology, the preheating and charging stages are optimized segmentally, the parameters in the preheating and charging processes are precisely controlled, rapid preheating and efficient charging are achieved, while reducing battery damage and significantly extending the service life of the battery.

[0087] 2. Intelligence and wide applicability: The optimization strategy of the present invention is based on deep reinforcement learning, with strong intelligence and self-adaptability. The proposed strategy has the ability to adjust the charging rate and heating loss weight according to user needs, realizing personalized charging and having wide popularization and application value.

[0088] 3. Solve the low-temperature charging problem: The present invention is specifically optimized for low-temperature environments, effectively solving the problem of low charging efficiency in low-temperature environments and improving the reliability and safety of charging.

[0089] In summary, the present invention significantly improves the charging efficiency of lithium-ion batteries through deep reinforcement learning technology, extends the battery life, and has intelligence and wide applicability. It is specifically optimized for low-temperature environments, solves the low-temperature charging problem, provides strong support for the progress of battery charging technology, and has significant application prospects and technical advantages. Brief Description of the Drawings

[0090] Figure 1 It is a schematic flow chart of the method for segmentally optimizing the low-temperature preheating charging of lithium-ion batteries based on reinforcement learning;

[0091] Figure 2(a) is a schematic diagram of the comparison results of the preheating pulse amplitudes of the method of the present invention under different SOCs;

[0092] Figure 2(b) is a schematic diagram of the comparison results of the frequencies of the method of the present invention under different SOCs;

[0093] Figure 3(a) is a schematic diagram of the comparison results of the battery temperatures of the method of the present invention under different switching temperatures;

[0094] Figure 3(b) is a schematic diagram of the comparison results of the SOCs of the method of the present invention under different switching temperatures. Detailed Embodiments

[0095] The present invention will be further defined below in conjunction with the accompanying drawings of the specification and embodiments, but is not limited thereto.

[0096] Embodiment 1

[0097] A segmented optimization method for low-temperature preheating charging of lithium-ion batteries based on reinforcement learning, as Figure 1 shown, includes:

[0098] Establish a thermoelectric coupling model of the battery, and establish state transition expressions for voltage, temperature, and SOC;

[0099] When the battery temperature is lower than the preset switching temperature, it is necessary to preheat and charge the battery, and start the high-frequency pulse preheating optimization process based on reinforcement learning; when the battery temperature reaches the preset switching temperature, enter the heating and charging process based on reinforcement learning.

[0100] Embodiment 2

[0101] The segmented optimization method for low-temperature preheating charging of lithium-ion batteries based on reinforcement learning according to Embodiment 1 is characterized in that:

[0102] The thermoelectric coupling model of the battery is as follows:

[0103]

[0104] Among them, m represents the mass of the battery; C p represents the specific heat capacity of the battery; t is time; Q represents the heat generation rate of the battery; h represents the heat transfer coefficient of the battery; S represents the surface area of the battery; T represents the temperature of the battery; T 0 represents the ambient temperature.

[0105] The state transition expressions for voltage, temperature, and SOC are as follows:

[0106] The voltage of the lithium-ion battery is calculated by the following method:

[0107] U O = U OCV + I ac · Z total (T, f, SOC)

[0108] Among them, U O is the output voltage of the lithium-ion battery; U OCV is the OCV of the lithium-ion battery; I ac is the peak-to-peak value of the AC component; Z total is the total AC impedance at the current moment, obtained by interpolation;

[0109] The calculation of the heat generation rate Q in the preheating stage is expressed as:

[0110]

[0111] where Z Re is the real part of the AC impedance at the current moment, obtained by interpolation; SOC is defined as the ratio of the remaining capacity to the capacity at full charge state; f represents the frequency of the AC component;

[0112] Discretize the above formula to obtain the temperature T of the lithium-ion battery at time k+1 k+1 :

[0113]

[0114] where T s represents the time interval of the temperature;

[0115] The battery SOC is obtained by Coulomb counting; as follows:

[0116]

[0117] where SOC int is the initial SOC of the battery, Q bat is the actual available capacity of the battery, I is the current flowing through the battery, positive for charging and negative for discharging.

[0118] During the high-frequency pulse preheating optimization based on reinforcement learning, in order to achieve the optimal preheating strategy, the TD3 (Twin Delayed Deep Deterministic policy gradient) model is used to update the network parameters of the reinforcement learning model; the TD3 model includes two evaluation networks and one action network; the evaluation network is used to evaluate the value of the action, and the action network is used to select the action; these networks use deep neural networks to approximate the state-action value function and the policy function. The action network outputs according to the current voltage, temperature and SOC state, and according to the reward function, outputs the optimal pulse frequency as the action; the impedance information is obtained by interpolation, combined with the calculated heat generation rate Q, and the state transfer is carried out through the established state transfer expressions of voltage, temperature and SOC; by continuously interacting with the environmental state, the TD3 model agent learns the optimal preheating strategy, and the preheating strategy guides the selection of the pulse frequency to achieve the predetermined preheating effect; the whole optimization process is a loop operation until the preset number of iterations is reached.

[0119] During the high-frequency pulse preheating optimization based on reinforcement learning, the state space s(t) is defined as:

[0120] s(t) = {T(t)}

[0121] where T(t) represents the temperature of the battery and t represents the time step;

[0122] Define the action space a(t) as:

[0123] a(t) = {A(t), f(t)}

[0124] Wherein, A(t) represents the amplitude of the high-frequency pulse current waveform, and the value range is [-25A, 25A]; f(t) represents the frequency of the frequency pulse current waveform, and the value range is [1000Hz, 10000Hz];

[0125] Set the reward function r(t) as:

[0126] r(t) = ω 1 r temp (t) + ω 2 r targ (t) + ω 3 r volt (t)

[0127] Wherein, ω 1 ~ω 3 Are the weights of each part of the reward or punishment respectively; r temp (t) and r targ (t) are the target temperature difference penalty and the heating end reward respectively, and r volt (t) is the voltage overlimit penalty;

[0128] r temp (t) and r targ (t) expressions are respectively:

[0129] r temp (t) = T(t) - T target

[0130]

[0131] Wherein, T target Is the preset target temperature, and R is the heating end reward value; the target temperature difference penalty is the difference between the current temperature and the target temperature at each time step, and the penalty is negative, that is, the lower the current temperature, the greater the penalty, and vice versa. The heating end reward is a one-time reward given when the battery temperature is heated to the target temperature. The two work together to achieve the fastest heating to the target temperature.

[0132] r volt (t) expression is:

[0133]

[0134] Wherein, U limit Is the charging cut-off voltage, and U(t) is the current battery terminal voltage; if the current terminal voltage value is greater than the cut-off voltage, a voltage overlimit penalty is given to ensure safety during the battery heating process.

[0135] During the heating and charging process based on reinforcement learning,

[0136] An action space a(t)' is established as follows:

[0137] a(t)' = {(k, x k ) | k ∈ [1, 2, 3], x k ∈ [I ac (t), f(t), I dc (t)]}

[0138] Wherein, k represents different modes, and x k represents the action space, and the correspondence between discrete actions and continuous actions is:

[0139]

[0140] Among them, k = 1 represents selecting the heating mode, and the associated parameters are the pulse amplitude I ac (t) and the frequency f(t); k = 2 represents selecting the heating and charging mode, and the associated parameters are the pulse frequency f(t), the amplitude of the AC component I ac (t) and the amplitude of the DC component I dc (t); k = 3 represents selecting the charging mode, and the associated parameter is only the amplitude of the DC current I dc (t);

[0141] The state space s(t)' is defined as:

[0142] s(t)' = {T(t), SOC(t)}

[0143] Wherein, T(t) represents the temperature of the battery, and SOC(t) represents the state of charge of the battery;

[0144] State transition is performed through the established state transition expressions of voltage, temperature, and SOC;

[0145] The calculation of the heat generation rate Q is related to the mode selected at the current moment:

[0146]

[0147] Wherein, Z Re and Z tot are the real part of the AC impedance and the total impedance at the current moment, obtained through interpolation; R i is the Ohmic impedance at the current moment, obtained through interpolation; T represents the temperature of the battery, f represents the frequency of the AC current, and SOC refers to the state of charge of the battery.

[0148] The reward function is set as:

[0149] r(t)' = α 1 r SOC (t) + (1 - α 1 )r Qac (t) + α 2 r time (t) + α 3 r volt (t) + α 4 r dclim (t)

[0150] where α 1 to α 4 are the weights of each part of the reward or punishment respectively; r SOC (t) and r Qac (t) are the target SOC difference penalty and the AC heat generation penalty respectively, r time (t) is the charging duration penalty, r volt (t) and r dclim (t) are the voltage overlimit penalty and the DC component overlimit penalty respectively;

[0151] r SOC (t) and r Qac (t) are expressed as follows:

[0152] r SOC (t) = SOC(t) - SOC target

[0153]

[0154] where SOC target is the set SOC target value. The smaller the SOC value at the current moment, the greater the SOC difference penalty, which promotes the overall charging process to reach the SOC target value and be fully charged quickly; the AC heat generation penalty is the penalty related to the AC component heat generation given when choosing the heating or heating charging mode. K is a scaling factor to ensure that the values of the SOC difference penalty and the AC heat generation penalty are of the same order of magnitude;

[0155] r time (t) is expressed as:

[0156] r time (t) = -(1 - e -tT )

[0157] where T is the set maximum charging duration. As the charging progresses, this part of the reward gradually increases to complete the charging target faster;

[0158] r volt (t) and r dclim (t) are the voltage overlimit penalty and the DC component overlimit penalty respectively, and the expressions are:

[0159]

[0160]

[0161] I dc_lim (T t ,SOC t )=[U limit -U ocv (T t ,SOC t )] / R i (T t ,SOC t )

[0162] Among them, U limit is the charging cut-off voltage; if the current terminal voltage value is greater than the cut-off voltage, a penalty with a value equal to the difference between the cut-off voltage and the current terminal voltage is given; I dc_lim is the DC boundary. When selecting the heating charging or charging mode and the DC component is greater than the DC boundary, a penalty is given.

[0163] The heating charging process based on reinforcement learning includes:

[0164] Adopt a PDQN model with the ability to handle mixed action spaces.

[0165] Collect the initial state information of the current lithium battery, such as temperature, SOC, etc., and input this initial state information into the target policy network model; the target policy network model is a PDQN model;

[0166] The target policy network model is a deep neural network that has been trained to output an optimal heating charging strategy based on the current state. The PDQN model is a parameterized deep Q-network that learns how to evaluate the value of performing different actions in different states. At each time step, the PDQN model receives the current state as input and outputs the Q-values of each possible action; the agent selects an action to execute based on the Q-values and gives feedback according to the current state;

[0167] The PDQN model uses optimization algorithms such as gradient descent to update the network parameters of the PDQN model based on the reward value calculated by the reward function and the new state information; this process aims to enable the PDQN model to more accurately predict future rewards and optimize its estimation of the Q-value function accordingly. This process continues, the agent continuously interacts with the environmental state, accumulates experience, and gradually improves the heating charging strategy through the update of the PDQN model; when the preset number of iterations is met, the PDQN model has learned a better Q-value function.

[0168] Figure 2(a) is a schematic diagram of the battery temperature comparison results of the method of the present invention at different switching temperatures; Figure 2(b) is a schematic diagram of the SOC comparison results of the method of the present invention at different switching temperatures; it shows the preheating charging effects at switching temperatures of 10°C, 5°C, and 0°C when the ambient temperature is -20°C and the initial SOC value is 20%, and the temperature changes and SOC charging curves during the whole process are obtained. As can be seen from the figure, at different switching temperatures, the battery maintains a relatively high temperature during the heating process. Further observation shows that a higher switching temperature not only increases the average temperature during the heating and charging stage, but also speeds up the charging rate. Specifically, when the switching temperatures are set to 10°C, 5°C, and 0°C respectively, the SOC values at the end of charging are 0.5123, 0.5015, and 0.4840 respectively. Although the switching temperature of 0°C enables the heating and charging stage to start earlier, the relatively low average temperature during the whole process limits the allowable charging current, resulting in a slowdown in the SOC growth rate. However, a higher switching temperature also means that more energy is required during the preheating stage, thus providing a choice between lower heat consumption and a faster charging rate.

[0169] Figure 3(a) is a schematic diagram of the comparison results of the battery temperature of the method of the present invention and the traditional constant current heating followed by charging at a switching temperature of 10°C; Figure 3(b) is a schematic diagram of the comparison results of the SOC of the method of the present invention and the traditional constant pulse current heating followed by charging at a switching temperature of 10°C; it can be seen that the growth rate of SOC of the method of this patent is significantly higher than that of the traditional method within the same time, and the SOC values at the end of charging are 0.5123 and 0.3775 respectively, with an increase in the charging rate of about 35.71%.

[0170] Example 3

[0171] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of the method for segmental optimization of low-temperature preheating charging of lithium-ion batteries based on reinforcement learning described in Example 1 or 2.

[0172] Example 4

[0173] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps of the method for segmental optimization of low-temperature preheating charging of lithium-ion batteries based on reinforcement learning described in Example 1 or 2.

[0174] Example 5

[0175] A system for segmental optimization of low-temperature preheating charging of lithium-ion batteries based on reinforcement learning includes:

[0176] The thermoelectric coupling model establishment module is configured to: establish the thermoelectric coupling model of the battery and establish the state transition expressions of voltage, temperature, and SOC;

[0177] The preheating charging segmented optimization module is configured to: when the battery temperature is lower than the preset switching temperature, preheating charging of the battery is required, and start the high-frequency pulse preheating optimization process based on reinforcement learning; when the battery temperature reaches the preset switching temperature, enter the heating charging process based on reinforcement learning.

Claims

1. A segmented optimization method for low-temperature preheating and charging of lithium-ion batteries based on reinforcement learning, characterized in that: include: Establish a thermoelectric coupling model for the battery and establish state transition expressions for voltage, temperature and SOC; When the battery temperature is lower than the preset switching temperature, the high-frequency pulse preheating optimization process based on reinforcement learning is started; when the battery temperature reaches the preset switching temperature, the heating charging process based on reinforcement learning is entered.

2. The segmented optimization method for low-temperature preheating charging of lithium-ion batteries based on reinforcement learning according to claim 1 is characterized in that: The thermoelectric coupled model of the battery is shown below: Where m represents the mass of the battery; C p represents the specific heat capacity of the battery; t is the time; Q represents the heat generation rate of the battery; h represents the heat transfer coefficient of the battery; S represents the surface area of ​​the battery; T represents the temperature of the battery; T0 represents the ambient temperature.

3. The segmented optimization method for low-temperature preheating charging of lithium-ion batteries based on reinforcement learning according to claim 1 is characterized in that: The state transition expressions of voltage, temperature and SOC are as follows: The voltage of a lithium-ion battery is calculated as follows: U O =U OCV +I ac ·Z total (T,f,SOC) Among them, U O is the output voltage of the lithium-ion battery; U OCV is the OCV of lithium-ion battery; I ac is the peak-to-peak value of the AC component; Z total is the total impedance of the AC impedance at the current moment, obtained by interpolation; The calculation of the heat production rate Q in the preheating stage is expressed as: Among them, Z Re is the real part of the AC impedance at the current moment, obtained by interpolation; SOC is defined as the ratio of the remaining capacity to the capacity in the fully charged state; f represents the frequency of the AC component; Get the temperature T of the lithium-ion battery at time k+1 k+1 : Among them, T s The time interval for indicating the temperature; The battery SOC is obtained by coulomb counting; as shown below: Among them, SOC int is the initial SOC of the battery, Q bat is the actual available capacity of the battery, I is the current flowing through the battery, which is positive for charging and negative for discharging.

4. The segmented optimization method for low-temperature preheating charging of lithium-ion batteries based on reinforcement learning according to claim 1 is characterized in that: In the high-frequency pulse preheating optimization process based on reinforcement learning, the TD3 model is used to update the network parameters of the reinforcement learning model; the TD3 model includes two evaluation networks and one action network; the evaluation network is used to evaluate the value of the action, and the action network is used to select the action; the action network outputs the optimal pulse frequency as the action according to the current voltage, temperature and SOC state output according to the reward function; the impedance information is obtained by interpolation, combined with the calculated heat generation rate Q, and the state transition is performed through the established state transition expression of voltage, temperature and SOC; by constantly interacting with the environmental state, the TD3 model agent learns the optimal preheating strategy, and the preheating strategy guides the selection of pulse frequency to achieve the predetermined preheating effect; the entire optimization process is a cyclic operation until the preset number of iterations is reached.

5. The segmented optimization method for low-temperature preheating charging of lithium-ion batteries based on reinforcement learning according to claim 4 is characterized in that: In the high-frequency pulse preheating optimization process based on reinforcement learning, the state space s(t) is defined as: s(t)={T(t)} Where T(t) represents the temperature of the battery and t represents the time step; Define the action space a(t) as: a(t)={A(t),f(t)} Where A(t) represents the amplitude of the high-frequency pulse current waveform, and f(t) represents the frequency of the high-frequency pulse current waveform; Set the reward function r(t) to: r(t)=ω1r temp (t)+ω2r targ (t)+ω3r volt (t) Among them, ω1~ω3 are the weights of rewards or penalties for each part; r temp (t) and r targ (t) are target temperature difference penalty and heating end reward, r volt (t) is the voltage over-limit penalty; r temp (t) and r targ (t) The expressions are: r temp (t)=T(t)-T target Among them, T target is the preset target temperature, R is the heating end reward value; r volt (t) is expressed as: Among them, U limit is the charging cut-off voltage, and U(t) is the current battery terminal voltage.

6. The segmented optimization method for low-temperature preheating charging of lithium-ion batteries based on reinforcement learning according to claim 5 is characterized in that: In the heating and charging process based on reinforcement learning, Establish the action space a(t)′ as follows: a(t)′={(k,x k )|k∈[1,2,3],x k ∈[I ac (t),f(t),I dc (t)]} Among them, k represents different modes, x k Represents the action space, and the corresponding relationship between discrete actions and continuous actions is: Among them, k = 1 means selecting the heating mode, and the parameter associated with it is the pulse amplitude I ac (t) and frequency f(t); k = 2 indicates that the heating charging mode is selected, and the parameters associated with it are the pulse frequency f(t), the AC component amplitude I ac (t) and the DC component amplitude I dc (t); k = 3 indicates that the charging mode is selected, and the only parameter associated with it is the DC current amplitude I dc (t); Define the state space s(t)′ as: s(t)′={T(t),SOC(t)} Wherein, T(t) represents the temperature of the battery, and SOC(t) represents the state of charge of the battery; Perform state transition by establishing state transition expressions of voltage, temperature and SOC; The calculation of the heat generation rate Q is related to the mode selected at the current moment: Among them, Z Re and Z tot is the real part of the AC impedance and the total impedance at the current moment, obtained by interpolation; R i is the ohmic impedance at the current moment, obtained by interpolation; T represents the temperature of the battery, f represents the frequency of the AC current, and SOC refers to the state of charge of the battery; The reward function is set as: Among them, α1 to α4 are the weights of each part of the reward or punishment; r SOC (t) and r Qac (t) are the target SOC difference penalty and AC heat generation penalty, r time (t) is the charging time penalty, r volt (t) and r dclim (t) are the voltage over-limit penalty and DC component over-limit penalty respectively; r SOC (t) and r Qac (t) The expressions are: r SOC (t)=SOC(t)-SOC target Among them, SOC target is the set SOC target value, the AC heat generation penalty is the penalty related to the AC component heat generation given when the heating or heating charging mode is selected, and K is the scaling factor to ensure that the values ​​of the SOC difference penalty and the AC heat generation penalty are in the same order of magnitude; r time (t) is expressed as: r time (t)=-(1-e -tT ) Among them, T is the maximum charging time set. As the charging progresses, this part of the reward gradually increases, and the charging target is completed faster; r volt (t) and r dclim (t) are respectively the voltage over-limit penalty and the DC component over-limit penalty, and the expression is: I dc_lim (T t ,SOC t )=[U limit -U ocv (T t ,SOC t )] / R i (T t ,SOC t ) Among them, U limit is the charging cut-off voltage; if the current terminal voltage is greater than the cut-off voltage, a penalty equal to the difference between the cut-off voltage and the current terminal voltage is given; I dc_lim It is the DC boundary. When the heating charging or charging mode is selected and the DC component is greater than the DC boundary, a penalty is given.

7. The segmented optimization method for low-temperature preheating charging of lithium-ion batteries based on reinforcement learning according to claim 1 is characterized in that: The heating and charging process based on reinforcement learning includes: Collecting the initial state information of the current lithium battery and inputting the initial state information into the target strategy network model; the target strategy network model is a PDQN model; At each time step, the PDQN model receives the current state as input and outputs the Q value of each possible action; the agent selects an action to perform based on the Q value and gives feedback based on the current state; The PDQN model uses a gradient descent optimization algorithm to update the network parameters of the PDQN model based on the reward value calculated by the reward function and the new state information; this process continues, the agent constantly interacts with the environmental state, accumulates experience, and gradually improves the heating and charging strategy through the update of the PDQN model; when the preset number of iterations is met, the PDQN model has learned a better Q-value function.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the reinforcement learning-based lithium-ion battery low-temperature preheating charging segmented optimization method according to any one of claims 1 to 7 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the reinforcement learning-based lithium-ion battery low-temperature preheating charging segmented optimization method according to any one of claims 1 to 7 are implemented.

10. A lithium-ion battery low-temperature preheating charging segmentation optimization system based on reinforcement learning, characterized in that: include: The module for establishing the thermoelectric coupling model is configured to: establish the thermoelectric coupling model of the battery and establish the state transition expressions of voltage, temperature and SOC; The preheating charging segmented optimization module is configured as follows: when the battery temperature is lower than the preset switching temperature, the high-frequency pulse preheating optimization process based on reinforcement learning is started; when the battery temperature reaches the preset switching temperature, the heating charging process based on reinforcement learning is entered.

Citation Information

Patent Citations

  • Preheating charging control method for lithium ion battery

    CN113815494A

  • Low-temperature environment battery heating cooperative charging method

    CN115642673A

  • Method, device and equipment for controlling low-temperature charging and readable storage medium

    CN118478755A

  • Lithium ion battery low-temperature lossless pulse preheating method and system

    CN118712583A

  • Lithium ion battery low-temperature heating and charging collaborative optimization control method and system based on reinforcement learning

    CN119009203A

Cited By

  • Mobile phone pole charging heat-electricity multi-mode fusion safety protection system based on deep learning

    CN121172928A

  • Mobile phone extreme heat-electricity multi-modal fusion safety protection system based on deep learning

    CN121172928B