Charging management method and system based on depth deterministic strategy gradient, and medium

By constructing a thermoelectric coupling model and a deep deterministic strategy gradient algorithm, the battery aging and safety problems in traditional charging strategies are solved, and the battery safety and life is improved, and an intelligent fast charging management method is provided.

CN120354809APending Publication Date: 2025-07-22HEFEI UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510420173.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

Existing charging strategies are difficult to improve battery safety and extend battery life while shortening charging time. Traditional strategies have low charging efficiency and inability to monitor the internal temperature of the battery in real time, resulting in accelerated battery aging and safety risks.

Method used

A thermoelectric coupling model is constructed, combining the extended Kalman filtering algorithm and the depth deterministic strategy gradient algorithm, and the accurate estimation of battery charge state and optimal charging strategy optimization through the integer-order equivalent circuit model and the two-state lumped parameter thermal model.

Benefits of technology

It achieves the ability to shorten charging time, reduce energy loss, and extend battery life while ensuring battery safety, and provides intelligent fast charging management strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354809A_ABST
    Figure CN120354809A_ABST
Patent Text Reader

Abstract

The invention discloses a charging management method and system based on a depth deterministic strategy gradient, and a medium, and the method achieves the precise estimation of the SOC of a battery through constructing a thermoelectric coupling model covering the state parameters of the battery based on an extended Kalman filtering algorithm, and achieves the formulation and optimization of a charging strategy in combination with a DDPG algorithm. Reliable battery state information is provided for a system charging management strategy by combining an SOC state estimation method based on an EKF algorithm, a theory and a method which comprehensively consider the charging time, the thermal safety and the service life of the battery are expected to be established, and an intelligent charging management strategy method is provided for a fast charging technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of battery charging management, and particularly relates to a charging management method, system and medium based on deep deterministic policy gradient. Background Art

[0002] Compared with traditional fuel vehicles, new energy vehicles mainly face two major technical bottlenecks in popularization and application: insufficient charging infrastructure and limited driving range. Among them, users have an urgent need for fast charging technology. However, fast charging technology shortens the charging time by increasing the charging current and power. This process will exacerbate the side reactions inside the battery, leading to irreversible aging phenomena such as the thickening of the solid electrolyte interface film (SEI film) and the growth of lithium dendrites, significantly reducing the battery cycle life. More seriously, the Joule heat effect caused by high-current charging will cause the internal temperature of the battery to rise sharply. When the temperature exceeds the safety threshold, it may trigger a thermal runaway chain reaction, posing a serious safety hazard. Therefore, the development of lithium battery fast charging technology should find the best balance between charging time and battery safety.

[0003] Currently, the mainstream charging strategies can be mainly divided into four categories: model-free control strategies, strategies based on equivalent circuit models, data-driven strategies, and intelligent charging strategies based on machine learning. A large number of experimental studies have shown that the internal temperature field distribution of the battery is a key control parameter affecting the battery performance degradation rate, capacity attenuation degree, and the probability of thermal faults. Traditional fixed-rule-based charging strategies, such as the constant current-constant voltage (CC-CV) charging mode, although simple to implement, have inherent defects such as low charging efficiency and inability to real-time monitor the internal temperature of the battery, making it difficult to achieve precise control of the charging process. At the same time, such strategies often have overcharging problems, resulting in side reactions such as the destruction of the cathode material structure and the decomposition of the electrolyte, accelerating the battery aging process, and seriously affecting the service life and safety of the battery. Traditional charging strategies cannot balance charging time and the internal temperature of the battery, greatly affecting the charging experience of users.

[0004] In summary, how to find a charging strategy that can both shorten the charging time and improve battery safety and extend battery life is a problem that needs to be solved by the existing technology. Summary of the Invention

[0005] The purpose of the present invention is to overcome the above problems existing in the prior art, and provide a charging management method, system and medium based on deep deterministic policy gradient, which comprehensively considers the charging time, energy loss and internal temperature of the battery to achieve a safe and reliable charging strategy method.

[0006] To achieve the above technical objectives and reach the above technical effects, the present invention is realized through the following technical solutions:

[0007] A charging management method based on deep deterministic policy gradient, comprising:

[0008] Construct a thermoelectric coupling model composed of an integer-order equivalent circuit model and a two-state lumped parameter thermal model;

[0009] Estimate the state of charge of the battery based on the extended Kalman filter algorithm;

[0010] Analyze and obtain the optimal charging strategy based on the thermoelectric coupling model, the battery state of charge estimation result, and the deep deterministic policy gradient algorithm.

[0011] Further, the integer-order equivalent circuit model is a first-order RC equivalent circuit model for characterizing the polarization characteristics of the battery; the two-state lumped parameter thermal model is used to describe the dynamic changes of the battery core temperature and surface temperature, and the least squares method is used for parameter identification of the two-state lumped parameter thermal model.

[0012] Further, the integer-order equivalent circuit model is specifically as follows:

[0013] The internal voltage drop of the battery is defined as the system output variable. By performing Laplace transform on the mathematical expression of the first-order RC equivalent circuit model, the transfer function of the system is obtained:

[0014]

[0015] Let the sampling time interval be T s , and it is calculated that:

[0016] U t (k) = a1U t (k - 1) + b1I(k) + b2I(k - 1)

[0017] In the formula, a1, b1, and b2 are the parameters of the discretized model, and there is the following corresponding relationship between the discretized model parameters and the continuous model parameters R0, R p , C p :

[0018]

[0019] The standard expression of the least squares method is as follows:

[0020] U t (k) = [-U t (k - 1), I(k), I(k - 1)][a1, b1, b2] + r(k) = Φ T (k)θ(k) + δ(k)

[0021] By setting the initial parameter values, at each sampling period T sRecursively update the parameters a1, b1, and b2 internally to achieve the dynamic correction of the model parameters R0, R p , C p .

[0022] Furthermore, the two-state lumped parameter thermal model is specifically as follows:

[0023]

[0024] Furthermore, estimating the state of charge of the battery based on the extended Kalman filter algorithm includes:

[0025] Based on the Kalman filter algorithm, use the first-order Taylor expansion linearization and the Jacobian matrix. Based on the integer-order equivalent circuit model, realize the observation of the nonlinear discrete system to achieve the accurate estimation of the state of charge of the battery;

[0026] Among them, the Kalman filter algorithm combines the open-circuit voltage method and the ampere-hour integration method, and completes the estimation of the state of charge of the battery by continuously correcting the relationship between the open-circuit voltage and the state of charge of the battery.

[0027] Furthermore, the deep deterministic policy gradient algorithm includes setting a reward function, specifically as follows:

[0028] The charging optimization objective function J obj can be described as:

[0029] J obj = ω ct * J ct + ω el * J el + ω tr * J tr

[0030] Among them, J ct , J el , J tr are the loss functions of the battery charging time, energy loss, and internal temperature rise respectively. ω ct , ω el , ω tr are the weights of the loss functions of the battery charging time, energy loss, and internal temperature rise respectively; at the kth moment, the specific expression of the loss function is as follows:

[0031]

[0032] Among them, soc(k) is the value of SOC at the kth moment, and i b,mean is the average current;

[0033] The reward function is set as R = -J obj, so as to minimize the loss function, that is, the smaller the value of the charging objective function, the larger the value of the reward function.

[0034] Furthermore, the deep deterministic policy gradient algorithm includes setting a penalty mechanism, which is specifically as follows:

[0035] At time k, the charging constraint conditions are as follows:

[0036] i b,m i n <i b (k) ≤ i b,max

[0037] V t,min ≤ V t (k) ≤ V t,max

[0038] soc m i n ≤ soc(k) ≤ soc max

[0039] T i (k) ≤ T i,max

[0040] Among them, i b,min , i b,max are respectively the upper and lower limit constraints of the input current; V t,min , V t,max are respectively the upper and lower limit constraints of the input voltage; soc min , soc max are respectively the upper and lower limit constraints of SOC; T i,max is the upper limit constraint of the internal temperature of the battery;

[0041] When the battery state exceeds the charging constraint conditions, a corresponding penalty is imposed on the current reward.

[0042] Furthermore, the analysis shows that the optimal charging strategy includes: using historical data and real-time data to train the deep deterministic policy gradient algorithm, and continuously iterating to optimize the strategy to determine the best charging strategy.

[0043] The present invention also provides a charging management system based on the deep deterministic policy gradient, including:

[0044] A model construction module for constructing a thermoelectric coupling model composed of an integer-order equivalent circuit model and a two-state lumped parameter thermal model;

[0045] A state estimation module for estimating the battery charge state based on the extended Kalman filter algorithm;

[0046] A strategy analysis module for analyzing and obtaining an optimal charging strategy based on a thermoelectric coupling model, a battery state of charge estimation result, and a deep deterministic policy gradient algorithm.

[0047] The present invention also provides a computer storage medium having a computer program stored thereon, and when the computer program is executed, the above-mentioned charging management method is implemented.

[0048] The beneficial effects of the present invention are:

[0049] In order to ensure the charging efficiency of the battery while ensuring thermal safety, the present invention constructs a thermoelectric coupling model, uses the EKF algorithm to accurately estimate the SOC of the battery, and the DDPG algorithm continuously trains and iterates according to a reward function considering charging time, energy loss, and internal temperature to determine the optimal charging strategy; by combining the SOC state estimation method based on the EKF algorithm, reliable battery state information is provided for the system charging management strategy, and it is expected to establish a theory and method for a charging management strategy that comprehensively considers battery charging time, thermal safety, and battery life, providing an intelligent fast charging algorithm for fast charging technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] The drawings described herein are used to provide a further understanding of the present invention, and constitute a part of this application. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0051] Figure 1 is a flowchart of the method of the present invention;

[0052] Figure 2 is a schematic diagram of a first-order RC equivalent circuit model in the present invention;

[0053] Figure 3 is a schematic diagram of a two-state lumped parameter thermal model in the present invention;

[0054] Figure 4 is a system structure block diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0055] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0056] As Figure 1 shown, a charging management method based on deep deterministic policy gradient includes:

[0057] Step 1: Construct a thermoelectric coupling model composed of an integer-order equivalent circuit model and a two-state lumped parameter thermal model.

[0058] The integer-order equivalent circuit model is a first-order RC equivalent circuit model, and the polarization characteristics of lithium-ion batteries can be characterized by the first-order RC equivalent circuit model. The model uses the recursive least squares method for parameter identification, and its structural schematic diagram is as Figure 2 shown, where R0 is the ohmic internal resistance, Rp is the polarization resistance, and Cp is the polarization capacitance; Uoc is the ideal voltage source, Up is the voltage drop of the RC parallel link, and Ut is the output terminal voltage; I is the current.

[0059] In the model, the internal voltage drop U of the battery t = U oc (k) - U p (k) is defined as the system output variable. By performing Laplace transform on the mathematical expression of the first-order RC model, the transfer function of the system can be derived:

[0060]

[0061] To convert the transfer function of the continuous system into a discrete form, let the sampling time interval be T s , then the above formula can be transformed into the following form:

[0062] U t (k) = a1U t (k - 1) + b1I(k) + b2I(k - 1)

[0063] In the formula, a1, b1, and b2 are the parameters of the discretized model, and there is the following corresponding relationship between the parameters of the discretized model and the parameters R0, R p , C p of the continuous model:

[0064]

[0065] Considering the measurement noise commonly existing in the experimental environment, it is assumed that the system noise δ(k) follows a Gaussian distribution with a mean of zero. Combining the above formulas, the standard expression of the least squares method is as follows:

[0066] U t (k) = [-U t (k - 1), I(k), I(k - 1)][a1, b1, b2] + r(k) = Φ T (k)θ(k) + δ(k)

[0067] By setting the initial parameter values, the parameters a1, b1, and b2 are recursively updated within each sampling period T s to realize the model parameters R0, Rp , C p Dynamic correction of

[0068] To more accurately describe the battery characteristics, the present invention selects the parameter identification results under different states of charge (SOC) as typical values to construct a dynamic characteristic model of the battery module. This method can effectively reflect the dynamic response characteristics of the battery under different working conditions.

[0069] The two-state lumped parameter thermal model is a commonly used model for simulating the heat conduction behavior of the battery, which can describe the dynamic changes of the core temperature and the surface temperature of the battery. The parameters of the two-state lumped parameter thermal model are identified by the least squares method, and its structural schematic diagram is as Figure 3 shown, where T i , T s , T a are the internal temperature, the surface temperature and the ambient temperature of the battery respectively; R i , Ro represent the thermal resistance between the inside and the surface of the battery and the thermal resistance between the surface of the battery and the environment respectively; C e , C s represent the heat capacity coefficients of the internal material and the surface material of the battery respectively; Q g represents the heat generated inside the battery.

[0070] The state space expression of the two-state lumped parameter thermal model is as follows:

[0071]

[0072] Step 2: Accurately estimate the state of charge (SOC) of the battery based on the extended Kalman filter (EKF) algorithm.

[0073] Based on the Kalman filter algorithm, by using the first-order Taylor expansion linearization and the Jacobian matrix, and based on the integer-order equivalent circuit model, the observation of the nonlinear discrete system is realized, which can more accurately describe the battery model to achieve the accurate estimation of the state of charge of the battery; among them, the Kalman filter algorithm combines the open circuit voltage method and the ampere-hour integration method, and completes the estimation of the state of charge of the battery by continuously correcting the relationship between the open circuit voltage and the state of charge of the battery.

[0074] The SOC state estimation method based on the EKF algorithm can accurately estimate the SOC value of the battery under the current state. The SOC value range of the battery is from 10% to 90%, with each 5% as an interval. The entire charging process can be divided into 16 different states. Each state corresponds to an SOC interval.

[0075] Step 3: Analyze and obtain the optimal charging strategy based on the thermoelectric coupling model, the battery state of charge estimation result and the deep deterministic policy gradient (DDPG) algorithm.

[0076] Among them, the action range of the DDPG algorithm is defined as follows: taking the charging current as the action variable of the DDPG algorithm, and its value range is the action space.

[0077] The basic principle of the Deep Deterministic Policy Gradient (DDPG) algorithm can be summarized as follows:

[0078] Dual network structure: In the DDPG algorithm, both the policy function and the value function adopt a dual neural network architecture. Specifically, both the policy function and the value function include an online network and a target network. The online network is used to update the estimates of the policy and value in real time, and the target network is used to stabilize the training process and avoid training instability caused by rapid changes in network parameters. This dual network structure gradually synchronizes the parameters of the online network and the target network through a soft update mechanism to ensure the convergence and stability of the algorithm.

[0079] In the DDPG algorithm, the policy network is responsible for generating deterministic actions, while the value network evaluates the value (Q-value) of this action, that is, the long-term reward for executing a certain action in the current state. Their corresponding target networks gradually synchronize the parameters of the online network through a soft update mechanism to ensure the stability of training. This "actor-critic" structure enables DDPG to efficiently optimize the policy in a continuous action space.

[0080] Experience replay mechanism: The DDPG algorithm introduces an experience replay mechanism to store the experience data (including state, action, reward, and next state) generated during the interaction between the agent and the environment. These experience data are stored in the experience pool, and a batch of data is randomly sampled from it for learning during training. The main role of the experience replay mechanism is to break the temporal correlation between data, reduce sample dependence, and thus improve the stability and efficiency of training. In this way, DDPG can make full use of historical data for offline learning, improve the utilization rate of data, and avoid overfitting problems caused by continuous sampling;

[0081] Noise exploration mechanism: To enhance the exploration ability of the agent in the environment, the DDPG algorithm introduces a noise mechanism in the behavior policy. Commonly used noise types include time-correlated Ornstein-Uhlenbeck (OU) noise and Gaussian noise. OU noise has time correlation and can generate smooth random actions, which is suitable for continuous control tasks; while Gaussian noise increases the diversity of exploration by randomly perturbing the action values. The introduction of noise enables the agent to balance exploration and exploitation, avoid falling into local optimal solutions, and thus improve the robustness and generalization ability of the policy.

[0082] As a specific implementation manner of the present invention, the Deep Deterministic Policy Gradient algorithm includes setting a reward function, specifically as follows:

[0083] The core objective of charging management is to significantly shorten the charging duration. However, the use of high-power and high-current charging facilities can cause a series of negative effects: the side reactions inside the battery intensify, resulting in battery energy loss and internal temperature rise of the battery. Considering the safety aspect of battery use, excessive battery temperature may induce thermal faults and even pose a risk of thermal runaway in severe cases. Economically, energy loss not only increases the charging cost but also accelerates battery aging, affecting its service life and endurance performance. Moreover, the high cost of battery replacement will further increase the user's burden and affect the user experience. Based on the above considerations, the present invention innovatively establishes a multi-dimensional optimization objective function, incorporating charging time, energy loss, and internal temperature into a unified framework to achieve optimal control of the charging process.

[0084] The charging optimization objective function J obj can be described as:

[0085] J obj = ω ct * J ct + ω el * J el + ω tr * J tr

[0086] wherein, J ct , J el , J tr are the loss functions of battery charging time, energy loss, and internal temperature rise respectively. ω ct , ω el , ω tr are the weights of the loss functions of battery charging time, energy loss, and internal temperature rise respectively; at the k-th moment, the specific expressions of the loss functions are as follows:

[0087]

[0088] wherein, soc(k) is the value of SOC at the k-th moment, and i b,mean is the average current;

[0089] The reward function is set as R = -J obj , so as to minimize the loss function, that is, the smaller the value of the charging objective function, the larger the value of the reward function.

[0090] As a specific implementation manner of the present invention, the deep deterministic policy gradient algorithm includes setting a penalty mechanism, specifically as follows:

[0091] According to the charging constraint conditions, a penalty mechanism is set: to ensure the safety of the charging process, this solution adds constraint conditions to the objective function to limit the charging current, output voltage, internal temperature, and SOC (state of charge of the battery). The limitations on the charging current and output voltage can avoid the damage to the battery caused by pulsed current, the limitation on the internal temperature can effectively avoid the charging thermal fault problem of the battery, and the limitation on the SOC can avoid the overcharging and over-discharging problems of the battery.

[0092] At time k, the charging constraint conditions are as follows:

[0093] ib ,m i n <ib(k)≤ib ,max

[0094] V t,min ≤V t (k)≤V t,max

[0095] soc m i n ≤soc(k)≤soc max

[0096] T i (k)≤T i,max

[0097] Among them, i b,min , i b,max are the upper and lower limit constraints of the input current respectively; V t,min , V t,max are the upper and lower limit constraints of the input voltage respectively; soc min , soc max are the upper and lower limit constraints of the SOC respectively; T i,max is the upper limit constraint of the internal temperature of the battery;

[0098] When the battery state exceeds the charging constraint conditions, a corresponding penalty is imposed on the current reward.

[0099] As a specific implementation manner of the present invention, the analysis to obtain the optimal charging strategy specifically includes the following steps:

[0100] Initialize the OU noise, policy network, Q network and their target networks, and the experience replay buffer; set the observation period and training period.

[0101] During the observation period, the policy network randomly selects actions based on the OU noise and does not update the neural network.

[0102] Execute the action and update the state, calculate the current reward according to the reward function, and store the experience data (state, action, reward, next state) in the experience replay buffer.

[0103] Repeat the above steps until the SOC reaches the 16th state, end this round of training, initialize the battery-related variables, and enter the next round.

[0104] During the training period, the policy network generates deterministic actions based on the current state and adds OU noise to them to increase exploration.

[0105] Execute the action and update the state, calculate the current reward according to the reward function, and store the experience data (state, action, reward, next state) in the experience replay buffer.

[0106] Repeat the above steps until the SOC reaches the 16th state, end this round of training, initialize the battery-related variables, and enter the next round.

[0107] The DDPG algorithm regularly randomly samples a batch of data from the experience replay buffer, calculates the target Q value, calculates the losses of the policy network and the Q network, updates the parameters of the policy network and the Q network using the gradient descent method, and synchronizes the target network using the soft update mechanism.

[0108] Use historical data and real-time data to train the DDPG algorithm, and determine the optimal charging strategy by continuously iteratively optimizing the policy.

[0109] As Figure 4 shown, the second aspect of the present invention also provides a charging management system based on deep deterministic policy gradient, including:

[0110] A model construction module for constructing a thermoelectric coupling model composed of an integer-order equivalent circuit model and a two-state lumped parameter thermal model;

[0111] A state estimation module for estimating the battery charge state based on the extended Kalman filter algorithm;

[0112] A policy analysis module for analyzing and obtaining the optimal charging strategy based on the thermoelectric coupling model, the battery charge state estimation result, and the deep deterministic policy gradient algorithm.

[0113] The third aspect of the present invention also provides a computer storage medium, on which a computer program is stored, and when the computer program is executed, the above method is implemented. The storage medium may include: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.

[0114] In the description of this specification, the descriptions referring to terms such as "one embodiment", "example", "specific example", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0115] The above has shown and described the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and what is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements fall within the scope of the present invention claimed.

Claims

1. A charging management method based on deep deterministic policy gradient, characterized in that Including: Construct a thermoelectric coupling model composed of an integer-order equivalent circuit model and a two-state lumped parameter thermal model; Estimate the state of charge of the battery based on the extended Kalman filter algorithm; Analyze and obtain the optimal charging strategy based on the thermoelectric coupling model, the battery state of charge estimation result, and the deep deterministic policy gradient algorithm.

2. The charging management method based on deep deterministic policy gradient according to claim 1, wherein The integer-order equivalent circuit model is a first-order RC equivalent circuit model, which is used to characterize the polarization characteristics of the battery; the two-state lumped parameter thermal model is used to describe the dynamic changes of the battery core temperature and surface temperature, and the least squares method is used to identify the parameters of the two-state lumped parameter thermal model.

3. The charging management method based on deep deterministic policy gradient according to claim 2, wherein The integer-order equivalent circuit model is specifically as follows: The internal voltage drop of the battery is defined as the system output variable. By performing Laplace transform on the mathematical expression of the first-order RC equivalent circuit model, the transfer function of the system is obtained: Let the sampling time interval be T s , and it is calculated that: U t u(k) = a1U t (k - 1) + b1I(k) + b2I(k - 1) where a1, b1, b2 are the parameters of the discretized model, and there is the following correspondence between the parameters of the discretized model and the parameters of the continuous model R0, R p , C p : The standard expression of the least squares method is as follows: U t u(k) = [-U t (k - 1), I(k), I(k - 1)][a1, b1, b2] + r(k) = Φ T (k)θ(k) + δ(k) By setting the initial parameter values, within each sampling period T s the parameters a1, b1, b2 are recursively updated to achieve the dynamic correction of the model parameters R0, R p , C p .

4. The charging management method based on deep deterministic policy gradient according to claim 3, wherein The two-state lumped parameter thermal model is specifically as follows:

5. The charging management method based on deep deterministic policy gradient according to claim 1, characterized in that Estimating the state of charge of the battery based on the extended Kalman filter algorithm includes: Based on the Kalman filter algorithm, using the first-order Taylor expansion linearization and the Jacobian matrix, and based on the integer-order equivalent circuit model, the observation of the nonlinear discrete system is realized to achieve the accurate estimation of the battery state of charge; Among them, the Kalman filter algorithm combines the open-circuit voltage method and the ampere-hour integration method, and completes the estimation of the battery state of charge by continuously correcting the relationship between the open-circuit voltage and the battery state of charge.

6. A charging management method based on deep deterministic policy gradient according to any one of claims 1-5, characterized in that The deep deterministic policy gradient algorithm includes setting a reward function, specifically as follows: The charging optimization objective function J obj can be described as: J obj = ω ct * J ct + ω el * J el + ω tr * J tr Among them, J ct , J el , J tr are the loss functions of the battery charging time, energy loss, and internal temperature rise respectively. ω ct , ω el , ω tr are the weights of the loss functions of the battery charging time, energy loss, and internal temperature rise respectively; at the k-th moment, the specific expressions of the loss functions are as follows: Among them, soc(k) is the value of SOC at time k, and i b,mean is the average current; The reward function is set as R = -J obj , so as to minimize the loss function, that is, the smaller the charging objective function value, the larger the reward function value.

7. A charging management method based on deep deterministic policy gradient according to claim 6, characterized in that The deep deterministic policy gradient algorithm includes setting a penalty mechanism, specifically as follows: At time k, the charging constraint conditions are as follows: i b,m i n <i b (k)≤i b,max V t,min ≤ V t (k) ≤ V t,max soc m i n ≤soc(k)≤soc max T i (k) ≤ T i,max wherein, i b,min , i b,max are the upper and lower limit constraints of the input current respectively; V t,min , V t,max are the upper and lower limit constraints of the input voltage respectively; soc min , soc max are the upper and lower limit constraints of SOC respectively; T i,max is the upper limit constraint of the internal temperature of the battery; When the battery state exceeds the charging constraint conditions, the current reward is punished accordingly.

8. The charging management method based on deep deterministic policy gradient according to claim 7, wherein Analyzing and obtaining the optimal charging strategy includes: using historical data and real-time data for training of the deep deterministic policy gradient algorithm, and determining the best charging strategy by continuously iterating and optimizing the strategy.

9. A charging management system based on deep deterministic policy gradient, characterized in that, Including: A model construction module, which is used to construct a thermoelectric coupling model composed of an integer-order equivalent circuit model and a two-state lumped parameter thermal model; A state estimation module, which is used to estimate the state of charge of the battery based on the extended Kalman filter algorithm; A policy analysis module, which is used to analyze and obtain the optimal charging strategy based on the thermoelectric coupling model, the battery state of charge estimation result, and the deep deterministic policy gradient algorithm.

10. A computer storage medium having a computer program stored thereon, characterized in that: When the computer program is executed, it implements the charging management method described in any one of claims 1-8.