Strong-adaptability lithium ion battery state estimation method guided by deep reinforcement learning

Through the highly adaptable lithium-ion battery state estimation method guided by deep reinforcement learning, combined with the second-order RC equivalent circuit model and PID controller, the problem of insufficient accuracy and adaptability of existing SOC estimation methods is solved, and higher SOC estimation accuracy and adaptability are achieved.

CN119940112APending Publication Date: 2025-05-06CHONGQING UNIV

Patent Information

Application Number
CN202510015384.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing SOC estimation methods for lithium-ion batteries are difficult to achieve the expected accuracy and adaptability, and are affected by the uncertainty of the dynamics of the battery system, nonlinearity and strong coupling between battery parameters.

Method used

The strongly adaptable lithium-ion battery state estimation method guided by deep reinforcement learning is adopted, combined with the second-order RC equivalent circuit model and the PID controller, and the PID controller is automatically adjusted to match the real-time operating conditions through the dual-delay depth deterministic strategy gradient deep reinforcement learning algorithm and the dynamic composite reward function of adaptive constraints.

Benefits of technology

It significantly improves the accuracy, robustness and adaptability of SOC estimation of lithium-ion batteries, and reduces the impact of battery modeling uncertainty on PID controllers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940112A_ABST
    Figure CN119940112A_ABST
Patent Text Reader

Abstract

The invention relates to a high-adaptability lithium ion battery state estimation method guided by deep reinforcement learning, and belongs to the technical field of battery state estimation, and the method comprises the following steps: S1, selecting a second-order RC equivalent circuit model as a battery model, considering the deviation of an augmented model, building a PID controller, and estimating the state of charge (SOC) of a battery in real time; s2, constructing a double-delay depth deterministic strategy gradient depth reinforcement learning algorithm and a dynamic composite reward function with adaptive constraints; and S3, combining the PID controller with a double-delay depth deterministic strategy gradient depth reinforcement learning algorithm, carrying out model training, and then automatically adjusting the PID controller according to an optimal strategy determined by learning so as to match a real-time working condition. According to the state estimation method, the estimation precision, robustness and adaptivity are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of battery state estimation and relates to a highly adaptable lithium-ion battery state estimation method guided by deep reinforcement learning. Background Art

[0002] Accurate and robust battery state estimation is essential for the safe and reliable operation of battery systems. The battery management system estimates multiple states inside the battery based on measurable lithium-ion battery variables, and then uses these estimated states to achieve battery balancing, effective charging and discharging, thermal management, fault diagnosis and prediction. SOC is considered to be one of the most critical states. Given the uncertainty and complex nonlinear characteristics of battery modeling, existing methods are difficult to achieve the expected accuracy and adaptability. Therefore, developing a more accurate and adaptable battery SOC estimation method is more meaningful for the practical application of battery systems.

[0003] The SOC of a lithium-ion battery refers to the ratio of the remaining capacity of the battery to the current available capacity. Its function is equivalent to the fuel gauge of a fuel vehicle. It is manifested in the concentration of lithium ions in the cathode and anode at the microscopic level. It cannot be obtained by direct measurement and can only be estimated in a non-invasive way. Existing SOC estimation methods for lithium-ion batteries can be divided into four categories: direct calculation method, model-based method, data-driven method and model-data fusion method. The direct calculation method includes the table lookup method and the ampere-hour integration method. The biggest feature of this method is that it is easy to implement, but due to the uncertainty and open-loop characteristics of the measurement of related parameters, its accuracy is limited; the model-based method is based on a suitable battery model and combined with an advanced state estimation algorithm to realize the online estimation of the battery SOC. The battery model mainly refers to the electrochemical model and the equivalent circuit model. Among them, the equivalent circuit model uses a resistor-capacitor network to simulate the electrochemical behavior of the battery and can provide good accuracy over a wide frequency range. It can achieve a good balance between fidelity and complexity. This method is easy to implement, has sufficient accuracy, and is capable of online feedback correction, and has broad prospects for practical application; the data-driven method regards the battery as a black box model. This method does not need to capture any physical and chemical mechanisms related to the battery and has the advantages of high flexibility, high nonlinear matching, and strong adaptability. The effectiveness of this method depends largely on the quantity and quality of the battery dataset; the model-data based method achieves better estimation accuracy and robustness by integrating the advantages of each method, and has better fidelity and generalization ability.

[0004] The technical difficulty of lithium-ion battery state estimation lies in the uncertainty and nonlinearity of battery system dynamics and the strong coupling between battery parameters. In order to meet the growing safety requirements of lithium-ion batteries, control methods based on reinforcement learning have attracted attention due to their powerful adaptive learning capabilities. Reinforcement learning is a goal-oriented machine learning method that allows agents to interact with the environment through learning mechanisms. The pursuit of maximizing rewards guides agents to automatically learn optimal control. At present, deep deterministic policy gradient and double-delayed deep deterministic policy gradient algorithms are two important implementation methods of deep reinforcement learning algorithms. Summary of the invention

[0005] In view of this, the purpose of the present invention is to provide a highly adaptable lithium-ion battery state estimation method guided by deep reinforcement learning, combining a proportional-integral-derivative PID controller and deep reinforcement learning to design an enhanced battery state estimation framework to improve the accuracy, stability and adaptability of state of charge (SOC) estimation.

[0006] In order to achieve the above object, the present invention provides the following technical solutions:

[0007] A method for highly adaptive lithium-ion battery state estimation guided by deep reinforcement learning, comprising the following steps:

[0008] S1: Select the second-order RC equivalent circuit model as the battery model, consider the augmented model deviation, build a PID controller, and estimate the battery state of charge SOC in real time;

[0009] S2: Construct a double-delayed deep deterministic policy gradient deep reinforcement learning algorithm and a dynamic composite reward function with adaptive constraints;

[0010] S3: Combine the PID controller with the double-delayed deep deterministic policy gradient deep reinforcement learning algorithm and perform model training. Then, the PID controller is automatically adjusted to match the real-time operating conditions based on the optimal strategy determined by learning.

[0011] Further, the state space equation expression of the second-order RC model in step S1 is:

[0012]

[0013] Among them, C a is the charge transfer capacitance, R a is the charge transfer resistance, V a is the charge transfer polarization voltage, C b is the diffusion capacitance, R b is the concentration polarization resistance, V b is the voltage of the internal diffusion process of the battery, η i is the Coulomb efficiency, C tis the battery capacity, SOC is the battery state of charge, R ohm is the battery ohmic polarization internal resistance, V oc (s) is the open circuit voltage of the battery;

[0014] Considering the augmented model deviation, the state space equation is transformed into the state space equation form:

[0015]

[0016] Among them, V mb is the model bias;

[0017] The state space equation is organized into a state space expression:

[0018]

[0019] in, D=R ohm , y=V o , u=I,f(x)=V oc (soc)+V a +V b +V mb ;

[0020] The PID controller expression is as follows:

[0021]

[0022] Among them, K p , K i and K d are proportional gain matrix, integral gain matrix and differential gain matrix respectively, g and h are the integral and differential of system error respectively, K fi is the pre-gain for the control error integral.

[0023] Furthermore, the dual-delay deep deterministic policy gradient deep reinforcement learning algorithm adopts an actor-critic structure, in which the estimation of Q(s,a) uses the critic approximation function Q θ (s,a) and adjustable parameters θ, using a secondary frozen target network Q θ' (s,a) is updated, and the target y is kept constant through multiple updates. The expression is:

[0024] y=r+γQ θ' (s',a')

[0025] The dual-delay deep deterministic policy gradient deep reinforcement learning algorithm is built with six neural networks, a policy network With parameters For policy update, two value function networks critic are used to estimate Q(s,a), that is, Q with parameter θ1 θ1 and Q with parameter θ2 θ2 , each actor and critic corresponds to a target network, using Q θ1' and Q θ2' It also includes a double Q-learning algorithm, which takes the minimum of two estimates for target update. The update expression of the critic is:

[0026]

[0027] Among them, γ is the discount factor, r is the reward value, and N is the number of samples selected for conversion at each step;

[0028] Improve the critic update expression and limit the target to a small range. The expression is:

[0029]

[0030] Among them, ε~clip(N(0,σ),-c,c), N represents the standard normal distribution, σ is the noise standard deviation, and c is the noise limit;

[0031] Using the policy gradient method, the actor is updated by maximizing the Q function as follows:

[0032]

[0033] The soft update method is used to update the target network parameters as follows:

[0034] θ'←(1-τ)θ'+τθ

[0035]

[0036] where τ∈(0,1) is the learning rate.

[0037] Furthermore, the steps of constructing the dynamic composite reward function with adaptive constraints are as follows:

[0038] S21: The system output error is established based on the difference between the measured battery voltage and the estimated value. The expression is:

[0039] r1=w1·(1-w adr )·|e|

[0040] Among them, w1 is the weight coefficient, w adr is the adaptive weight;

[0041] The battery state error function is established based on the battery output voltage and the internal dynamics of the battery. The expression is:

[0042] r2=w2·w adr ·|e s |

[0043] Among them, w2 is the weight coefficient;

[0044] Establish an additional guided reward function, the expression is:

[0045] r3=w3·|e|·d1+w4·|e s |·d2

[0046] Among them, w3 and w4 are the output weight and state error weight respectively, and the values ​​of d1 and d2 are 0 or 1. The specific value selection method is as follows:

[0047]

[0048] S22: The dynamic composite reward function of the adaptive constraint is determined as the sum of the system output error, the battery state error function and the additional guidance reward function:

[0049] r=r1+r2+r3

[0050] Further, the PID controller is combined with the double-delayed deep deterministic policy gradient deep reinforcement learning algorithm described in step S3, specifically including:

[0051] The PID controller passes the system state space to the policy network actor and the value function network critic, and passes the batch processed data to the neural network layer of the double-delayed deep deterministic policy gradient deep reinforcement learning algorithm; the policy network passes the relevant information to the PID controller; finally, the PID controller automatically adjusts the PID controller to match the real-time working conditions based on the optimal strategy determined by learning;

[0052] The state space of the PID controller output contains the normalized battery current I cr , the time derivative of current dI cr / dt, measured battery voltage V o , the time derivative of voltage dV o / dt, estimated battery voltage Voltage output error e, voltage output error time derivative de / dt, estimated battery internal state and the time derivative of the battery internal state The state space representation of the PID controller output is:

[0053]

[0054] Among them, I cr =I / C t ,

[0055] Select PID control gain and adaptive reward gain as the action space, and the action space is expressed as:

[0056] A s ={K p1 ,K p2 ,K p3 ,K p4 ,K i1 ,K i2 ,K i3 ,K i4 ,K d1 ,K d2 ,K d3 ,K d4 ,w adr}

[0057] in, is the proportional gain, integral gain and differential gain under different battery states, 0<w adr <1 is the weight of the adaptive reward function;

[0058] The battery model parameters are updated to the optimal values ​​at the current stage by minimizing the loss function. The expression of the minimization loss function is:

[0059]

[0060] Further, the model training process described in step S3 is as follows:

[0061] The state space of the PID controller is used as the input of the double-delay deep deterministic policy gradient deep reinforcement learning algorithm. The actor network generates an action in the limited action space. The PID controller uses the double-delay deep deterministic policy gradient deep reinforcement learning algorithm agent action to estimate the internal state of the battery, and then uses the dynamic composite reward function with adaptive constraints to determine the reward value of the quality of the behavior. These transformations are stored in the circular buffer R. When enough experience is collected, the agent is trained. During the training process, small samples of N cycles are sampled, and the critic is updated through the critic update expression. The actor is updated after d iterations according to the policy gradient method. The target network is updated according to the soft update method. The iterative process continues to the maximum position N e .

[0062] The beneficial effects of the present invention are:

[0063] (1) Taking full advantage of the advantages of each algorithm, the PID controller is integrated with the double-delayed deep deterministic policy gradient deep reinforcement learning algorithm, which can adaptively adjust the PID control parameters according to the operating conditions and significantly improve the accuracy, robustness and adaptability of lithium-ion battery SOC estimation;

[0064] (2) Based on the second-order RC equivalent circuit model, an augmented deviation model is introduced to realize real-time model deviation compensation and reduce the impact of battery modeling uncertainty on the PID controller.

[0065] Other advantages, objectives and features of the present invention will be described in the following description to some extent, and to some extent, will be obvious to those skilled in the art based on the following examination and study, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below in conjunction with the accompanying drawings, wherein:

[0067] Figure 1 This is a schematic diagram of a highly adaptable lithium-ion battery state estimation method guided by deep reinforcement learning of the present invention;

[0068] Figure 2 It is a schematic diagram of the second-order RC equivalent circuit model of the present invention;

[0069] Figure 3 This is a schematic diagram of the principle of the double-delay deep deterministic policy gradient deep reinforcement learning algorithm of the present invention;

[0070] Figure 4 A flowchart for the specific implementation of the dynamic composite reward function with adaptive constraints of the present invention;

[0071] Figure 5 It is a schematic diagram of the principle of the dynamic composite reward function of the adaptive constraint of the present invention;

[0072] Figure 6 A block diagram of a PID controller regulated by a double-delayed deep deterministic policy gradient deep reinforcement learning algorithm for lithium-ion battery state estimation according to the present invention;

[0073] Figure 7 It is a complete training flow chart of the model of the present invention;

[0074] Figure 8 The figure is a flow chart for obtaining the OCV-SOC relationship of the battery and related battery parameters of the present invention. DETAILED DESCRIPTION

[0075] The following describes the embodiments of the present invention by specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner, and the following embodiments and features in the embodiments can be combined with each other without conflict.

[0076] It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and thus the drawings only show components related to the present invention rather than being drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component may be changed arbitrarily, and the component layout may also be more complicated.

[0077] In the following description, numerous details are discussed to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.

[0078] See also Figure 1 , a deep reinforcement learning-guided strongly adaptive lithium-ion battery state estimation method, comprising the following steps:

[0079] S1: Select the second-order RC equivalent circuit model as the battery model, consider the augmented model deviation, build a PID controller, and estimate the battery state of charge SOC in real time;

[0080] like Figure 2 , the state space equation expression of the second-order RC model in step S1 is:

[0081]

[0082] Among them, C a is the charge transfer capacitance, R a is the charge transfer resistance, V a is the charge transfer polarization voltage, C b is the diffusion capacitance, R b is the concentration polarization resistance, V b is the voltage of the internal diffusion process of the battery, η i is the Coulomb efficiency, C t is the battery capacity, SOC is the battery state of charge, R ohm is the battery ohmic polarization internal resistance, Voc (s) is the open circuit voltage of the battery;

[0083] Considering the augmented model deviation, the state space equation is transformed into the state space equation form:

[0084]

[0085] Among them, V mb is the model bias.

[0086] The state space equation is organized into a state space expression:

[0087]

[0088] in, D=R ohm , y=V o , u=I,f(x)=V oc (soc)+V a +V b +V mb .

[0089] The PID controller expression is as follows:

[0090]

[0091] Among them, K p , K i and K d are proportional gain matrix, integral gain matrix and differential gain matrix respectively, g and h are the integral and differential of system error respectively, K fi is the pre-gain for the control error integral.

[0092] S2: Construct a double-delayed deep deterministic policy gradient deep reinforcement learning algorithm and a dynamic composite reward function with adaptive constraints;

[0093] like Figure 3 In step S2, the double-delayed deep deterministic policy gradient deep reinforcement learning algorithm adopts an actor-critic structure, in which the estimate of Q(s,a) uses the critic approximation function Q θ (s,a) and adjustable parameters θ, using a secondary frozen target network Q θ' (s,a) is updated, and the target y is kept constant through multiple updates. The expression is:

[0094] y=r+γQ θ' (s',a')

[0095] The dual-delay deep deterministic policy gradient deep reinforcement learning algorithm is built with six neural networks, a policy network With parameters For policy update, two value function networks critic are used to estimate Q(s,a), that is, Q with parameter θ1 θ1 and Q with parameter θ2 θ2 , each actor and critic corresponds to a target network, using Q θ1' and Q θ2' It also includes a double Q-learning algorithm, which takes the minimum of two estimates for target update. The update expression of the critic is:

[0096]

[0097] Among them, γ is the discount factor, r is the reward value, and N is the number of samples selected for conversion at each step;

[0098] Improve the critic update expression and limit the target to a small range. The expression is:

[0099]

[0100] Among them, ε~clip(N(0,σ),-c,c), N represents the standard normal distribution, σ is the noise standard deviation, and c is the noise limit;

[0101] Using the policy gradient method, the actor is updated by maximizing the Q function as follows:

[0102]

[0103] The soft update method is used to update the target network parameters as follows:

[0104] θ'←(1-τ)θ'+τθ

[0105]

[0106] where τ∈(0,1) is the learning rate.

[0107] like Figure 4 and Figure 5 , the steps for constructing the dynamic composite reward function with adaptive constraints are as follows:

[0108] S21: The system output error is established based on the difference between the measured battery voltage and the estimated value. The expression is:

[0109] r1=w1·(1-w adr )·|e|

[0110] Among them, w1 is the weight coefficient, w adr is the adaptive weight.

[0111] The battery state error function is established based on the battery output voltage and the internal dynamics of the battery. The expression is:

[0112] r2=w2·w adr ·|e s |

[0113] Among them, w2 is the weight coefficient.

[0114] In order to make the output result closer to the expected value, an additional guided reward function is established, the expression is:

[0115] r3=w3·|e|·d1+w4·|e s |·d2

[0116] Among them, w3 and w4 are the output weight and state error weight respectively, and the values ​​of d1 and d2 are 0 or 1. The specific value selection method is as follows:

[0117]

[0118] S22: The dynamic composite reward function of the adaptive constraint is determined as the sum of the system output error, the battery state error function and the additional guidance reward function:

[0119] r=r1+r2+r3

[0120] S3: Combine the PID controller with the double-delayed deep deterministic policy gradient deep reinforcement learning algorithm and perform model training. Then, the PID controller is automatically adjusted to match the real-time working conditions based on the optimal policy determined by learning.

[0121] like Figure 6 In S3, the method of combining the PID controller with the double-delay deep deterministic policy gradient deep reinforcement learning algorithm is that the PID controller passes the system state space to the policy network actor and the value function network critic, and passes the batch processed data to the neural network layer of the double-delay deep deterministic policy gradient deep reinforcement learning algorithm; then, the policy network passes the relevant information to the PID controller; finally, the PID controller automatically adjusts the PID controller to match the real-time working conditions according to the optimal strategy determined by learning.

[0122] The state space design of the PID controller output minimizes the error between the SOC estimate and the actual value and continuously tracks the battery voltage. The state space contains the normalized battery current I cr , the time derivative of current dI cr / dt, measured battery voltage V o , the time derivative of voltage dV o / dt, estimated battery voltage Voltage output error e, voltage output error time derivative de / dt, estimated battery internal state and the time derivative of the battery internal state The PID state space is expressed as:

[0123]

[0124] Among them, I cr =I / C t ,

[0125] Since multiple internal state control gains of the battery affect the estimation performance, PID control gain and adaptive reward gain are selected as the action space, and the action space is expressed as:

[0126] A s ={K p1 ,K p2 ,K p3 ,K p4 ,K i1 ,K i2 ,K i3 ,K i4 ,K d1 ,K d2 ,K d3 ,K d4 ,w adr}

[0127] in, is the proportional gain, integral gain and differential gain under different battery states, 0<w adr <1 is the weight of the adaptive reward function.

[0128] The augmented model deviation can compensate for the instantaneous uncertainty, model defects and disturbances during the operation of lithium-ion batteries. However, as time goes on, when the battery response differs greatly from the model value identified by the augmented model deviation, the RLS function is activated and the battery model parameters are updated to the optimal value at the current stage by minimizing the loss function. The expression for minimizing the loss function is:

[0129]

[0130] like Figure 7 , the model training process described in step S3 is as follows:

[0131] The state space of the PID controller is used as the input of the double-delay deep deterministic policy gradient deep reinforcement learning algorithm. The actor network generates an action in the limited action space. The PID controller uses the double-delay deep deterministic policy gradient deep reinforcement learning algorithm agent action to estimate the internal state of the battery, and then uses the dynamic composite reward function with adaptive constraints to determine the reward value of the quality of the behavior. These transformations are stored in the circular buffer R. When enough experience is collected, the agent is trained. During the training process, small samples of N cycles are sampled, and the critic is updated through the critic update expression. The actor is updated after d iterations according to the policy gradient method. The target network is updated according to the soft update method. The iterative process continues to the maximum position N e .

[0132] Finally, experimental and real vehicle data are collected to verify the accuracy and versatility of the estimation method.

[0133] like Figure 8 As shown, before obtaining battery data through experiments, all batteries are capacity tested and screened, and batteries with capacity fluctuations less than 3% are used as experimental objects to obtain the battery's OCV-SOC relationship and battery parameters;

[0134] The OCV-SOC relationship of the battery is obtained by constant current constant voltage and trickle charge and discharge method. First, the battery is fully charged, and then discharged with a low rate current of C / 30 until they reach the lowest voltage. The battery is left to stand for two hours every 10% SOC to obtain the open circuit voltage of the battery during discharge. The battery is charged to the maximum voltage using the same low rate current of C / 30. The battery is left to stand for two hours every 10% SOC to obtain the open circuit voltage of the battery during charging. Since the extremely low charge and discharge current greatly weakens the hysteresis effect and polarization effect, the battery terminal voltage during charging and discharging is:

[0135]

[0136] Among them, V o,ch Indicates the battery terminal voltage during charging, V o,dis Indicates the battery terminal voltage during discharge;

[0137] The open circuit voltage of the battery can be calculated by taking the average value of the battery terminal voltage during charging and the battery terminal voltage during discharging:

[0138]

[0139] In order to determine the battery parameter values, the battery is placed in a variety of temperatures for specific hybrid pulse power characteristics HPPC tests. Under each specific temperature HPPC test, the battery data is processed by curve fitting and optimization tools, and the battery parameter values ​​at different temperatures are calculated, including battery resistance and capacitance.

[0140] Error analysis methods include maximum absolute error, root mean square error, and mean absolute error;

[0141] The maximum absolute error expression is:

[0142]

[0143] in, and Represent the measured state value and the estimated state value respectively;

[0144] The root mean square error expression is:

[0145]

[0146] The mean absolute error expression is:

[0147]

[0148] In the above embodiments, the description's reference to "this embodiment" indicates that a particular feature, structure, or characteristic described in conjunction with the embodiment is included in at least some embodiments, but not necessarily all embodiments. Multiple occurrences of "this embodiment" do not necessarily all refer to the same embodiment.

[0149] In the above-described embodiments, although the invention has been described in conjunction with specific embodiments of the invention, many substitutions, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art based on the foregoing description. For example, other storage structures (e.g., dynamic RAM (DRAM)) may use the embodiments discussed. Embodiments of the invention are intended to encompass all such substitutions, modifications, and variations that fall within the broad scope of the appended claims.

[0150] This embodiment further provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, any one of the methods in this embodiment is implemented.

[0151] This embodiment also provides an electronic terminal, including: a processor and a memory;

[0152] The memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the terminal executes any one of the methods in this embodiment.

[0153] The computer-readable storage medium in this embodiment can be understood by ordinary technicians in this field: all or part of the steps of implementing the above-mentioned method embodiments can be completed by hardware related to the computer program. The aforementioned computer program can be stored in a computer-readable storage medium. When the program is executed, the steps of the above-mentioned method embodiments are executed; and the aforementioned storage medium includes: ROM, RAM, magnetic disk or optical disk and other media that can store program codes.

[0154] The electronic terminal provided in this embodiment includes a processor, a memory, a transceiver and a communication interface. The memory and the communication interface are connected to the processor and the transceiver and complete communication with each other. The memory is used to store computer programs, the communication interface is used to communicate, and the processor and the transceiver are used to run computer programs so that the electronic terminal executes each step of the above method.

[0155] In this embodiment, the memory may include a random access memory (RAM), and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.

[0156] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0157] The present invention can be used in many general or special computing system environments or configurations, such as personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like.

[0158] The present invention may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0159] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solution, which should be included in the scope of the claims of the present invention.

Claims

1. A highly adaptive lithium-ion battery state estimation method guided by deep reinforcement learning, characterized in that: The following steps are involved: S1: Select the second-order RC equivalent circuit model as the battery model, consider the augmented model deviation, build a PID controller, and estimate the battery state of charge SOC in real time; S2: Construct a double-delayed deep deterministic policy gradient deep reinforcement learning algorithm and a dynamic composite reward function with adaptive constraints; S3: Combine the PID controller with the double-delayed deep deterministic policy gradient deep reinforcement learning algorithm and perform model training. Then, the PID controller is automatically adjusted to match the real-time operating conditions based on the optimal strategy determined by learning.

2. The method for highly adaptive lithium-ion battery state estimation guided by deep reinforcement learning according to claim 1, characterized in that: The state space equation expression of the second-order RC model in step S1 is: Among them, C a is the charge transfer capacitance, R a is the charge transfer resistance, V a is the charge transfer polarization voltage, C b is the diffusion capacitance, R b is the concentration polarization resistance, V b is the voltage of the internal diffusion process of the battery, η i is the Coulomb efficiency, C t is the battery capacity, SOC is the battery state of charge, R ohm is the battery ohmic polarization internal resistance, V oc (s) is the open circuit voltage of the battery; Considering the augmented model deviation, the state space equation is transformed into the state space equation form: Among them, V mb is the model bias; The state space equation is organized into a state space expression: Among them, D = R ohm , y = V o , u = I, f(x) = V oc (soc) + V a + V b + V mb ; The PID controller expression is as follows: Among them, K p , K i and K d are proportional gain matrix, integral gain matrix and differential gain matrix respectively, g and h are the integral and differential of system error respectively, K fi is the pre-gain for the control error integral.

3. The method for highly adaptive lithium-ion battery state estimation guided by deep reinforcement learning according to claim 1, characterized in that: The dual-delay deep deterministic policy gradient deep reinforcement learning algorithm adopts an actor-critic structure, in which the estimate of Q(s,a) uses the critic approximation function Q θ (s,a) and adjustable parameters θ, using a secondary frozen target network Q θ' (s,a) is updated, and the target y is kept constant through multiple updates. The expression is: y=r+γQ θ' (s',a') The dual-delay deep deterministic policy gradient deep reinforcement learning algorithm is built with six neural networks, a policy network With parameters For policy update, two value function networks critic are used to estimate Q(s,a), that is, Q with parameter θ1 θ1 and Q with parameter θ2 θ2 , each actor and critic corresponds to a target network, using Q θ1' and Q θ2' It also includes a double Q-learning algorithm, which takes the minimum of two estimates for target update. The update expression of the critic is: Where γ is the discount factor, r is the reward value, and N is the number of samples selected for conversion at each step; Improve the critic update expression and limit the target to a small range. The expression is: Among them, ε~clip(N(0,σ),-c,c), N represents the standard normal distribution, σ is the noise standard deviation, and c is the noise limit; Using the policy gradient method, the actor is updated by maximizing the Q function as follows: The soft update method is used to update the target network parameters as follows: θ'←(1-τ)θ'+τθ where τ∈(0,1) is the learning rate.

4. The method for highly adaptive lithium-ion battery state estimation guided by deep reinforcement learning according to claim 1, characterized in that: The steps for constructing the dynamic composite reward function with adaptive constraints are as follows: S21: The system output error is established based on the difference between the measured battery voltage and the estimated value. The expression is: r1=w1·(1-w adr )·|e| Among them, w1 is the weight coefficient, w adr is the adaptive weight; The battery state error function is established based on the battery output voltage and the internal dynamics of the battery. The expression is: r2=w2 w adr ·|e s | Among them, w2 is the weight coefficient; Establish an additional guided reward function, the expression is: <h2 style=";text-align:left;direction:ltr">r3 = w3|e|d1+w4|e<h2 style=";text-align:left;direction:ltr"> s <h2 style=";text-align:left;direction:ltr"> |·d2 Among them, w3 and w4 are the output weight and state error weight respectively, and the values ​​of d1 and d2 are 0 or 1. The specific value selection method is as follows: S22: The dynamic composite reward function of the adaptive constraint is determined as the sum of the system output error, the battery state error function and the additional guidance reward function: r=r1+r2+r3。 5. The method for highly adaptive lithium-ion battery state estimation guided by deep reinforcement learning according to claim 1, characterized in that: The step S3 combines the PID controller with the double-delayed deep deterministic policy gradient deep reinforcement learning algorithm, specifically including: The PID controller passes the system state space to the policy network actor and the value function network critic, and passes the batch processed data to the neural network layer of the double-delayed deep deterministic policy gradient deep reinforcement learning algorithm; the policy network passes the relevant information to the PID controller; finally, the PID controller automatically adjusts the PID controller to match the real-time working conditions based on the optimal strategy determined by learning; The state space of the PID controller output contains the normalized battery current I cr , the time derivative of current dI cr / dt, measured battery voltage V o , the time derivative of voltage dV o / dt, estimated battery voltage Voltage output error e, voltage output error time derivative de / dt, estimated battery internal state and the time derivative of the battery internal state The state space representation of the PID controller output is: Among them, I cr =I / C t , Select PID control gain and adaptive reward gain as the action space, and the action space is expressed as: A s ={K p1 ,K p2 ,K p3 ,K p4 ,K i1 ,K i2 ,K i3 ,K i4 ,K d1 ,K d2 ,K d3 ,K d4 ,w adr } in, j=1,2,3,4 are the proportional gain, integral gain and differential gain under different battery states, 0<w adr <1 is the weight of the adaptive reward function; The battery model parameters are updated to the optimal values ​​at the current stage by minimizing the loss function. The expression of the minimization loss function is:

6. The method for highly adaptive lithium-ion battery state estimation guided by deep reinforcement learning according to claim 1, characterized in that: The model training process described in step S3 is as follows: The state space of the PID controller is used as the input of the double-delay deep deterministic policy gradient deep reinforcement learning algorithm. The actor network generates an action in the limited action space. The PID controller uses the double-delay deep deterministic policy gradient deep reinforcement learning algorithm agent action to estimate the internal state of the battery, and then uses the dynamic composite reward function with adaptive constraints to determine the reward value of the quality of the behavior. These transformations are stored in the circular buffer R. When enough experience is collected, the agent is trained. During the training process, small samples of N cycles are sampled, and the critic is updated through the critic update expression. The actor is updated after d iterations according to the policy gradient method. The target network is updated according to the soft update method. The iterative process continues to the maximum position N e .

Citation Information

Patent Citations

  • Model and data fusion driven enhanced vehicle-mounted battery state-of-charge estimation method

    CN118205445A

  • Rule and double depth q-network-based hybrid vehicle energy management method

    WO2022252559A1

Cited By

  • Battery management system

    CN120200355A

  • Lithium ion battery multi-state estimation method based on deep reinforcement learning of perceived physical information

    CN120652317A

  • Battery SOC online estimation method based on deep reinforcement learning and related equipment

    CN121522465A

  • Multi-target adaptive control method for new energy automobile battery pack thermal management system

    CN121608653A