A lithium ion battery multi-state estimation method based on perception physical information deep reinforcement learning

CN120652317BActive Publication Date: 2026-09-08CHONGQING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510965402.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-14
Publication Date
2026-09-08
Estimated Expiration
2045-07-14

AI Technical Summary

Technical Problem

然而,现有方法存在显著局限:

Benefits of technology

[0029] (1) This invention significantly improves the accuracy and efficiency of multi-state collaborative estimation, completely overcoming the limitation of existing methods that only focus on a single state. By constructing a multi-input multi-output electrothermal coupling model, this model integrates the electrochemical dynamics and thermodynamic behavior of the battery, simultaneously capturing the strong coupling relationship between the state of charge, energy state, and temperature state, thus avoiding error accumulation. Combined with a proportional-integral-differential observer, it achieves joint optimization estimation of the internal temperature state, state of charge, and energy state of the battery, and significantly reduces computational complexity by utilizing the correlation between states.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120652317B_ABST
    Figure CN120652317B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of lithium ion battery multi-state estimation method based on perception physical information depth reinforcement learning, belong to battery management system technical field.For the precision insufficient and low robustness problem caused by single state estimation limitation, poor adaptability of manual parameter adjustment and core temperature neglect of existing method, technical scheme includes: first, establish multi-input multi-output electrothermal coupling model;Second, the observability of system is analyzed;Then, proportional-integral-derivative observer is constructed to jointly estimate battery temperature state, state of charge and energy state;Then, multi-constraint peak power state estimation strategy is formulated;The multi-state estimation framework is converted into multi-agent deep reinforcement learning control problem;Finally, the effectiveness of the method is verified by experimental data, hardware-in-the-loop test and real vehicle operation data.The technical effect is to improve the multi-state collaborative estimation accuracy, enhance the system robustness of intelligent control, ensure the safety of battery and prolong the service life, and realize high adaptability in diversified dynamic scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of battery management system technology and relates to a multi-state estimation method for lithium-ion batteries based on deep reinforcement learning of perceived physical information. Background Technology

[0002] Lithium-ion batteries, as the most widely used energy storage medium, are an indispensable core component in electrified transportation vehicles and renewable energy storage systems aiming for net-zero emissions. However, lithium-ion batteries possess complex nonlinear reaction dynamics and are accompanied by potential safety risks, necessitating an intelligent battery management system to ensure their safe and reliable operation. This system estimates key internal states, including state of charge, state of health, state of energy, and state of power, using measurable parameters. Accurate estimation of these states directly impacts the efficiency of management tasks such as charge / discharge control, battery balancing, thermal management, and power prediction.

[0003] Current battery state estimation algorithms are mainly divided into four categories: direct measurement methods, data-driven methods, electrochemical model-based methods, and equivalent circuit model-based methods. Among them, the equivalent circuit model method simulates the battery's electrochemical dynamic response through electrical components such as resistors and capacitors, achieving a good balance between model accuracy and complexity, and has become the mainstream approach. However, existing methods have significant limitations:

[0004] (1) Most methods are designed for a single state and cannot adapt to the complex dynamic characteristics of strong coupling of multiple states in real-world scenarios, making it difficult to meet the high requirements of battery management for robustness and accuracy.

[0005] (2) Due to the uncertainty, strong internal coupling and highly nonlinear behavior of the battery system under critical operating conditions, traditional methods are difficult to obtain analytical solutions for the control law.

[0006] (3) Existing joint estimation frameworks generally ignore the impact of battery core temperature on state estimation, which leads to a significant performance degradation in real-world environments.

[0007] (4) In the face of dynamic operation scenarios, manual parameter tuning is difficult to achieve effective control of the state estimator, which exacerbates the fluctuation of system performance.

[0008] Currently, a joint estimation framework integrating state of charge, energy, and temperature needs to be established to improve accuracy by leveraging the coupling relationships between states. A physically based intelligent control strategy needs to be introduced to address the adaptability issues of manual parameter tuning in complex scenarios. The core temperature must be incorporated into the estimation system to avoid thermal runaway risks and ensure system safety. Summary of the Invention

[0009] In view of this, the purpose of this invention is to provide a multi-state estimation method for lithium-ion batteries based on deep reinforcement learning of perceived physical information.

[0010] To achieve the above objectives, the present invention provides the following technical solution:

[0011] A multi-state estimation method for lithium-ion batteries based on deep reinforcement learning of perceived physical information, comprising the following steps:

[0012] S1: Establish an electrothermal coupling model for a lithium-ion battery with a multiple-input multiple-output (MIMO) structure for each battery cell;

[0013] S2: Perform observability analysis on the electrothermal coupling model to determine the observability of the system;

[0014] S3: Based on the model of S1, construct a proportional-integral-derivative (PID) observer to jointly estimate the battery's state of temperature (SOT), state of charge (SOC), and state of energy (SOE).

[0015] S4: Considering SOC, voltage, temperature and current constraints, formulate a multi-constraint peak power state (SOP) estimation strategy;

[0016] S5: Transform the multi-state estimation framework into a multi-agent deep reinforcement learning (DRL) control problem, extract physical quantities from the measurement data and input them into the DRL agent to learn the optimal policy;

[0017] S6: Verify the effectiveness of the method through experimental data, hardware-in-the-loop (HIL) testing, and real vehicle data.

[0018] Optionally, in S1, a second-order RC equivalent circuit model is established for lithium iron phosphate (LFP) batteries to describe the internal reaction process of lithium-ion batteries; considering the influence of temperature on battery operation, a two-state thermal model that takes into account the surface temperature and core temperature of the battery is established, and a lithium-ion battery electrothermal coupling model is coupled in place.

[0019] Optionally, in S2, the established lithium-ion battery model is highly nonlinear, and the system observability analysis considers the uncertainty of the system to calculate whether the observability matrix is ​​full rank, and judges the observability of the electrothermal coupling model under any condition.

[0020] Optionally, in S3, based on the constructed battery electrothermal coupling model, a MIMO PID observer is constructed to perform joint estimation of the battery cell SOT, SOC and SOE.

[0021] Optionally, in S4, the battery's power capability SOP is described by the available peak power, the battery voltage is kept within a specified range, the peak current is calculated under the constraints of voltage, SOC, temperature and a specified charge / discharge current range, and the maximum available peak power is estimated.

[0022] Optionally, in S5, a multi-agent DRL structure that takes into account the characteristics of perceived physical information is used to learn the optimal control of the state estimator. In the proposed framework, meaningful physical quantities, such as battery OCV, core temperature, and desired SOC neighbor values, are extracted from the measurement data and provided as input to the DRL agent along with the measured battery current and voltage. Physics-based constraints are also included in the design of the action space and reward function to help the agent effectively learn the optimal control strategy.

[0023] Optionally, in step S6, the verification scenario includes:

[0024] High-current dynamic operating conditions specifically include the Federal Urban Driving Schedule (FUDS), Dynamic Stress Test (DST), Federal Supplemental Test Procedure (US06), and Beijing Dynamic Stress Test (BJDST).

[0025] -5℃ low temperature environment test;

[0026] HIL platform dynamic operating condition simulation;

[0027] Verification through comparison with real vehicle operation data.

[0028] The beneficial effects of this invention are as follows:

[0029] (1) This invention significantly improves the accuracy and efficiency of multi-state collaborative estimation, completely overcoming the limitation of existing methods that only focus on a single state. By constructing a multi-input multi-output electrothermal coupling model, this model integrates the electrochemical dynamics and thermodynamic behavior of the battery, simultaneously capturing the strong coupling relationship between the state of charge, energy state, and temperature state, thus avoiding error accumulation. Combined with a proportional-integral-differential observer, it achieves joint optimization estimation of the internal temperature state, state of charge, and energy state of the battery, and significantly reduces computational complexity by utilizing the correlation between states.

[0030] (2) This invention innovatively introduces a deep reinforcement learning-driven intelligent control mechanism, effectively overcoming the adaptation defects of manual parameter tuning in dynamic scenarios. By extracting physical quantities such as open-circuit voltage, core temperature, and adjacent values ​​of the desired state of charge from measurement data as inputs, and combining a dual-agent interactive architecture, the action space of each agent is embedded in the state space of the other agent, achieving adaptive optimization of model parameters. The physical constraint reward function guides the agents to quickly converge to the optimal strategy that conforms to battery dynamics, ensuring robustness under complex operating conditions.

[0031] (3) This invention significantly enhances battery safety and lifespan assurance capabilities through a multi-constraint peak power state estimation strategy. This strategy comprehensively considers multi-dimensional boundaries such as state of charge, voltage, temperature, and current, calculates peak current and power under constraints, actively suppresses the risk of thermal runaway, and prevents overcharging and discharging. Under extreme conditions such as low temperature and high current, this mechanism maintains stable operation, ensuring the reliability of the battery management system under voltage fluctuations and thermal stress, and extending the overall service life.

[0032] Other advantages, objectives, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination, or may be learned from practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description

[0033] To make the objectives, technical solutions, and advantages of the present invention clearer, the preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, wherein:

[0034] Figure 1 This is the overall flowchart of the present invention;

[0035] Figure 2 This is a schematic diagram of the battery electrothermal coupling model in this invention;

[0036] Figure 3 This is a flowchart illustrating the observability verification of the electrothermal coupling model and the SOE estimation using a joint PID observer in this invention.

[0037] Figure 4 This is a flowchart of SOP estimation in this invention;

[0038] Figure 5 This is a simplified diagram of the working mechanism of the dual-agent DRL structure established in this invention;

[0039] Figure 6 This is a design illustration of the experimental verification part of the present invention. Detailed Implementation

[0040] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0041] The accompanying drawings are for illustrative purposes only and are schematic diagrams, not actual pictures. They should not be construed as limiting the invention. To better illustrate the embodiments of the invention, some parts in the drawings may be omitted, enlarged, or reduced, and do not represent the actual product dimensions. It is understandable to those skilled in the art that some well-known structures and their descriptions may be omitted in the drawings.

[0042] In the accompanying drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components. In the description of the present invention, it should be understood that if terms such as "upper," "lower," "left," "right," "front," and "rear" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, they are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, the terms used to describe positional relationships in the drawings are only for illustrative purposes and should not be construed as limiting the present invention. For those skilled in the art, the specific meaning of the above terms can be understood according to the specific circumstances.

[0043] Please see Figure 1 A multi-state estimation method for lithium-ion batteries based on deep reinforcement learning of perceived physical information includes the following steps:

[0044] S1: Establish an electrothermal coupling model for a lithium-ion battery with a multi-input multi-output (MIMO) structure for a single battery cell;

[0045] S2: Perform observability analysis on nonlinear battery systems to determine the observability of the electrothermal coupling model;

[0046] S3: Based on the established battery model, a PID observer is constructed to jointly estimate the battery's state of temperature (SOT), state of charge (SOC), and state of energy (SOE).

[0047] S4: Considering SOC, voltage, temperature and current constraints, formulate a multi-constraint peak power state (SOP) estimation strategy to ensure battery safety and extend battery life;

[0048] S5: Transform the multi-state estimation framework into a multi-agent DRL control problem, extract meaningful physical quantities from the measured battery data and provide them to the DRL agent to learn the optimal policy;

[0049] S6: Design experiments to demonstrate the practicality of the proposed design in various scenarios through experimental data, hardware-in-the-loop (HIL) test bench verification, and real vehicle operation data.

[0050] Please see Figure 2 In S1, a second-order RC equivalent circuit model is selected for modeling, where R a For ohmic internal resistance, R b For polarization internal resistance, C b For polarization capacitors, V oc If (s) is the open-circuit voltage and I is the input current, then the model can be expressed as:

[0051]

[0052] Where η represents the coulombic efficiency of a single battery cell, C t 's' represents the total battery capacity, and 's' represents the battery state of charge.

[0053] The thermal behavior of a battery can be described by the direct relationship between circuit elements and internal electrochemical phenomena, expressed by the following equation:

[0054]

[0055] Among them, R c R is the thermal resistance between the battery surface and the core. e C is the convective resistance between the battery surface and the external environment. c and C s T represents the heat capacity of the battery cell material and the battery casing material, respectively. c T represents the core temperature of the battery. s T represents the surface temperature of the battery. env The heat Q generated inside the battery cell is at ambient temperature. gen It can be given by the following formula:

[0056]

[0057] The first term represents the irreversible heat generated inside the battery due to ohmic losses. The second term represents the reversible heat generated by entropy change, T = (T c +T s ) / 2 is the average temperature. The entropy coefficient is obtained by integrating the thermal and electrical components of the model, expressed as V. o and T sIf the output is the system output, then this coupled model can be represented by the following multi-input multi-output state space:

[0058]

[0059] Where u = [IT env ] T The subscript T represents the temperature dependence of the cell parameters, which is updated in real time based on the temperature estimate. Therefore, the state-space expression of the model can be rewritten as:

[0060]

[0061] Please see Figure 3 In S2, for this battery electrothermal coupling model, its observability can be determined using the Lie derivative tool theorem. Let the system state equation be:

[0062]

[0063] The gradient matrix of the Lie derivative of this system is:

[0064] Q(x)=[dh1(x)...dL k Ax h1(x)dL B k h1(x)h2(x)...dh2L k Ax (x)] T

[0065]

[0066] L Ax k h(x) = L Ax (dL Ax k-1 h(x)); L B k h(x) = L b (dL B k-1 h(x))

[0067] By calculating the Lie derivative of the measured values, the observability matrix can be obtained as follows:

[0068]

[0069] in,

[0070] when When Q(x) is a full-rank matrix, the system is observable. The battery OCV is a nonlinear function of SOC, and has... If the condition holds true, then the electrothermal coupling model remains observable under any circumstances.

[0071] Please see Figure 3 In S3, the PID observer can be mathematically described as follows:

[0072]

[0073] Among them, K p =[K p1 K p2 K p3 K p4 K p5 ] T ∈R 5×1 K i =[K i1 K i2 K i3 K i4 K i5 ] T ∈R 5×1 and K d =[K d1 K d2 K d3 K d4 K d5 ] T ∈R 5×1 These are the proportional, integral, and differential gain matrices, respectively, where g represents K. fi Multiplying by the integral of the error signal, h represents the differential of the difference between the measured and true values ​​of the battery cell's output voltage. The battery's state of energy (SOE) is the remaining usable energy (E) of the battery. R With maximum available energy E M The ratio can be calculated using the following formula:

[0074]

[0075] Please see Figure 4 In S4, SOP refers to the battery power capability, which is described by the available peak power. Its calculation formula is as follows:

[0076]

[0077] in,

[0078] Peak current is the maximum charge / discharge current of the battery under safety constraints. It estimates the optimal available peak power by considering constraints such as SOC, voltage, temperature, and current, ensuring battery safety and extending battery life. Maintaining the battery voltage within the specified range (V) max V is the upper cutoff voltage. minIf the lower cutoff voltage is used, then the voltage-constrained peak current is:

[0079]

[0080] To maintain the battery's state of charge (SOC) within the specified range, the peak charge / discharge current constrained by the SOC is:

[0081]

[0082] To ensure that the peak current generated does not cause the lithium-ion battery to heat up beyond the upper temperature limit T. max The peak current, considering temperature constraints, can be calculated as follows:

[0083]

[0084] To ensure that the peak current is within the safe range of battery charging and discharging current specified by the battery manufacturer (I max(PD) I is the maximum current of the pulse discharge. max(Ch) Given the maximum charging current, the peak current considering current constraints can be calculated as follows:

[0085]

[0086] Please see Figure 5 In step S5, a multi-agent DRL structure with physically-informed features is used to learn the optimal control of the state estimator. In the proposed framework, meaningful physical quantities, such as battery OCV, core temperature, and desired SOC neighbor values, are extracted from measurement data and provided as inputs to the DRL agent along with measured battery current and voltage. Physics-based constraints are also included in the action space and reward function design to help the agent effectively learn the optimal policy. An improved DRL mechanism is implemented using two TD3 agents, where each agent's action space is contained within the state space of the other agent. Both agents effectively learn the optimal policy for their respective parts of the model.

[0087] Please see Figure 6In step S6, a battery experiment is designed, employing a high-current dynamic operating condition consisting of the Federal Urban Driving Cycle (FUDS), Dynamic Stress Test (DST), Federal Supplemental Test Procedure (US06), and Beijing Dynamic Stress Test (BJDST). The battery is placed in a -5°C temperature chamber with an initial SOC of 80%, subjected to a specified composite current condition. Then, under the same temperature conditions, a specified maximum rated continuous discharge current is applied to the battery. The cell temperature, peak power, and various internal states are measured under both conditions to evaluate and verify the effectiveness of the DRL mechanism. A hardware-in-the-loop (HIL) setup is developed to execute multiple verification tests. The HIL test bench connects the controller to real-time experimental data, simulating lithium-ion batteries in a real-world scenario. Under maintained temperature conditions, the cells are placed indoors, and with the aid of a battery tester, customized dynamic current conditions are tested to verify the practicality of the method under dynamic operating conditions. Based on the standard driving curves generated by the test bench and real-vehicle operating data, the method is systematically compared with the battery state estimation based on the physical information-based neural network PINN to verify the accuracy, robustness, and adaptability of the method.

[0088] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A multi-state estimation method for lithium-ion batteries based on deep reinforcement learning of perceived physical information, characterized in that: Includes the following steps: S1: Establish an electrothermal coupling model for a lithium-ion battery with a multi-input multi-output (MIMO) structure for a single battery cell; S2: Perform observability analysis on the electrothermal coupling model to determine the observability of the system; S3: Construct a proportional-integral-derivative PID observer based on the electrothermal coupling model to jointly estimate the battery temperature state SOT, state of charge SOC and state of energy SOE; S4: Considering SOC, voltage, temperature and current constraints, formulate a multi-constraint peak power state (SOP) estimation strategy; S5: The multi-state estimation framework is transformed into a deep reinforcement learning (DRL) control problem based on two TD3 agents, namely an electrical model TD3 agent and a thermal model TD3 agent; the battery open-circuit voltage, core temperature, and adjacent values ​​of the desired SOC are extracted from the measurement data and used as inputs to the two TD3 agents along with the measured battery current and voltage; physics-based constraints are set in the action space and reward function, and the action space of each TD3 agent is embedded into the state space of the other TD3 agent, so that the two TD3 agents learn the optimal policies of the electrical model part and the thermal model part of the electrothermal coupling model, respectively; S6: The effectiveness of the proposed lithium-ion battery multi-state estimation method is verified through experimental data, hardware-in-the-loop (HIL) testing, and real-vehicle data.

2. The multi-state estimation method for lithium-ion batteries based on deep reinforcement learning of perceptual physical information according to claim 1, characterized in that: In S1: A second-order RC equivalent circuit model is established for lithium iron phosphate (LFP) batteries, including ohmic internal resistance. Polarization internal resistance Polarized capacitors ; A two-state thermal model is established to describe the battery surface temperature. With core temperature Relationship: in To generate heat for the battery, , These are the core heat capacity and the outer casing heat capacity, respectively. For internal thermal resistance, For surface convection thermal resistance, The ambient temperature.

3. The multi-state estimation method for lithium-ion batteries based on deep reinforcement learning of perceptual physical information according to claim 2, characterized in that: In S2, observability is verified using the Lie derivative: Constructing the observability matrix ,when hour Full rank, system observable.

4. The multi-state estimation method for lithium-ion batteries based on deep reinforcement learning of perceptual physical information according to claim 1, characterized in that: In S3, the expression for the PID observer is: in , , These are the proportional gain matrix, integral gain matrix, and differential gain matrix, respectively. For the error integral term, For the error differential term, This is the state estimate. y is the output estimate, and y is the output measurement.

5. The multi-state estimation method for lithium-ion batteries based on deep reinforcement learning of perceptual physical information according to claim 1, characterized in that: In S4, the multi-constraint peak power state (SOP) estimation strategy includes: Calculation of peak current under voltage constraints: SOC-constrained peak current calculation: Temperature-constrained peak current calculation: Current constraints must satisfy .

6. The multi-state estimation method for lithium-ion batteries based on deep reinforcement learning of perceptual physical information according to claim 1, characterized in that: In S6, the verification scenarios include: High-current dynamic operating conditions, including Federal Urban Driving Cycle (FUDS), Dynamic Stress Test (DST), Federal Supplemental Test Procedure (US06), and Beijing Dynamic Stress Test (BJDST). -5℃ low temperature environment test; HIL platform dynamic operating condition simulation; And verification through comparison with actual vehicle operation data.