Cooperative control method and device for energy storage system

By establishing a coupled physical model of electricity, heat, and aging and training a charging and thermal management agent, the problem of independent design of charging and thermal management in energy storage systems was solved, achieving coordinated control and improving system efficiency and battery life.

CN121485249APending Publication Date: 2026-02-06ZHEJIANG LEAPENERGY TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511612874.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-05
Publication Date
2026-02-06

AI Technical Summary

Technical Problem

In existing energy storage systems, the charging and thermal management systems are usually designed independently, resulting in a lack of coordination, which leads to energy loss and accelerated battery aging, and makes it impossible to achieve optimal energy and temperature control.

Method used

A coupled physical model of the energy storage system, involving electricity, heat, and aging, is constructed. By using a reinforcement learning environment model and a neural network, charging agents and thermal management agents are trained to achieve coordinated control of charging and thermal management.

Benefits of technology

The charging and thermal management of the energy storage system has been optimized, reducing energy loss, extending battery life, and improving the overall efficiency and safety of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121485249A_ABST
    Figure CN121485249A_ABST
Patent Text Reader

Abstract

The invention discloses a cooperative control method and device for an energy storage system, and belongs to the technical field of energy storage control, and the method comprises the steps: building an electricity-heat-aging coupling physical model of the energy storage system, and determining a reinforcement learning environment model according to the electricity-heat-aging coupling physical model; determining a charging agent and a thermal management agent according to the reinforcement learning environment model and a first preset neural network model; training the charging agent and the thermal management agent according to a preset model training algorithm to determine a target double-agent model; the target double-agent model is used for realizing cooperative control of charging and thermal management of the energy storage system. According to the method, the charging agent and the thermal management agent are constructed, and training optimization is performed on the charging agent and the thermal management agent through the preset model training algorithm, so that the optimal target double-agent model is obtained, and cooperative control of charging and thermal management of the energy storage system is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of energy storage control technology, specifically to a collaborative control method and device for an energy storage system. Background Technology

[0002] In existing technologies, battery charging management systems and thermal management systems are typically designed as independent subsystems. For example, the charging strategy focuses on quickly and safely charging the battery, while the thermal management strategy passively responds to the temperature rise caused by charging. This decoupled control mode leads to a lack of coordination: when formulating a high-current charging strategy, the charging system does not proactively consider the energy cost and capability limitations of the thermal management system in suppressing temperature rise; conversely, the thermal management system cannot actively influence the charging process to reduce heat generation at the source. This non-cooperative working mode results in the energy storage system being in a suboptimal state, causing unnecessary energy loss and accelerated battery aging. Summary of the Invention

[0003] A method and apparatus for coordinated control of an energy storage system are provided to achieve coordinated control of charging and thermal management of the energy storage system.

[0004] Firstly, a collaborative control method for an energy storage system is provided, including: Establish an electro-thermal-aging coupled physical model of the energy storage system, and determine the reinforcement learning environment model based on the electro-thermal-aging coupled physical model; The charging agent and the thermal management agent are determined based on the reinforcement learning environment model and the first preset neural network model. The charging agent and the thermal management agent are trained according to the preset model training algorithm to determine the target dual-agent model; the target dual-agent model is used to realize the coordinated control of charging and thermal management of the energy storage system.

[0005] In some embodiments, the reinforcement learning environment model includes global observation; the first preset neural network model includes a first gated recurrent unit, a second gated recurrent unit, a graph attention network, a first decision layer, and a second decision layer; The charging agent and thermal management agent are determined based on the reinforcement learning environment model and the first preset neural network model, including: First and second time-series features are obtained from global observations; wherein, the first time-series feature is the observation data sequence of the charging agent, and the second time-series feature is the observation data sequence of the thermal management agent. The first historical feature and the second historical feature are determined based on the first gated cycle unit, the second gated cycle unit, the first timing feature, and the second timing feature; wherein, the first historical feature is the historical state feature of charging, and the second historical feature is the historical state feature of thermal management. The charging agent and the thermal management agent are determined based on the first historical feature, the second historical feature, the graph attention network, the first decision layer, the second decision layer, the first temporal feature, and the second temporal feature.

[0006] In some embodiments, the first gated loop unit includes a first encoder; the second gated loop unit includes a second encoder; determining the first historical feature and the second historical feature based on the first gated loop unit, the second gated loop unit, the first timing feature, and the second timing feature includes: The first temporal feature is input into the first encoder to obtain the first historical feature, and the second temporal feature is input into the second encoder to obtain the second historical feature.

[0007] In some embodiments, determining the charging agent and the thermal management agent based on a first historical feature, a second historical feature, a graph attention network, a first decision layer, a second decision layer, a first temporal feature, and a second temporal feature includes: The target attention weight coefficients are determined based on the first historical feature, the second historical feature, the graph attention network, the preset norm function, and the first preset activation function. The target charging characteristics and target thermal management characteristics are determined based on the target attention weight coefficient, the second preset activation function, the first historical characteristics, and the second historical characteristics. The charging agent is determined based on the target charging characteristics, the first decision layer, and the first time-series characteristics; the thermal management agent is determined based on the target thermal management characteristics, the second decision layer, and the second time-series characteristics.

[0008] In some embodiments, determining the target attention weight coefficients based on a first historical feature, a second historical feature, a graph attention network, a preset norm function, and a first preset activation function includes: The importance of the target is determined based on the first historical feature, the second historical feature, the graph attention network, and the preset norm function; The target attention weight coefficient is determined based on the target importance and the first preset activation function.

[0009] In some embodiments, a charging agent is determined based on target charging characteristics, a first decision layer, and a first temporal characteristic; a thermal management agent is determined based on target thermal management characteristics, a second decision layer, and a second temporal characteristic, including: The first target action is determined based on the target charging characteristics and the first decision-making layer. The second target action is determined based on the target thermal management characteristics and the second decision layer; wherein, the first target action is used to characterize the charging target action, and the second target action is used to characterize the thermal management target action; The charging agent is determined based on the first target action and the first temporal characteristics; The thermal management agent is determined based on the second objective action and the second temporal characteristics.

[0010] In some embodiments, the charging agent includes a first target action, and the thermal management agent includes a second target action; The charging agent and the thermal management agent are trained according to a preset model training algorithm to determine the target dual-agent model, including: Obtain the energy flow direction during the charging of the energy storage system; A hybrid reward system is determined based on energy flow and global observation. The global evaluation dual-strategy network is determined based on the first historical feature, the second historical feature, the first target action, the second target action, and the second preset neural network model. The charging agent and the thermal management agent are trained based on a hybrid reward system, a global evaluation dual-policy network, and a pre-set training architecture to determine the target dual-agent model.

[0011] In some embodiments, the second preset neural network model includes a third encoder, a fourth encoder, an attention network, a third activation function, and an evaluation network; The global evaluation dual-strategy network is determined based on the first historical feature, the second historical feature, the first target action, the second target action, and the second preset neural network model, including: The first potential feature is determined based on the first historical feature, the first target action, and the third encoder; The second potential feature is determined based on the second historical feature, the second target action, and the fourth encoder; The target attention weights are determined based on the first latent feature, the second latent feature, the attention network, and the third activation function. The joint action value function is determined based on the target attention weight, the first latent feature, the second latent feature, and the evaluation network, and the global evaluation dual-policy network is determined based on the joint action value function.

[0012] In some embodiments, the hybrid reward system includes charging efficiency benefits, energy storage system health benefits, target benefits, charging local costs, and thermal management local costs.

[0013] Secondly, this application also provides a collaborative control device for an energy storage system, comprising: A module is established to create a coupled physical model of the energy storage system's electrical-thermal-aging processes. The first determination module is used to determine the reinforcement learning environment model based on the electro-thermal-aging coupled physical model; The second determining module is used to determine the charging agent and the thermal management agent based on the reinforcement learning environment model and the first preset neural network model. The third determining module is used to train the charging agent and the thermal management agent according to the preset model training algorithm to determine the target dual agent model; wherein, the target dual agent model is used to realize the coordinated control of charging and thermal management of the energy storage system.

[0014] Beneficial Effects: This application provides a collaborative control method and apparatus for an energy storage system. The collaborative control method includes: establishing an electro-thermal-aging coupled physical model of the energy storage system, and determining a reinforcement learning environment model based on the electro-thermal-aging coupled physical model; determining a charging agent and a thermal management agent based on the reinforcement learning environment model and a first preset neural network model; training the charging agent and the thermal management agent according to a preset model training algorithm to determine a target dual-agent model; wherein, the target dual-agent model is used to realize the collaborative control of charging and thermal management of the energy storage system. The collaborative control method for the energy storage system provided in this application constructs a charging agent and a thermal management agent, and trains and optimizes the charging agent and the thermal management agent through a preset model training algorithm to obtain the optimal target dual-agent model, thereby realizing the collaborative control of charging and thermal management of the energy storage system. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 This is a flowchart of a collaborative control method for an energy storage system provided in the embodiments of this application; Figure 2 This is a schematic diagram of a second-order equivalent circuit structure provided in the embodiments of this application; Figure 3 This is a schematic diagram of an agent policy network structure provided in the embodiments of this application; Figure 4 This is a schematic diagram of a charging strategy evaluation network structure within a global evaluation dual-strategy network provided in an embodiment of this application; Figure 5 This is a schematic diagram of a multi-threaded concurrent architecture provided in the embodiments of this application; Figure 6 This is a schematic diagram of the overall process of a collaborative control method for an energy storage system provided in the embodiments of this application; Figure 7 This is a schematic diagram of the model framework of a collaborative control method for an energy storage system provided in the embodiments of this application; Figure 8This is a schematic diagram of the collaborative control device for an energy storage system provided in the embodiments of this application. Detailed Implementation

[0017] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0018] In the description of this application, it should be understood that the terms "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientation or positional relationships based on the orientation or positional relationships shown in the accompanying drawings, are used only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this application. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of the stated features. In the description of this application, "a plurality of" means two or more, unless otherwise explicitly specified.

[0019] "A and / or B" includes the following three combinations: A only, B only, and a combination of A and B.

[0020] The use of "applies to" or "configured to" in this application implies open and inclusive language, which does not exclude the applicability to or configuration to devices performing additional tasks or steps. Additionally, the use of "based on" implies openness and inclusivity, because processes, steps, calculations, or other actions "based on" one or more of the stated conditions or values ​​may in practice be based on additional conditions or values ​​beyond those stated.

[0021] In this application, the term "exemplary" is used to mean "used as an example, illustration, or description." Any embodiment described as "exemplary" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use this application. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that this application can be made without using these specific details. In other instances, well-known structures and processes are not described in detail to avoid obscuring the description of this application with unnecessary detail. Therefore, this application is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed in this application.

[0022] Lithium-ion batteries, with their superior electrochemical characteristics such as high energy density, high operating voltage, long cycle life, and low self-discharge rate, have become indispensable core energy storage components in various high-power, high-energy-demand applications. However, their performance, lifespan, and safety are highly sensitive to operating conditions. Any operation deviating from their optimal operating range, such as excessively high operating temperature, excessively high charging cut-off voltage, or excessive operating current, can trigger irreversible internal side reactions, leading to capacity decay, increased internal resistance, and even thermal runaway in extreme cases, posing serious safety risks. Currently, the industry-standard lithium-ion battery charging method is a static open-loop control strategy with a constant current-constant voltage protocol. This involves charging with a constant current initially, then switching to a constant voltage charging mode once the voltage reaches its upper limit until the current decays to a preset threshold. This constant current and voltage threshold cannot dynamically adapt to the real-time characteristic evolution of the battery due to aging, changes in ambient temperature, and differences in initial state. This limitation is particularly pronounced in applications pursuing high efficiency and high-rate charging, often resulting in excessively long charging times or hidden damage to the battery.

[0023] Due to unpredictable changes in conditions and nonlinear factors such as fluctuating market prices, mathematical models can provide a more accurate description, but they are limited by computational costs, making them less suitable for real-time applications. Therefore, low-cost real-time algorithms are of great value.

[0024] Research combining reinforcement learning and deep learning is more suitable for real-time scenarios. Reinforcement learning does not require a predefined system model and can adapt to constantly changing conditions and uncertainties. Optimizing the real-time charging and discharging schedule of electric vehicles based on model-free deep reinforcement learning (DRL) algorithms can be represented as a Markov decision process (MDP), enabling adaptive control and achieving a better balance between speed and lifespan.

[0025] Lithium-ion battery charging technologies are essentially static open-loop control methods, whose control parameters cannot adaptively adjust to the dynamic evolution of battery characteristics caused by aging and changes in ambient temperature. Linear programming control attempts to find the optimal charging path by solving a linear objective function, but its performance is limited by model mismatch issues caused by linear simplification of the highly nonlinear battery system. Model predictive control uses a dynamic battery model to predict future states and iteratively optimizes the charging current within a finite time window, thus proactively handling constraints such as voltage and temperature. However, its performance is also heavily dependent on the accuracy of the model, and the online optimization process places stringent demands on computational resources. Furthermore, these traditional methods optimize the charging current outside the boundaries of an independent, passively responding thermal management system.

[0026] Reinforcement learning-based charging strategies model the charging process as a Markov decision process, no longer relying on precise mechanistic models. However, their action space is limited to the charging current, treating the thermal management system as an environment to adapt to rather than an actively cooperating agent. Therefore, their optimization results are also limited to the local optimum of passive heat dissipation capabilities.

[0027] In view of this, embodiments of this application provide a collaborative control method and apparatus for an energy storage system. This application constructs a charging agent and a thermal management agent, and trains and optimizes the charging agent and the thermal management agent through a preset model training algorithm to obtain the optimal target dual agent model, thereby realizing the collaborative control of charging and thermal management of the energy storage system.

[0028] It should be noted that the energy storage system in this application embodiment includes battery systems, new energy vehicle battery systems, etc. For example, in the technical solution of this application embodiment, a battery is used as an example for description, and the same will not be repeated below.

[0029] Thermal management refers to the process of controlling system temperature through active or passive methods. In battery systems, thermal management aims to maintain battery temperature within its optimal operating range to ensure battery performance, safety, and lifespan. Effective thermal management strategies need to consider multiple aspects such as heat generation, heat transfer, and heat dissipation, and strike a balance between temperature control effectiveness, energy consumption, and system complexity.

[0030] Figure 1 This is a flowchart illustrating a coordinated control method for an energy storage system provided in an embodiment of this application. This method is applicable to energy storage control systems, enabling coordinated control of the charging and thermal management of the energy storage system. The method can be executed by a coordinated control device for the energy storage system, which can be implemented in software and / or hardware and can be configured within the processor or controller of the energy storage control system. Please refer to [link to relevant documentation]. Figure 1 The method includes the following steps: Step 110: Establish an electro-thermal-aging coupled physical model of the energy storage system, and determine the reinforcement learning environment model based on the electro-thermal-aging coupled physical model.

[0031] First, the coupled electro-thermal-aging physical model is derived from the battery's second-order equivalent circuit, two-state thermal model, and aging model considering capacity decay and rising internal resistance (hereinafter referred to as the aging model). The second-order equivalent circuit reflects the battery's electrical characteristics and is used to simulate its charge and discharge response. The two-state thermal model reflects the battery's thermodynamic behavior and is used to simulate its heat transfer relationships. The aging model reflects the battery's aging mechanism and is used to simulate the aging degradation trend.

[0032] Figure 2 This is a schematic diagram of a second-order equivalent circuit structure provided in an embodiment of this application. For an example, please refer to [link to example diagram]. Figure 2 The second-order equivalent circuit includes a first resistor. Second resistor First capacitor Second capacitor and internal resistance The electrical behavior of the battery is modeled using a second-order equivalent circuit, and its formula is as follows: ; ; ; ; in, This is the open-circuit voltage; The internal resistance represents the battery's instantaneous voltage response. First resistor Second resistor First capacitor Second capacitor and the first resistor First voltage at both ends and the second resistor The second voltage at both ends For RC parallel network; Terminal voltage, Here, I represents the battery capacity, state of charge (SOC), and current. The first resistor... Second resistor and internal resistance The unit is ohms; first capacitor Second capacitor and battery capacity The unit is farad; open circuit voltage First voltage Second voltage and terminal voltage The unit is volt; the unit of current I is ampere; the unit of state of charge is percentage; the values ​​of each resistor, capacitor, and open-circuit voltage are affected by the battery temperature and battery aging state, and are usually obtained by fitting experimental data.

[0033] The two-state thermal model divides the battery into two nodes: an internal core and an external surface. This allows for a more accurate description of the temperature difference between the battery's interior and surface, thus characterizing its thermal behavior. Specifically, the temperature change of the battery's internal core is influenced by internal heat generation and heat conduction to the surface, while the battery surface temperature is affected by ambient temperature, cell temperature, adjacent cell surface temperatures, and coolant temperature. The coolant outlet temperature is influenced by the absorption of heat from the battery and its own cooling effect.

[0034] The calculation formulas for the changes in battery core temperature over time, battery surface temperature over time, and battery coolant outlet temperature over time are as follows: ; ; ; in, For the core temperature of the battery, The surface temperature of the battery. The conduction resistance for heat transfer between the battery cell and the surface. The equivalent heat capacity of the battery core, For ambient temperature, The thermal resistance for convective heat transfer from the surface to the environment. It is the surface equivalent capacitance. The equivalent thermal resistance between batteries. The surface temperature of the neighboring battery. This refers to the inlet temperature of the coolant. The equivalent thermal resistance for heat transfer from the battery surface to the coolant. Coolant outlet temperature, For coolant quality, For quality flow, This refers to the specific heat capacity of the coolant.

[0035] The aging model decomposes capacity loss and internal resistance increase into three parts: an empirical constant factor, a temperature factor, and a usage factor. The formulas for capacity loss and internal resistance increase are modeled as follows: ; ; in, This is the amount of capacity loss. These are empirical constants obtained by fitting a large amount of battery aging experimental data. The activation energy for capacity decay, Let be the ideal gas constant. To accumulate throughput per ampere-hour, For time index, This is the increase in internal resistance. The activation energy is the increase in internal resistance.

[0036] Then, a reinforcement learning environment model is established based on the electro-thermal-aging coupled physical model, defining the global state and action space. The reinforcement learning environment model provides the following global observations: ; in, This represents the global state of the environment (i.e., global observation). This is the highest temperature of the battery pack. This is the lowest temperature for the battery pack. For battery pack temperature difference, This refers to the temperature at the coolant inlet. This refers to the coolant outlet temperature. For ambient temperature, This refers to the battery's SOC value. The target SOC value for the battery. Battery current, The water pump speed, For thermal management power consumption, Power consumed for charging This refers to the real-time electricity price.

[0037] Among them, the joint actions provided by the reinforcement learning environment model for: ; in, This is the charging current. For thermal management mode, For water pump flow rate, Request the temperature for the water inlet.

[0038] Based on actual physical constraints or requirements, each control variable is normalized, limiting its range to meet battery safety requirements. Among these, 0 represents stopping charging, and 1 represents the maximum allowable charging current obtained by looking up a table based on battery temperature and SOC. Therefore, the actual charging current is... ; Here, -1, 0, and 1 represent cooling, idle, and heating, respectively. 1, 2 and 3 correspond to different pump speed levels. The unit is degrees Celsius, which limits the temperature range of the coolant at the inlet.

[0039] Step 120: Determine the charging agent and the thermal management agent based on the reinforcement learning environment model and the first preset neural network model.

[0040] The first preset neural network model is a hybrid network structure of gated recurrent unit (GRU) and graph attention network (GAT).

[0041] GAT is an advanced graph neural network that introduces a self-attention mechanism, allowing the model to dynamically assign different importance weights to each neighbor of a node when processing graph-structured data. When aggregating neighbor information, GAT no longer simply takes an average, but performs a weighted sum based on the learned attention weights, enabling the model to focus on more relevant nodes and thus capture the complex dependencies between nodes in the graph more accurately and profoundly.

[0042] The first preset neural network model includes a first gated recurrent unit, a second gated recurrent unit, a graph attention network, a first decision layer, and a second decision layer.

[0043] Figure 3 This is a schematic diagram of an agent policy network structure provided in an embodiment of this application. For an example, please refer to [link to example]. Figure 3 The agent policy network comprises a dual-agent system consisting of a charging agent and a thermal management agent. The charging agent determines the charging current, aiming to complete the charging task as quickly and cost-effectively as possible while optimizing energy efficiency and mitigating battery degradation.

[0044] Among them, the observation vector of the charging agent With action space The definitions are as follows: ; ; The charging agent can partially observe local environmental information such as the highest and lowest temperatures of the battery pack, the battery SOC value, the target SOC value, and the real-time electricity price, and control the battery charging current to regulate the charging rate.

[0045] Among them, the observation vector of the thermal management agent With action space The definitions are as follows: ; ; The thermal management agent can partially observe local environmental information such as the highest battery pack temperature, the lowest battery pack temperature, the ambient temperature, the coolant inlet temperature, the coolant outlet temperature, the battery SOC value, the battery current, and the water pump speed. It controls three control vectors: thermal management mode, water pump speed, and inlet temperature request, to regulate the battery temperature.

[0046] In some embodiments, the reinforcement learning environment model includes global observation; the first preset neural network model includes a first gated recurrent unit, a second gated recurrent unit, a graph attention network, a first decision layer, and a second decision layer; determining the charging agent and the thermal management agent based on the reinforcement learning environment model and the first preset neural network model includes the following steps: Step 1: Obtain the first and second time-series features from global observations; wherein, the first time-series feature is the observation data sequence of the charging agent, and the second time-series feature is the observation data sequence of the thermal management agent.

[0047] Specifically, each agent's GRU receives partial global observations from the reinforcement learning environment model. Each component independently extracts its own temporal features (i.e., its own observation data sequence). For example, the GRU of the charging agent extracts its own temporal features from global observations. Extracting the first temporal features independently, for example in, Natural numbers (referred to as observation sequence) The GRU, a thermal management agent, observes from a global perspective. Independently extract the second temporal features, for example (referred to as observation sequence) ).

[0048] Step 2: Determine the first historical feature based on the first gated loop unit, the second gated loop unit, the first timing feature, and the second timing feature. Second historical features Among them, the first historical feature Historical state characteristics of charging, second historical characteristics This refers to the historical characteristics of thermal management.

[0049] Specifically, the first historical feature is obtained based on the first gated recurrent unit (GRU1) and the first temporal feature. The second historical feature is obtained based on the second gated cyclic unit (GRU2) and the second temporal feature. .

[0050] In some embodiments, the first gated loop unit includes a first encoder; the second gated loop unit includes a second encoder; determining the first historical feature and the second historical feature based on the first gated loop unit, the second gated loop unit, the first timing feature, and the second timing feature includes: inputting the first timing feature into the first encoder to obtain the first historical feature, and inputting the second timing feature into the second encoder to obtain the second historical feature.

[0051] Specifically, each agent's GRU encoder receives partial global observations from the reinforcement learning environment model. Each hidden state vector is independently extracted from its temporal features, and then these features are input into their respective encoders to obtain their final hidden state vectors, thus achieving temporally independent temporal encoding. For example, see [link to relevant documentation]. Figure 3 Observation sequence The input is fed into the first encoder of the GRU1, where it is encoded and outputs the final hidden state vector, i.e., the first historical features. Similarly, the observation sequence The input is fed into the second encoder of the GRU2, where it encodes the data and outputs the final hidden state vector, i.e., the second historical features. .

[0052] Step 3: Determine the charging agent and the thermal management agent based on the first historical feature, the second historical feature, the graph attention network, the first decision layer, the second decision layer, the first temporal feature, and the second temporal feature.

[0053] Specifically, the first and second historical features are fused through a graph attention network to integrate charging and thermal management. The fused result is then used by the first and second decision layers to make collaborative decisions on charging and thermal management, resulting in a charging agent and a thermal management agent.

[0054] In some embodiments, determining a charging agent and a thermal management agent based on a first historical feature, a second historical feature, a graph attention network, a first decision layer, a second decision layer, a first temporal feature, and a second temporal feature includes: determining a target attention weight coefficient based on the first historical feature, the second historical feature, the graph attention network, a preset norm function, and a first preset activation function; determining a target charging feature and a target thermal management feature based on the target attention weight coefficient, the second preset activation function, the first historical feature, and the second historical feature; determining a charging agent based on the target charging feature, the first decision layer, and the first temporal feature; and determining a thermal management agent based on the target thermal management feature, the second decision layer, and the second temporal feature.

[0055] The first preset activation function is the softmax function, and the second preset activation function is the sigmoid activation function.

[0056] Among them, the target charging feature is a charging feature that includes dynamic thermal management information. The target thermal management feature is a thermal management feature that includes dynamic charging information.

[0057] The system consists of two decision layers: a first decision layer for charging and a second decision layer for thermal management. Both decision layers are fully connected decision-making network layers.

[0058] In some embodiments, determining the target attention weight coefficient based on the first historical feature, the second historical feature, the graph attention network, the preset norm function, and the first preset activation function includes: determining the target importance based on the first historical feature, the second historical feature, the graph attention network, and the preset norm function; and determining the target attention weight coefficient based on the target importance and the first preset activation function.

[0059] Among them, the importance of the target includes the first importance. Second importance Third importance and the fourth importance Of these, the most important is This indicates the importance of thermal management status to the state of charge, second most important. This indicates the importance of the charging state to the thermal management state, ranking third in importance. This indicates the importance of the charging state itself, the fourth importance. This indicates the importance of the thermal management status itself.

[0060] For example, please refer to Figure 3 According to the first historical characteristics Second historical features Graph construction and attention interaction are performed. Graph construction connects independent analysis and collaborative decision-making, organizing the separate, unstructured agent information into a structured, physically intuitive graph. Its node feature matrix H is a concatenation of historical summary vectors, and edges E define the direction of communication between two agents. The node feature matrix H and edge E are as follows: ; ; The calculation process for GAT attention interaction includes attention scoring, normalization, and weighted aggregation. First, the first historical features... Second historical features The initial features of the two nodes are input into the GAT layer. The GAT layer utilizes a shared parameter. The linear mapping is used to increase the dimensionality of the feature, thereby enhancing the feature. Then, the enhanced features are concatenated and used as vectors. Mapping to real numbers, the resulting correlation coefficient characterizes the correlation between vertices, i.e., the importance of the target (including first importance). Second importance Third importance and the fourth importance ).

[0061] Among them, the first importance Second importance Third importance and the fourth importance The calculation formulas are as follows: ; ; ; ; Then, the target attention weight coefficients are determined based on the target importance and the first preset activation function. Specifically, GAT allocates attention weight coefficients by weighted aggregation of information from itself and its neighboring nodes to obtain the target attention weight coefficients. The target attention weight coefficients include the first target attention weight coefficients. Second target attention weight coefficient Third target attention weight coefficient and the fourth target attention weight coefficient Among them, the first target attention weight coefficient The attention weighting coefficients representing the thermal management state to the charging state, and the second target attention weighting coefficients. The attention weighting coefficients representing the charging state to the thermal management state, and the third objective attention weighting coefficients. The attention weight coefficient representing the charging state itself, and the attention weight coefficient representing the fourth target. This represents the attention weighting coefficient of the thermal management state itself.

[0062] Among them, the first target attention weight coefficient Second target attention weight coefficient Third target attention weight coefficient and the fourth target attention weight coefficient The calculation formulas are as follows: ; ; ; ; in, It is a non-linear activation function. It is an exponential function.

[0063] In some embodiments, determining target charging features and target thermal management features based on target attention weight coefficients, a second preset activation function, a first historical feature, and a second historical feature includes: determining target charging features based on a first target attention weight coefficient, a third target attention weight coefficient, a first historical feature, a second historical feature, and a second preset activation function; and determining target thermal management features based on a second target attention weight coefficient, a fourth target attention weight coefficient, a first historical feature, a second historical feature, and a second preset activation function.

[0064] Among them, target charging characteristics The calculation formula is: ; in, This is the Sigmoid activation function.

[0065] Among them, target thermal management features The calculation formula is: ; In some embodiments, determining a charging agent based on target charging characteristics, a first decision layer, and a first temporal characteristic, and determining a thermal management agent based on target thermal management characteristics, a second decision layer, and a second temporal characteristic, includes: determining a first target action based on the target charging characteristics and the first decision layer; determining a second target action based on the target thermal management characteristics and the second decision layer; wherein the first target action characterizes a charging target action, and the second target action characterizes a thermal management target action; determining a charging agent based on the first target action and the first temporal characteristic; and determining a thermal management agent based on the second target action and the second temporal characteristic.

[0066] Specifically, will and target thermal management characteristics The data is fed into their respective fully connected decision layer networks. That is, the target charging features are... Input the target thermal management characteristics into the charging decision layer. The input is to the thermal management decision layer. Based on this highly condensed and fully interactive information, the fully connected decision layer network outputs the final action values. Specifically, the charging decision layer outputs the first objective action (e.g., Gaussian current distribution and charging current), and the thermal management decision layer outputs the second objective action (e.g., water temperature distribution, water pump distribution, and thermal management action), thus achieving information fusion and collaborative decision-making. For example, when the battery temperature is too high, the charging agent should pay close attention to the state of the thermal management agent (e.g., "heat dissipation has reached its limit") and postpone the decision to charge with high current. Finally, the action space and observation vector of the charging agent are obtained based on the first objective action and the first temporal features, and the action space and observation vector of the thermal management agent are obtained based on the second objective action and the second temporal features.

[0067] Therefore, it can be seen that the embodiments of this application construct a charging-thermal management dual-agent system by using a hybrid network structure based on GRU-GAT and a reinforcement learning environment model. This system can achieve coordinated control of battery charging and thermal management, decoupling the complex joint control problem of charging and thermal management into a charging agent and a thermal management agent, thereby achieving a high level of coordinated cooperation.

[0068] Step 130: Train the charging agent and the thermal management agent according to the preset model training algorithm to determine the target dual agent model; wherein, the target dual agent model is used to realize the coordinated control of charging and thermal management of the energy storage system.

[0069] In some embodiments, the charging agent includes a first target action, and the thermal management agent includes a second target action; training the charging agent and the thermal management agent according to a preset model training algorithm to determine a target dual-agent model includes: acquiring the energy flow direction during charging of the energy storage system; determining a hybrid reward system based on the energy flow direction and global observation; determining a global evaluation dual-policy network based on a first historical feature, a second historical feature, a first target action, a second target action, and a second preset neural network model; and training the charging agent and the thermal management agent according to the hybrid reward system, the global evaluation dual-policy network, and the preset training architecture to determine the target dual-agent model.

[0070] The hybrid reward system sets completely consistent team goals for the two agents while also taking responsibility for their own costs, thus enabling them to learn to make complex and intelligent trade-offs and collaborate for a common goal.

[0071] The hybrid reward system includes charging efficiency benefits, energy storage system health benefits, target benefits, charging local costs, and thermal management local costs.

[0072] Among them, charging efficiency benefits, energy storage system health benefits, and target benefits are global variables, while charging local costs and thermal management local costs are local variables.

[0073] Specifically, the dual-agent hybrid reward system is constructed following the principles of shared objectives and independent costs. Shared rewards include energy strategies that optimize charging efficiency. Battery health benefits and delayed lifespan degradation Target revenue encourages charging targets and temperature control targets. Local costs are used to reward or penalize each agent for the direct costs of its own actions, including local charging costs. and local costs of thermal management The charging local cost includes electricity price costs to minimize user electricity expenses and current smoothing costs to encourage smooth control. The thermal management local cost includes thermal management energy consumption costs to minimize energy consumption.

[0074] Considering the energy flow and distribution during parking and charging, the total energy provided by the power grid The energy flowing to the on-board charger is Energy flowing to the on-board charger After conversion efficiency The converted energy, a portion of which flows to the thermal management system and other onboard equipment (ignored), is defined as the energy flowing to the thermal management system and other onboard equipment. Another portion is converted into direct current and supplied to the battery pack (i.e., the energy entering the battery pack). Energy entering the battery pack. It is further divided into effective energy that is truly stored in the form of chemical energy. And the energy dissipated as heat due to the battery's internal resistance This is the internal loss during the charging process. The overall goal of the system is to maximize revenue. Minimize all energy costs expended and .

[0075] Therefore, the construction of the hybrid reward system (i.e., the definition of the reward function) is as follows: ; ; ; ; ; ; in, All are constant coefficients corresponding to each item. This is the difference between the battery temperature and the ideal temperature range. The energy actually stored in the battery within a time step is the effective output of the entire system. This refers to the total energy capacity of the battery. The difference in SOC within a time step; The total electrical energy fed into the battery pack includes the energy that is ultimately stored. and loss in the form of heat ; Energy consumption for thermal management and These represent the percentage of capacity loss and the percentage of internal resistance increase at the point of maximum degradation, respectively, characterizing the degree of battery degradation. This represents the difference in charging current between the two points in time. Power consumption for thermal management; Power is consumed during charging.

[0076] In some embodiments, the second preset neural network model includes a third encoder, a fourth encoder, an attention network, a third activation function, and an evaluation network; determining the global evaluation dual-strategy network based on the first historical features, the second historical features, the first target action, the second target action, and the second preset neural network model includes: determining a first latent feature based on the first historical features, the first target action, and the third encoder; determining a second latent feature based on the second historical features, the second target action, and the fourth encoder; determining a target attention weight based on the first latent feature, the second latent feature, the attention network, and the third activation function; determining a joint action value function based on the target attention weight, the first latent feature, the second latent feature, and the evaluation network; and determining the global evaluation dual-strategy network based on the joint action value function.

[0077] The third activation function is the Softmax function.

[0078] Figure 4 This is a schematic diagram of the charging strategy evaluation network structure within a global evaluation dual-strategy network provided in this application embodiment. Specifically, the global evaluation dual-strategy network is designed based on an attention interaction mechanism. During training, the evaluation networks of the two agents observe the global state and output the expected value of the long-term reward obtained under the current state and current action to judge the quality of the joint action. It can automatically learn and allocate attention weights to achieve dynamic credit allocation among multiple agents.

[0079] The charging strategy evaluation network structure of the global evaluation dual-strategy network is as follows: Figure 4 As shown, firstly, through parallel encoder modules (i.e., the third and fourth encoders), the concatenated vectors of local historical observations and actions of each agent are encoded into latent features of a fixed dimension. For example, the first historical features of the charging agent are... and the first target action The first concatenation vector is obtained by concatenation. The second historical feature of thermal management agents Second target action ( The second concatenated vector is obtained by concatenating the two vectors. Then, the first concatenated vector... The input is fed into the third encoder, Encoder C, to obtain the first latent feature. , concatenate the second vector The input is fed into the fourth encoder, Encoder T, to obtain the second latent feature. Among them, the first potential feature Second potential feature They are respectively: ; ; Secondly, an attention module operates on the latent features of all agents and outputs a set of normalized target attention weights via a Softmax layer. ,in This weight The aim is to explicitly model the marginal contribution of an agent in the current joint state-action space. The target attention weights include the first attention weight. Second attention weight .

[0080] In some embodiments, determining a joint action value function based on target attention weights, a first latent feature, a second latent feature, and an evaluation network includes: determining a joint action value function based on the first attention weights, the second attention weights, the first latent feature, the second latent feature, and the evaluation network.

[0081] Specifically, through weighted summation This integrates the latent features of all agents into a global context representation. This characterization Ultimately, it is mapped to a scalar joint action value function by a neural network (such as a charging evaluation network). This estimates the long-term returns of the charging agent. Similarly, the thermal management strategy evaluation network uses attention modules of the same structure to act on latent features. , It also makes independent attention weight inferences and calculates joint action value functions to estimate the long-term returns of the thermal management agent.

[0082] In some embodiments, the charging agent and the thermal management agent are trained according to a hybrid reward system, a global evaluation dual-policy network, and a preset training architecture to determine the target dual-agent model.

[0083] The preset training architecture is a multi-agent deep deterministic policy gradient (MADDPG) algorithm, which adopts a centralized training framework with multi-scenario progression, noise simulation, parameter drift and domain randomization.

[0084] Specifically, the model training incorporates multi-scenario progression, noise simulation, parameter drift, and domain randomization. First, a multi-scenario progression approach is adopted, allowing the charging agent and thermal management agent to interact in an environment without an aging model, learning to balance energy saving and temperature control. Building upon the basic charging-temperature control linkage, a battery aging model and corresponding penalty terms are introduced, enabling the agents to learn to balance charging speed, temperature control intensity, and battery life. Furthermore, to address the risk of the model making decisions outside the training data distribution, noise simulation, parameter drift, and domain randomization are introduced. The process of introducing noise (e.g., sensor noise) simulation involves injecting random Gaussian noise into temperature and voltage / current readings in the simulation environment to improve the resilience of the dual agents' decisions. The process of introducing environmental parameter drift simulates the performance of the dual agents when internal model parameters deviate over time. The process of domain randomization introduces environmental models with varying parameters in different rounds, such as different temperatures, initial SOCs, or health states, to avoid overfitting and improve the robustness and generalization ability of the dual agents.

[0085] A dual-agent model is trained using the MADDPG training framework. Centralized value evaluation continuously adjusts the policy network parameters and attention weights of each agent, gradually bringing their collaborative strategy closer to the optimal level. First, the charging agent and the thermal management agent each initialize their policy-evaluation network, target policy-evaluation network, and a shared experience replay pool. During the execution phase, each agent makes independent decisions based on local observations. In the centralized training phase, the algorithm samples from the shared experience replay pool, calculates the loss using a hybrid reward and evaluation network to update the evaluation network, uses the policy gradient of the evaluation network to guide the update of the action network, and ensures convergence through a soft update mechanism in the target network, ultimately achieving efficient distributed collaborative decision-making. The evaluation network loss... Calculate and update gradients using the action network. and target network parameters , The updated formula is as follows, for the agent , For instant rewards, As a discount factor, For the next step of global observation, The next action calculated by the agent based on the next state. To evaluate the network, The target evaluation network has the same structure as the evaluation network. It receives the next global observation and the next action of all agents, and outputs the Q-value prediction. M is the sampling batch size. For the action policy network, the parameters are: Its input agent's local observation Output the action taken , To output the gradient of the action network with respect to its parameters, To evaluate the gradient of the network with respect to the agent's actions, where the actions are generated by the action network, For action networks, For average network parameters, For target action network parameter, Evaluation network for target parameter, This is the soft update coefficient.

[0086] ; ; ; ; ; Therefore, this embodiment employs a centralized training framework that incorporates multi-scenario progression, noise simulation, parameter drift, and domain randomization, combined with a hybrid reward system and a global evaluation dual-policy network to continuously train and optimize the dual-agent strategy. Thus, by adhering to the core paradigm of centralized training and distributed execution, this embodiment allows the charging agent and the thermal management agent to share global information during the training phase to learn complex collaborative strategies, thereby mitigating environmental non-stationarity and reputation allocation problems. Through multiple iterative training iterations, the optimal charging-thermal management dual-agent model (i.e., the target dual-agent model) is obtained.

[0087] After training and optimization, a robust and secure distributed execution and deployment framework is built based on a multi-threaded concurrent architecture. During the execution phase, each agent makes efficient distributed decisions based on local observations.

[0088] Figure 5 This is a schematic diagram of a multi-threaded concurrent architecture provided in an embodiment of this application. For an example, please refer to [link to example]. Figure 5This application achieves decoupling and efficient execution of its functional modules through a multi-threaded concurrent architecture. This architecture includes a communication transceiver module, a model inference module, and a real-time data stream visualization module. The communication transceiver module includes message filtering (e.g., identifier (ID) filtering based on real-time vehicle operating condition data sources), signal extraction, validity judgment, decoding and updating, message decoding, and historical pool management. This module is responsible for receiving data related to the vehicle's infotainment system bus and performing preprocessing operations such as filtering, cleaning, decoding, conversion, and storage to ensure higher quality data for subsequent stages.

[0089] The model inference module includes operations such as condition judgment, observation vector generation, time series representation, feature extraction, model inference, and encoding and transmission. The model inference module is responsible for analyzing and calculating battery state data under different operating conditions based on a deep reinforcement learning model to send optimal charging and thermal management control signals.

[0090] The real-time data stream visualization module includes a lightweight backend data stream, indicator statistics and evaluation, and a frontend real-time visualization. Based on lightweight data transmission and reception between the frontend and backend, the module is responsible for displaying real-time battery status changes through the frontend page and providing overall status display and indicator statistics after the operating conditions are completed.

[0091] It should be noted that this multi-threaded concurrent architecture control system supports switching between the physical controller area network (CAN) bus and the virtual socket controller area network (SocketCAN) interface, facilitating offline simulation testing and online real-vehicle deployment. The system's startup and shutdown processes are well-designed, enabling secure management of thread lifecycles and hardware resources, ensuring stability and reliability under complex automotive operating conditions.

[0092] Figure 6 This is a schematic diagram of the overall process of a collaborative control method for an energy storage system provided in the embodiments of this application. Figure 7 This is a schematic diagram of the model framework of a collaborative control method for an energy storage system provided in an embodiment of this application. For an example, please refer to [link / reference needed]. Figure 6 and Figure 7 The overall implementation process of the collaborative control method for this energy storage system is as follows: Step 1: Construct a reinforcement learning environment model based on the battery electro-thermal-aging coupled physical model. In this embodiment, a coupled physical model is established based on a second-order equivalent circuit model, a two-state thermal model, and an aging model considering capacity decay and rising internal resistance. A reinforcement learning environment model is then built based on this coupled physical model. The global observations of the reinforcement learning environment model include the battery pack's highest temperature, lowest temperature, temperature difference, coolant inlet temperature, coolant outlet temperature, and ambient temperature. The action space includes charging current, thermal management mode, water pump flow rate, and inlet temperature request.

[0093] Step two involves designing a GRU-GAT network and constructing a charging-thermal management dual-agent system. This application sets the charging agent observation as... The action is Thermal management intelligent agent observation is The action is In the GRU-GAT network, the input is historical data information with 10 steps, and the output is a Gaussian distribution of current and inlet water temperature, a discrete distribution of pattern and pump level, and random sampling is performed.

[0094] Step 3: Design a dual-agent hybrid reward system and a global evaluation dual-policy network. Based on energy flow, multi-objective coupling, and attention interaction, a hybrid reward system and a global evaluation dual-policy network are designed. The reward function of this application includes shared rewards for charging efficiency, lifetime degradation, charging, and temperature control objectives; local penalties for charging agents related to charging cost and smooth control; and local penalties for thermal management agents related to energy consumption control. This application designs a global evaluation dual-policy network structure based on an attention interaction mechanism, taking global state observations as input and outputting the expected value of long-term rewards.

[0095] Step four involves centralized training based on the optimized MADDPG algorithm. A centralized training framework employing multi-scenario progression, noise simulation, parameter drift, and domain randomization is used to continuously optimize the dual-agent strategy. This application uses the MADDPG algorithm to optimize the decision network.

[0096] Step 5: Construct a robust and secure distributed execution deployment framework based on a multi-threaded concurrent architecture. This application employs CAN_FD (Controller Area Network) communication with flexible data rate and Intel little-endian encoding.

[0097] It should be noted that this application utilizes the MADDPG multi-agent algorithm to train the agents, but does not limit the specific algorithm. This application designs a GRU-GAT network to implement information interaction between policies, but does not limit the specific network structure that enables interaction. This application designs a multi-dimensional reward function to guide the agents to reduce energy consumption and lifespan degradation while completing charging and thermal management tasks, but does not limit the specific items of the reward or penalty and their weights.

[0098] The embodiments of this application can solve the following technical problems: First, the relevant technologies suffer from a lack of coordination between charging and thermal management strategies: Within the existing technological framework, battery charging management systems and thermal management systems are typically designed as independent subsystems. The charging strategy focuses primarily on quickly and safely charging the battery, while the thermal management strategy passively responds to the temperature rise caused by charging. This decoupled control model leads to a lack of coordination: when formulating high-current charging strategies, the charging system does not proactively consider the energy cost and capability limitations of the thermal management system in suppressing temperature rise; conversely, the thermal management system cannot actively influence the charging process to reduce heat generation at the source. This non-cooperative operating mode results in the system as a whole being in a suboptimal state, causing unnecessary energy loss and accelerated battery aging.

[0099] Secondly, there is the issue of static and non-adaptive control strategies under varying operating conditions: Battery management strategies in related technologies heavily rely on experience-based rules and offline-calibrated fixed parameters. However, the optimal operating range of a battery is a dynamic variable influenced by both internal and external factors. Traditional static strategies cannot make online, adaptive adjustments to these time-varying boundary conditions, leading to significant performance degradation under non-ideal conditions such as hot summers, cold winters, or the end of the battery's lifespan, failing to achieve optimal control across all operating conditions and the entire battery lifespan. Data-driven reinforcement learning-based solutions can achieve predictive, self-learning, adaptive optimization control, offering greater flexibility and robustness.

[0100] Third, existing control strategies rely on preset static modes, making it difficult to perform online, non-fixed multi-objective dynamic optimization based on real-time changing operating conditions and demands among conflicting factors such as charging speed (user experience), energy efficiency (operating costs), and battery life (asset value). Reinforcement learning schemes, by assigning weights to multiple objectives through a reward system to find and operate the optimal solution under the current conditions, can achieve a multi-objective dynamic optimization control scheme.

[0101] It is understood that the embodiments of this application combine the collaborative control of charging and thermal management through dual-agent cooperation. By introducing a graph neural network into the policy network to model the interaction between the two agents, an attention mechanism is used to dynamically capture the weight allocation of key information, a hybrid reward mechanism guides the optimization direction, and training is conducted under the CTDE paradigm to achieve the goal of collaborative control. Finally, a reliable deployment architecture ensures the feasibility of engineering applications, thereby balancing aging rate and energy efficiency to achieve an integrated and globally optimal charging-thermal management strategy. This application aims to establish a unified electric-thermal collaborative optimization model by deeply coupling charging and thermal management strategies. This electric-thermal collaborative optimization model can dynamically optimize within a complex performance space composed of charging speed, energy efficiency, battery life, safety constraints, and control smoothness, thereby maximizing the comprehensive performance and economic benefits of the lithium-ion battery system throughout its entire life cycle.

[0102] Figure 8 This is a schematic block diagram of a collaborative control device for an energy storage system provided in an embodiment of this application. This application also provides a collaborative control device for an energy storage system; please refer to [link / reference]. Figure 8 The collaborative control device 100 for the energy storage system includes: a modeling module 101 for establishing an electro-thermal-aging coupled physical model of the energy storage system; a first determining module 102 for determining a reinforcement learning environment model based on the electro-thermal-aging coupled physical model; a second determining module 103 for determining a charging agent and a thermal management agent based on the reinforcement learning environment model and a first preset neural network model; and a third determining module 104 for training the charging agent and the thermal management agent according to a preset model training algorithm to determine a target dual-agent model; wherein the target dual-agent model is used to realize the collaborative control of charging and thermal management of the energy storage system.

[0103] The technical solution of this application embodiment provides a collaborative control device for an energy storage system. By constructing a charging agent and a thermal management agent, and training and optimizing the charging agent and the thermal management agent through a preset model training algorithm, the optimal target dual agent model is obtained, thereby realizing the collaborative control of charging and thermal management of the energy storage system.

[0104] In some embodiments, the reinforcement learning environment model includes global observation; the first preset neural network model includes a first gated recurrent unit, a second gated recurrent unit, a graph attention network, a first decision layer, and a second decision layer; The second determining module 103 is further configured to: obtain a first temporal feature and a second temporal feature from global observations; wherein the first temporal feature is the observation data sequence of the charging agent, and the second temporal feature is the observation data sequence of the thermal management agent; determine a first historical feature and a second historical feature based on a first gated loop unit, a second gated loop unit, the first temporal feature, and the second temporal feature; wherein the first historical feature is the historical state feature of charging, and the second historical feature is the historical state feature of thermal management; and determine the charging agent and the thermal management agent based on the first historical feature, the second historical feature, the graph attention network, the first decision layer, the second decision layer, the first temporal feature, and the second temporal feature.

[0105] In some embodiments, the first gated loop unit includes a first encoder; the second gated loop unit includes a second encoder; the second determining module 103 is further configured to: input a first timing feature into the first encoder to obtain a first historical feature, and input a second timing feature into the second encoder to obtain a second historical feature.

[0106] In some embodiments, the second determining module 103 is further configured to: determine the target attention weight coefficient based on the first historical feature, the second historical feature, the graph attention network, the preset norm function, and the first preset activation function; The target charging characteristics and target thermal management characteristics are determined based on the target attention weight coefficient, the second preset activation function, the first historical characteristics, and the second historical characteristics. The charging agent is determined based on the target charging characteristics, the first decision layer, and the first time-series characteristics; the thermal management agent is determined based on the target thermal management characteristics, the second decision layer, and the second time-series characteristics.

[0107] In some embodiments, the second determining module 103 is further configured to: determine the importance of the target based on the first historical features, the second historical features, the graph attention network, and the preset norm function; The target attention weight coefficient is determined based on the target importance and the first preset activation function.

[0108] In some embodiments, the second determining module 103 is further configured to: determine a first target action based on the target charging characteristics and the first decision layer; The second target action is determined based on the target thermal management characteristics and the second decision layer; wherein, the first target action is used to characterize the charging target action, and the second target action is used to characterize the thermal management target action; The charging agent is determined based on the first target action and the first temporal characteristics; The thermal management agent is determined based on the second objective action and the second temporal characteristics.

[0109] In some embodiments, the charging agent includes a first target action, and the thermal management agent includes a second target action; The third determining module 104 is also used to: obtain the energy flow direction during the charging of the energy storage system; determine the hybrid reward system based on the energy flow direction and global observation; determine the global evaluation dual-strategy network based on the first historical feature, the second historical feature, the first target action, the second target action, and the second preset neural network model; and train the charging agent and the thermal management agent based on the hybrid reward system, the global evaluation dual-strategy network, and the preset training architecture to determine the target dual agent model.

[0110] In some embodiments, the second preset neural network model includes a third encoder, a fourth encoder, an attention network, a third activation function, and an evaluation network; The third determining module 104 is also used to: determine the first potential feature based on the first historical feature, the first target action, and the third encoder; The second potential feature is determined based on the second historical feature, the second target action, and the fourth encoder; The target attention weights are determined based on the first latent feature, the second latent feature, the attention network, and the third activation function. The joint action value function is determined based on the target attention weight, the first latent feature, the second latent feature, and the evaluation network, and the global evaluation dual-policy network is determined based on the joint action value function.

[0111] In some embodiments, the hybrid reward system includes charging efficiency benefits, energy storage system health benefits, target benefits, charging local costs, and thermal management local costs.

[0112] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0113] The above provides a detailed description of a collaborative control method and apparatus for an energy storage system provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A collaborative control method for an energy storage system, characterized in that, include: An electro-thermal-aging coupled physical model of the energy storage system is established, and a reinforcement learning environment model is determined based on the electro-thermal-aging coupled physical model. The charging agent and the thermal management agent are determined based on the reinforcement learning environment model and the first preset neural network model. The charging agent and the thermal management agent are trained according to a preset model training algorithm to determine a target dual-agent model; wherein the target dual-agent model is used to realize the coordinated control of charging and thermal management of the energy storage system.

2. The method according to claim 1, characterized in that, The reinforcement learning environment model includes global observation; the first preset neural network model includes a first gated recurrent unit, a second gated recurrent unit, a graph attention network, a first decision layer, and a second decision layer; The step of determining the charging agent and the thermal management agent based on the reinforcement learning environment model and the first preset neural network model includes: A first temporal feature and a second temporal feature are obtained from the global observations; wherein, the first temporal feature is the observation data sequence of the charging agent, and the second temporal feature is the observation data sequence of the thermal management agent; The first historical feature and the second historical feature are determined based on the first gated loop unit, the second gated loop unit, the first timing feature, and the second timing feature; wherein, the first historical feature is the historical state feature of charging, and the second historical feature is the historical state feature of thermal management; The charging agent and the thermal management agent are determined based on the first historical feature, the second historical feature, the graph attention network, the first decision layer, the second decision layer, the first temporal feature, and the second temporal feature.

3. The method according to claim 2, characterized in that, The first gated loop unit includes a first encoder; the second gated loop unit includes a second encoder; the step of determining the first historical feature and the second historical feature based on the first gated loop unit, the second gated loop unit, the first timing feature, and the second timing feature includes: The first temporal feature is input into the first encoder to obtain the first historical feature, and the second temporal feature is input into the second encoder to obtain the second historical feature.

4. The method according to claim 2, characterized in that, The step of determining the charging agent and the thermal management agent based on the first historical feature, the second historical feature, the graph attention network, the first decision layer, the second decision layer, the first temporal feature, and the second temporal feature includes: The target attention weight coefficient is determined based on the first historical feature, the second historical feature, the graph attention network, the preset norm function, and the first preset activation function; The target charging characteristics and target thermal management characteristics are determined based on the target attention weight coefficient, the second preset activation function, the first historical characteristics, and the second historical characteristics. The charging agent is determined based on the target charging characteristics, the first decision layer, and the first time-series characteristics; the thermal management agent is determined based on the target thermal management characteristics, the second decision layer, and the second time-series characteristics.

5. The method according to claim 4, characterized in that, The step of determining the target attention weight coefficient based on the first historical feature, the second historical feature, the graph attention network, the preset norm function, and the first preset activation function includes: The importance of the target is determined based on the first historical feature, the second historical feature, the graph attention network, and the preset norm function; The target attention weight coefficient is determined based on the target importance and the first preset activation function.

6. The method according to claim 4, characterized in that, The step of determining the charging agent based on the target charging characteristics, the first decision layer, and the first time-series characteristics, and determining the thermal management agent based on the target thermal management characteristics, the second decision layer, and the second time-series characteristics, includes: The first target action is determined based on the target charging characteristics and the first decision layer; A second target action is determined based on the target thermal management characteristics and the second decision layer; wherein, the first target action is used to characterize the charging target action, and the second target action is used to characterize the thermal management target action; The charging agent is determined based on the first target action and the first timing feature; The thermal management agent is determined based on the second target action and the second timing feature.

7. The method according to claim 2, characterized in that, The charging agent includes a first target action, and the thermal management agent includes a second target action; The step of training the charging agent and the thermal management agent according to a preset model training algorithm to determine the target dual-agent model includes: Obtain the energy flow direction during the charging of the energy storage system; A hybrid reward system is determined based on the energy flow direction and the global observation. A global evaluation dual-strategy network is determined based on the first historical feature, the second historical feature, the first target action, the second target action, and the second preset neural network model; The charging agent and the thermal management agent are trained based on the hybrid reward system, the global evaluation dual-policy network, and the preset training architecture to determine the target dual-agent model.

8. The method according to claim 7, characterized in that, The second preset neural network model includes a third encoder, a fourth encoder, an attention network, a third activation function, and an evaluation network; The step of determining the global evaluation dual-strategy network based on the first historical feature, the second historical feature, the first target action, the second target action, and the second preset neural network model includes: A first potential feature is determined based on the first historical feature, the first target action, and the third encoder; The second potential feature is determined based on the second historical feature, the second target action, and the fourth encoder; The target attention weights are determined based on the first latent feature, the second latent feature, the attention network, and the third activation function; The joint action value function is determined based on the target attention weight, the first latent feature, the second latent feature, and the evaluation network, and the global evaluation dual-policy network is determined based on the joint action value function.

9. The method according to claim 7, characterized in that, The hybrid reward system includes charging efficiency benefits, energy storage system health benefits, target benefits, charging local costs, and thermal management local costs.

10. A collaborative control device for an energy storage system, characterized in that, include: A module is established to create an electrical-thermal-aging coupled physical model of the energy storage system. The first determining module is used to determine the reinforcement learning environment model based on the electro-thermal-aging coupled physical model. The second determining module is used to determine the charging agent and the thermal management agent based on the reinforcement learning environment model and the first preset neural network model. The third determining module is used to train the charging agent and the thermal management agent according to a preset model training algorithm to determine the target dual agent model; wherein, the target dual agent model is used to realize the coordinated control of charging and thermal management of the energy storage system.